Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The project Entangled Histories used early modern printed normative texts. The computer used to have significant problems being able to read Dutch Gothic print, which is used in the vast majority of the sources. Using the Handwritten Text Recognition suite Transkribus (v.1.07-v.1.10), we reprocessed the original scans that had poor quality OCR, obtaining a Character Error Rate (CER) much lower than our initial expectations of <5% CER. This result is a significant improvement that enables the searching through 75,000 pages of printed normative texts from the seventeen provinces, also known as the Low Countries. The books of ordinances are compilations; thus, segmentation is essential to retrace the individual norms. We have applied – and compared – four different methods: ABBYY, P2PaLA, NLE Document Recognition and a custom rule-based tool that combines lexical features with font recognition. Each text (norm) in the books concerns one or more topics or categories. A selection of normative texts was manually labelled with internationally used (hierarchical) categories. Using Annif, a tool for automatic subject indexing, the computer was trained to apply the categories by itself. Automatic metadata makes it easier to search relevant texts and allows further analysis. Text recognition, segmentation and categorisation of norms together constitute the datafication of the Early Modern Ordinances. Our experiments for automating these steps have resulted in a provisional process for datafication of this and similar collections.
This paper develops a hierarchical account of concept formation that distinguishes the manifestation of originary qualitative feel from the public, norm-governed concepts through which subjects later articulate experience. Its central claim is not merely that concepts differ in abstractness, but that they differ in constraint profile: some remain tightly corrigible by recurrent felt manifestations delivered through relatively stable experiential channels, whereas others depend far more heavily on social reweighting, comparison, and narrative organization. To capture this difference, I distinguish four strata: (1) the manifestation of originary qualitative feel, (2) channel-bound concepts, (3) integrative judgments, and (4) narrative concepts. The pivotal case is smart. In some contexts, smart functions as a third-level integrative judgment anchored by a relatively stable cue-cluster such as rapid learning, flexible problem solving, and quick comprehension. In other contexts, the same lexical item is drawn into fourth-level use, where application depends on broader evaluations of which forms of knowledge, achievement, or tradition count as genuinely worthwhile. This shift shows that semantic drift does not abolish embodiment; it redistributes which embodied, affective, and inferential resources are recruited as operative sense changes. The paper then reorganizes its dialogue with contemporary theory around this problem of functional migration. Barsalou and Dove help explain embodied anchoring but do not sufficiently account for the reweighting of application rules within public evaluative space. Borghi and collaborators illuminate the role of language, sociality, inner speech, and conversational coordination in abstract concepts, yet their frameworks can be sharpened by distinguishing lexical stabilization from discursive reorganization and by connecting internal cue-stabilization with sentence-level and discourse-level renegotiation. Brandom and Jorem and Lohr clarify the public inferential dimension of concept use, but require a stronger account of differential experiential corrigibility. Searle helps explain how high-level concepts approach recognition and status, but not every narrative concept is thereby reduced to a full institutional fact. The conclusion is that the decisive issue is not whether a concept is social, but how lower-level cues, public norms, and narrative frames are combined in use.
The present study offers an original multi-level linguistic analysis of child speech in the 3-5 age group, combining a structured observational methodology with a qualitative-quantitative framework developed by the author. Unlike existing survey-based accounts, this paper proposes a three-tier Phonetic–Lexical–Grammatical (PLG) Developmental Profile as a practical diagnostic tool for assessing speech maturity in preschool children. The research was conducted on the basis of 120 spontaneous speech samples collected from 40 Kyrgyz-Russian bilingual children aged 3 to 5 years attending preschool institutions in Bishkek. The study identifies age-specific patterns of phonological simplification, lexical growth strategies, and syntactic complexity, and introduces a composite Speech Maturity Index (SMI) that integrates quantitative indicators across all three levels. The findings demonstrate that bilingual children in the studied cohort display a systematically accelerated grammatical development compared to monolingual norms while exhibiting specific phonological transfer patterns not previously described in the literature on Kyrgyz–Russian bilingualism at preschool age. The proposed PLG Profile and SMI are presented as replicable instruments suitable for use by preschool educators and speech therapists.
This article presents an extended linguocultural analysis of the emotion of joy as expressed through the concepts of “smile” and “laughter” in English, Russian, and Uzbek. The study integrates approaches from cognitive linguistics, cultural linguistics, pragmatics, and discourse analysis to investigate how emotional experience is encoded, structured, and interpreted across languages and cultures. Special attention is given to lexical semantics, phraseology, metaphorical models, and communicative norms. The findings reveal that while smile and laughter are universal physiological and emotional responses, their linguistic representation is shaped by culturally specific values, social expectations, and communicative traditions. English tends toward conventionalized politeness and frequent smiling, Russian demonstrates emotional authenticity and restraint, and Uzbek reflects collectivist warmth and hospitality. The article contributes to cross-cultural linguistics by demonstrating that emotional concepts are culturally mediated and linguistically structured. The research also has implications for intercultural communication, translation studies, and language teaching.
The present article examines the content and linguistic features of the realization of the concept of “happiness” in the English language. Drawing on semantic, pragmatic, and cognitive approaches, the study explores how happiness is encoded through lexical units, phraseological expressions, and contextual usage. Special attention is given to the interaction between language and culture, as well as to the role of discourse in shaping the interpretation of emotional concepts. The findings demonstrate that happiness is a multifaceted concept that reflects not only individual emotional states but also cultural values, social norms, and communicative intentions.
There has been a long felt need to investigate what the study of poetic language contributes to an understanding of ordinary language, and how posing the question in this way may indeed shift some of the assumptions about the way ordinary language works. Metaphor remains an important case in point: How do we get from “He kicked the bucket?” to “He died”? In this metaphorical idiom, the process is unavailable, without going into etymological hypotheses, because the meaning is already given – it’s lexicalized, a “dead metaphor” -- no semantic construal is involved and the literal meaning doesn’t play a role in understanding the meaning of the expression. A poet might break up the idiom and say, “He kicked the bucket and broke it to pieces” – defamiliarizing the automatic reading of the idiom as “he died” and opening up suggestions, for example, of overcoming death; thus lending plasticity to and making present both the figurative and literal meaning. But true access to the process of metaphorical meaning-making is only available in cases of novel poetic metaphor, where the domains brought together are often disparate. I argue that there are good reasons for seeing poetry, fiction and literature in general not as speech acts but as representations, or semblances of speech acts (Searle’s notion of a pretend speech act doesn’t do this justice). Poems stage dramatic situations in which both poetic speaker (that is, the lyrical “I,”) and addressee, if any, are part of a fictional world, rather than the biographical poets themselves addressing a reader. The often ironic gaps and dialogical tensions between the norms and values projected by the poem itself and those expressed by any one of the voices in the poem – including the lyrical ‘I’s perspective – have a crucial expressive force unto themselves. As Searle realized, the indirect speech act model for poetry does not work; whereas “Could you pass the salt?” is a request masquerading as a question, there is no such one-to-one relationship between the poetic utterance and a single determinate speech act hiding behind it. Searle’s and Austin’s approaches couldn’t fully succeed, on my view, because there are key aspects of the language-game of poetry that distinguish poetry from ordinary language speech acts. Let’s say these are three aspects of a Gricean poetic contract between speaker and hearer: 1) the distinction between the lyrical ‘I’ and the biographical poet; 2) what I call the "quasipropositionality" of utterance in poetry (and fiction as well); and 3) poetic form and the ways in which it makes meaning through textual strategies of mediation, like point of view, parody, irony, enjambment, rhyme, meter, rhythm, sound-play, allusion, and metaphor. These forms of aesthetic mediation make it very problematic to isolate the serious or not-serious utterance of a single-line as a primary semantic unit.
This paper takes a short sample of student spoken interaction and analyses it in detail to identify some of the features of the talk which may be at variance with what would be recognized as more proficient and fluent interaction.The goal is to identify points of interactional practice that can be judged as areas for consciousness raising and explicit instruction by the teacher.Several of the points raised in the analysis are suggested to stem from a complex set of influences including transfer of lexical, grammatical, and interactional practices from the L1.There also may be a habituation to classroom discourse when speaking in the L2 which is then used unconsciously as a template for non-institutional mundane social interactions.Recognition of the special nature of learner interactions alongside an understanding of the possible causes of this kind of speaking can, it is suggested, inform focused and empirically based teaching that develops learners' interactional competence.The competence versus performance duality posited by Chomsky (1965) is based indirectly on a notion of native speaker intuition.That is, when a native speaker of a language hears an utterance in that language, they can make an immediate judgment of its acceptability in grammatical terms.Deviations from the norm in
This article deals with the study of similarities and differences in expression of the correspondence of the norm to English and Tatar linguistic cultures using the material of phraseological units with a gender component. The paper defines the concept of norm in language and phraseology. One singles out the main groups of phraseological units according to whether they conform to the norm or not. The gender direction in linguistics, the subject of which is the interrelation of language and gender as a social factor, considers the concepts such as “genderâ€, “femalenessâ€, “malenessâ€. Gender is expressed in semantics and in the grammar of the language, forming a linguistic image of the world, which in turn depends on the conceptual image. The gender image of the world is not biologically determined, and the concepts of femaleness and maleness are determined by cultural and historical factors, in particular, by language stereotypes in different cultures and language communities. Gender metaphor also influences the formation of a conceptual and linguistic image of the world. A gender metaphor is understood as “the transfer not only of the physical, but also of the totality of spiritual qualities and properties, united by the nominations of femininity and masculinity to the objects that are not connected with sexâ€. In different language communities, femininity and masculinity referents often do not coincide, which creates difficulties in intercultural communication and translation.
This article explores the relationship between language and gender identity in English and Uzbek social media discourse. It examines how linguistic choices, including lexical items, pronouns, and stylistic markers, reflect and construct gender identities in online communication. By analyzing social media posts, comments, and interactions, the study highlights patterns of gendered language use and the cultural influences shaping these practices. The findings suggest that social media provides a dynamic space where users perform, negotiate, and challenge traditional gender norms through language. Cross-linguistic comparison reveals both similarities and differences in how English and Uzbek speakers encode gender identities in digital communication.
This article explores the linguocultural features of proverbs and sayings in three genetically unrelated languages: Kazakh, English, and Chinese. Proverbs and sayings reflect a nation’s worldview, spiritual values, historical experience, culture, and social norms. In this respect, they are not merely linguistic units, but complex linguocultural phenomena that represent the identity, mentality, and way of thinking of a particular ethnic group. Specifically, the article analyzes proverbs rooted in nomadic traditions in Kazakh culture, Anglo-Saxon pragmatism in English culture, and Confucian philosophy in Chinese culture using linguocultural and comparative analysis methods. This allows the identification of how each nation's worldview and cultural values are reflected in their proverbs. Additionally, the article focuses on specific themes across the three languages – such as family values, education, patience, and labor – and highlights culturally marked lexical units specific to each theme. The linguocultural features of the proverbs are compared based on a model developed from scholarly analysis. The concepts of “proverb” and “linguoculture” are clarified from a theoretical perspective. The scientific objective of the article is to identify, compare, and analyze the linguocultural characteristics of proverbs in Kazakh, English, and Chinese, and to determine their similarities and differences. The article concludes with suggestions for future research.
This study explores the construction and translation of the paradoxical identity in Sahar Khalifeh&rsquo;s novel &ldquo;The End of Spring&rdquo; and its English translation. Adopting a Descriptive Translation Studies (DTS) framework, the paper applies Gideon Toury&rsquo;s (1995) norm-based model to analyze how the inherent contradictions of Palestinian life under occupation are negotiated during translation. The analysis is conducted in two distinct phases: a micro-linguistic level focusing on operational norms, such as dialectal dissonance, semantic oxymorons, and lexical paradoxes, and a macro-conceptual level addressing preliminary and initial norms related to socio-political contradictions and religious ambivalence. Findings show a tension between Adequacy and Acceptability. Since the translator often employs Standardization to handle dialectal dissonance and uses titular oxymorons to improve target-culture fluency, the translation largely maintains the intense, authentic essence of internal stereotypes and metaphysical despair. According to Polysystem Theory, the study concludes that the English translation occupies a peripheral but innovative position within the Anglophone polysystem. By preserving the sharpest edges of Khalifeh&rsquo;s internal critiques and religious ambivalence, the text resists binary simplification and functions as a Primary Model of Paradox, presenting a multilayered, contradictory Palestinian identity within the Anglophone literary system, bridging the gap between the &ldquo;humanity&rdquo; experience and the &ldquo;labels&rdquo; imposed by conflict.
Terms denoting human body parts form one of the most archaic lexical layers, closely linked to sensory-functional aspects of human existence. They reflect anthropo-cultural specifics of representatives of various linguistic communities. This lexical group is called somatic – it characterizes human body elements and organism functions – and represents one of the most attractive lexical-semantic categories of Tungusic-Manchu languages. Somatic vocabulary is part of the basic dictionary stock formed over millennia, expressing not only the speakers' knowledge of the surrounding world but also their self-understanding and perception of their own body. Small aphoristic forms of Even folklore – riddles, prohibition-talismans, omens – attracted less researcher attention in past decades than prose narrative genres. Records of Even riddles from the 1860s and 1930s–1950s fell out of scientific analysis for a long time, and later editions included few new samples. Additionally, different aphoristic genres of Even folklore received unequal study: riddles, proverbs, customs were used for teaching children their native language and ethnic culture immersion, while prohibitions-talismans and omens, persisting among Evens today, remained little-known and became subjects of systematic collection only recently. The lexeme "head" in Even regularly acts as a key folklore component, reflecting traditional somatic concepts. The use of "head – dyl" in Even folklore, especially proverbs, riddles, omens, prohibitions, unites bodily and spiritual personality comprehension, emphasizing body-consciousness-social role links. Such constructions express ethnic identity and traditional body-spirit views. Somatic expressions significantly aid cultural norms and spiritual values transmission through language. The article defines somatisms and details proverbs, prohibitions, omens, riddles with "head" lexeme.
Social networks have become one of the most intensive environments for everyday written communication in Russian. Unlike traditional print media, social platforms combine speed, conversational interactivity, algorithmic visibility, and multimodal expression, thereby reshaping how users select words, build sentences, signal stance, and negotiate norms. This article examines the influence of social networks on the Russian language as a dynamic interaction among technological affordances, communicative practices, and socio-cultural values. Using a mixed design that integrates (a) discourse-linguistic observation of social media genres, (b) comparative analysis of forms typical for networked communication versus standard written Russian, and (c) interpretation within established frameworks of computer-mediated communication and sociolinguistics, the study synthesizes key tendencies of contemporary Russian online speech. The results indicate that social networks stimulate accelerated lexical innovation (slang, expressive neologisms, borrowings, and semantic shifts), normalize a hybrid “written-oral” style marked by compressed syntax and dialogic structures, intensify pragmatic markers of evaluation and identity, and expand punctuation and графическое оформление into a system of affective and interpersonal cues. At the same time, social networks also generate counter-trends: heightened metalinguistic reflection, new prescriptive micro-norms inside communities, and the diffusion of editorial practices through influencer culture and platform moderation. The discussion highlights that the influence of social networks is not a linear “degradation” of Russian but a reconfiguration of registers, where variability, expressive economy, and community norms coexist with standard language ideologies. The paper concludes that the most consequential change is not the emergence of isolated slang items but the stabilization of new communicative conventions that redefine the boundaries between colloquial and written Russian.
This study explores how gender is expressed both explicitly and implicitly in English paremiological units, particularly proverbs and sayings. Drawing on contemporary gender linguistics and an anthropocentric approach, the research views language as a reflection of socially and culturally constructed gender roles. Proverbs are treated as stable cultural artifacts that preserve traditional values, moral judgments and stereotypes accumulated over generations. The analysis reveals that English proverbs often reflect gender asymmetry, with a noticeable tendency toward male-centered perspectives rooted in historical and patriarchal social structures. At the same time, these linguistic units demonstrate both universal patterns of gender representation and culturally specific features shaped by the English-speaking context. Special attention is given to the ways gender is encoded through lexical choices, semantic nuances, and metaphorical associations. By combining qualitative interpretation with elements of linguistic statistical analysis, the study highlights how proverbs contribute to the construction and transmission of gender norms. Ultimately, the findings suggest that English paremiological units function as a complex intersection of language, culture, and ideology, preserving both enduring stereotypes and culturally specific understandings of gender relations.
Nigerian English (NigE) has developed into a unique variety of English, shaped by the interplay between speakers’ creative use of morphology and the influence of indigenous Nigerian languages. This study explores how NigE demonstrates morphological productivity and lexical borrowing, using a corpus-based approach to capture authentic language patterns. A carefully balanced corpus of 500,000 words was compiled from newspapers, online media, and recorded spoken interactions. Analyses focused on derivational processes, compounding, and the adaptation of loanwords, highlighting the strategies speakers employ to create new forms and meanings. The findings reveal that NigE exhibits robust morphological innovation, particularly in verb and noun formation, where affixation and compounding are frequently employed. Borrowed words, mainly sourced from Yoruba, Igbo, and Hausa, are often modified phonologically and morphologically to align with English norms, producing hybrid forms that enrich the NigE lexicon. This study underscores the dynamic relationship between English and indigenous languages in Nigeria, showing how speakers actively manipulate linguistic resources to meet social and communicative demands. The findings carry significant implications for sociolinguistic research, language teaching, and lexicography, advocating for recognition of NigE’s creative morphological processes in both academic study and pedagogical practice. By highlighting the innovative and adaptive nature of NigE, the study provides insights into how global English interacts with local linguistic ecologies.
The article analyzes the definitional status of the lexical unit "training" within the Russian education system. The authors examine the problem of terminological ambiguity arising from the absence of a clear legislative definition for this concept in the Federal Law "On Education in the Russian Federation." The semantic instability of the lexical unit (denoting both the process and the result) is transferred from professional discourse into the legal domain thus causing systematic logical conflicts within the law. Through analysis of key articles the study demonstrates how the inconsistent use of the concepts "student training," "quality of education," "educational activities," and "quality of training" complicates systematic legal regulation. The research includes comparative analysis of lexicographic sources, legal provisions, and theoretical works. It concludes that there is a necessity for the legislative consolidation of a consistent definition of the lexical unit "training" in order to eliminate legal uncertainty and ensure the correct application of norms in by-laws and local regulatory documents.
This article analyzes the relationship between gender linguistics and slang in English and Uzbek languages, focusing on how gender influences speech styles, lexical choices, and the formation of informal language. It examines the sociolinguistic factors that shape gendered communication patterns and explores how slang functions as a marker of identity, group belonging, and social interaction, particularly among younger speakers. The study also considers the impact of globalization, digital technologies, and social media platforms on the development and spread of slang in both linguistic contexts. Special attention is given to how traditional gender norms influence language use in Uzbek society, while English demonstrates comparatively more flexible and less rigid gender distinctions in informal communication. Furthermore, the article highlights the increasing convergence of slang usage across genders due to the influence of online communication, where linguistic boundaries are becoming more fluid. The comparative analysis reveals both similarities and differences in how gender and slang interact in English and Uzbek, showing that while cultural and social factors continue to shape language use, modern digital environments are gradually reducing traditional linguistic constraints.
DISSILEX is a controlled vocabulary in the form of a manually built lexico-semantic network of medieval Latin verbs and verbal expressions, featuring a detailed valency lexicon and connections to a large set of Latin and modern-English concepts, presented here as a single SQLite file (dissilex.db) that is readable by the sqlite3 command-line tool, any SQLite browser, or Python's built-in sqlite3 module. Rooted in the domain of inquisitorial records, DISSILEX covers general as well as more subject-specific meanings, with both standard (synonym, hypernym, etc.) and less canonical relations. DISSILEX is a product of Computer-Assisted Semantic Text Modelling (CASTEMO; Zbíral et al. 2026 - see README.md for full references), an approach to modeling statements as a four-slot structure of subject(s), predicate(s) and two objects, creating a thickly connected network of data points. Coverage is richest for human-interaction verbs (testimony, accusation, belief, religious practice) and the legal vocabulary of heresy trials, making it a machine-operable resource for modelling Latin textual data on dissent, resistance, repression, and resilience. We distinguish two entry types: Actions (verbs and verbal expressions, each associated with a three-slot valency frame specifying entity type, morphosyntactic, and semantic valencies) and Concepts (single- and multi-word expressions for other parts of speech). The network is connected through a set of 11 relation types, including superclass (hypernym) membership, synonymy, antonymy, verb-to-noun mappings, and valency-specific relations. Each relation connects two entities, and can be unidirectional or bidirectional. As part of an ongoing effort to position DISSILEX within the Linguistic Linked Open Data (LLOD) cloud, many entries contain IDs to external sources stored in the database, specifically to the LiLa (Linking Latin) Lemma Bank and Princeton WordNet (PWN) 3.0 and 3.1 synsets. We applied the Collaborative Interlingual Index (CILI) to map between the two versions of the PWN for entries where only one of the IDs has been added. We also indicate cases where no equivalent for a DISSILEX lemma exists ("NA"). Via the LiLa SPARQL endpoint, it is possible to use the linked LiLa lemmas, which feature as the central unit of linking sources in the Latin LLOD cloud, to retrieve data from several resources including dictionaries, corpora, treebanks, and various NLP tools. We have made use of this opportunity to enrich the database file with lemmas from the LiLa Lemma Bank, while also supplying LatinCy-generated lemmas for most Actions (model: la_core_web_lg). This release contains: dissilex.db: SQLite database, which can be readily queried dissilex_schema.md / dissilex_schema.pdf: schema documentation README.md: full dataset description, statistics, and SQL examples ATTRIBUTION.md: license and attribution notices. LICENSE-DATA: Full CC BY-SA 4.0 license text. Funding, attribution and licence DISSILEX is developed by the Dissident Networks research group (DISSINET) at Masaryk University and has received funding from the European Research Council (ERC) under the European Union’s Horizon 2020 research and innovation programme, grant agreement No. 101000442, project “Networks of Dissent: Computational Modelling of Dissident and Inquisitorial Cultures in Medieval Europe”, and from the European Regional Development Fund, grant agreement No. CZ.02.01.01/00/22_008/0004595, project “Beyond Security: Role of Conflict in Resilience-Building”. DISSILEX is released under CC BY-SA 4.0. It incorporates data from external resources: LiLa Lemma Bank: CIRCSE, Università Cattolica del Sacro Cuore (Milan). Licensed CC BY-SA 4.0. The database redistributes a subset of LiLa lemma forms (subset-selected and format-converted, not otherwise modified); the ShareAlike clause is honoured by this release's CC BY-SA 4.0 licence. URL: https://lila-erc.eu/ Princeton WordNet 3.0 / 3.1: We distribute offset IDs for both versions, and gloss text from WordNet 3.0 only. (An entry with a 3.1 identifier carries the corresponding 3.0 gloss, mapped via CILI.) WordNet 3.0 Copyright 2006 by Princeton University. All rights reserved. WordNet License. https://wordnet.princeton.edu/ LatinCy/spaCy: We redistribute output from the LatinCy model. The `spacy_lemma` field contains lemmas generated by the LatinCy spaCy pipeline `la_core_web_lg` (Patrick J. Burns). The model is MIT-licensed ( ); the spaCy library is MIT-licensed ( ). Full notices, including the Collaborative Inter-Lingual Index (CILI) and Latin WordNet, are in ATTRIBUTION.md. Version 1.1.0 is mostly a quality improvement of existing entries, but also adds and removes entries and links to external resources. The schema is unchanged, so queries written against version 1.0.0 continue to work.
Raw data including responses from 200 participants (100 musicians) on perceived emotion, as well as valence and arousal ratings, for 116 musical stimuli. The dataset also includes responses to the Beck Anxiety Inventory (BAI), Beck Depression Inventory (BDI), and a musical background questionnaire. The dataset further includes the musical stimuli selected and validated for the Musical Emotion Evaluation Test (MEET), supporting transparency, replication, and reuse.
The COVID-19 pandemic has radically transformed the English language introducing new lexical terms, semantic changes, and pragmatic norms that are still effective in the post-pandemic world. This scientific review explores new tendencies in writing English after the pandemic and concentrates on the definition of vocabulary, discourse, and pragmatic adjustments in social, professional, educational, and digital environments. The paper summarizes recent linguistic studies to examine the way in which communication driven by crisis made the formation of neologisms, borrowing, compounding, and semantic recontextualization faster. Keywords of health, risk, remote interaction, and digitalization became used not only in specialized registers but also daily and the meaning of already existing words has been broadened or metaphorical. In addition to the use of lexis, the review identifies shifts in pragmatic practices such as the changes in politeness strategies, the manifestation of uncertainty and empathy, and the taming of crisis sensitive discourse in institutional and interpersonal communication. The heightened use of digital platforms has continued to affect turn taking, modality and interpersonal stance, transforming the rules of formality and interaction. As well, the English after the pandemic is indicative of the wider sociocultural change including increased attentiveness to mental health, work-life separation, and group responsibility, which linguistically are marked by evaluative and stance-marking categories. There are also pedagogical and applied implications in the review that the English language teaching and the professional communication training programs need to consider such changing norms. In spite of increasing attention, studies in this field are still disjointed, with little longitudinal and cross-cultural studies. This review finds that the post-pandemic English is a new stage of linguistic adaptation under the influence of the world crisis experience, which can contribute to the useful conclusions about the interconnection between the linguistic change, social unsteadiness, and communicative stability.
The aim of the study is to gain a systematic understanding of the linguocultural aspects of the linguistic mechanisms that construct social norms, power relations, and models of sociality in the folk song texts of the village of Krasny Yar, Ufa District, Republic of Bashkortostan, revealing their connection to the discursive and social practices of the local community. The article examines the specifics of critical discourse analysis (CDA) in the tradition of N. Fairclough as a tool for linguistic analysis that links the linguistic features of a text with its sociocultural context. The subject of the study comprises lexical-semantic, syntactic, and discursive means involved in the representation of social hierarchies and identities. The novelty of the research lies in the first attempt to integrate the classical methodology of CDA and linguocultural tools with the corpus of folk songs from a specific local tradition, which allows moving from a philological and ethnographic description to an analysis of folklore as an active discursive practice that transmits social reality. As a result, the key discourses represented in the songs have been systematized: the discourse of patriarchal power, the discourse of social stratification, and the discourse of the community’s interaction with external institutions; the role of folk songs as a tool for the symbolic maintenance of order and the articulation of latent social tensions has been determined.
This article analyzes the regional vocabulary of the Provence region found in regional print media focusing on sports. The subject of the study is the regional lexical units characteristic of the Provence region. The object of the research is the functioning of regional vocabulary in the texts of print media on sports. The author examines in detail how the most significant sporting events in the region are described and which lexical units are used. The author's attention is directed solely to lexical units, as an analysis of language units at other levels (phonetic and grammatical) based on written texts is not possible for a number of reasons: a detailed study of the phonetic features of the regional language requires a corpus of audio and video texts. At the grammatical level, no differences between the literary and regional languages are identified, as the texts of print media are composed in accordance with literary norms, while the researcher's focus is not on colloquial forms but on the spoken language of educated speakers of this region's language. For comparison, texts from the most popular national and regional publications covering the same sporting events were selected and the lexical units used in their descriptions were analyzed. A method of complete sampling and contextual analysis was employed for this purpose. The novelty of this research is due to the fact that texts from print media are a very important source for analyzing the national and cultural features of the regional variant of the French language. Existing lexicographic sources at this stage do not provide reliable information about the functioning of linguistic units in everyday speech. Therefore, studying regional language features based on media material has become relevant. It is not by chance that sports themes were chosen for the analysis of lexical units, as they are the most akin to colloquial speech. The conducted analysis showed that regional print media texts on sports are characterized by a wide integration of regional lexical units. This leads to the conclusion that these lexical units are indeed used in the written language of educated speakers and serve as a cultural marker of the French language variant in the Provence region.
This article investigates the linguo-culturological parameters of herb names (phytonyms) in English, Russian, and Kazakh, focusing on their general and nationally specific characteristics. The study is grounded in linguocultural theory and examines plant names as linguistic units that reflect both botanical knowledge and culturally marked meanings. Phytonyms are analyzed as components of the lexical system that encode cognitive, semantic, and symbolic representations shaped by historical experience and national worldview. This research is based on a comparative analysis of dictionary definitions, phraseological units, proverbs, folklore texts, and works of fiction in three different languages. At the definitional level, English and Russian dictionaries tend to include not only botanical descriptions but also figurative and evaluative meanings. In contrast, Kazakh lexicographic sources primarily emphasize conceptual and functional characteristics. The study identifies common semantic features in phytonyms, such as classification, habitat, physical attributes, and practical uses (including medicinal, culinary, and decorative), while also revealing differences in metaphorization and symbolic associations. Phraseological units containing plant components demonstrate both shared conceptual meanings and nationally specific imagery. Although equivalent expressions exist across languages, their figurative bases and lexical composition often differ. Proverbs and sayings similarly reflect universal themes; family resemblance, moral education, and life difficulties—while preserving distinct cultural codes and value systems. Folklore and literary texts further illustrate how phytonyms function as metaphors, symbols of beauty, morality, abundance, or danger, and as markers of ethnic identity. The findings confirm that phytonyms constitute an important part of the linguistic worldview in each culture. Through comparative linguo-cultural analysis, the study demonstrates how plant names embody collective memory, mythological beliefs, aesthetic ideals, and social norms, thereby highlighting both universal patterns and culturally specific conceptualizations of nature in English, Russian, and Kazakh linguistic traditions.
Abstract Retranslation creates new versions of previously translated texts, documenting shifts in linguistic preferences, market needs, and ideological environments over time. Self-retranslation—when translators revise their own prior work—remains an uncommon and understudied phenomenon. This research examines five English novels that received second Chinese translations by their original translators 8–27 years after initial publication. Using an AI-assisted annotation system, we identified 89,175 changes across lexical, syntactic, semantic, pragmatic, and orthographic dimensions, and developed two measurement tools: the Fidelity Index and Audience-Accommodation Index. The data shows newer translations typically increase source-text fidelity, supporting the Retranslation Hypothesis, though with significant variations between works. We propose the Iterative Self-Retranslation Process (ISRP) model to explain these differences, connecting revision patterns to five factors: translator expertise development, changing linguistic norms and technologies, market influences, reader response, and sociopolitical environments. The study's methodology, along with the developed indices and model, offers a replicable framework for future research and equips researchers, translators, and publishers with practical tools for editorial planning.
Abstract. The article investigates the phenomenon of gender-sensitive language in modern English from linguistic and sociolinguistic perspectives, focusing on its historical development, theoretical foundations, contemporary transformations, and pedagogical implications. The study aims to determine how gender-inclusive forms function in present-day English, how they are perceived by younger speakers, and what role education plays in their dissemination and normalization. The research combines theoretical analysis and empirical investigation. The theoretical component includes a critical review of linguistic and feminist scholarship on language and gender, androcentrism, discourse theory, and language reform. Special attention is given to recent academic publications of the last five years addressing inclusive language practices. The empirical component is based on a structured questionnaire administered to university students. Quantitative and qualitative methods were used to analyse awareness, frequency of use, contextual variation, and attitudinal responses to gender-sensitive language forms. The findings demonstrate that gender-sensitive language in English is no longer marginal but increasingly integrated into academic and professional discourse. The singular they, the honorific Mx., and gender-neutral professional titles (e.g., firefighter, chairperson) are recognized by the majority of respondents. While systematic usage remains context-dependent, attitudes toward inclusive language are predominantly positive. Awareness correlates with academic background and exposure to digital media. The results confirm that language reflects broader sociocultural transformations toward equality and inclusivity. Gender-sensitive language represents not merely lexical innovation but a structural shift in linguistic norms influenced by social justice movements, institutional policies, and generational change. Its pedagogical integration is essential for sustainable implementation. Further research is needed to explore long-term normative stabilization, cross-cultural comparisons, and the cognitive impact of inclusive linguistic forms.
This chapter examines the theoretical foundations of dependency grammar, a framework that explains sentence structure through direct relations between words rather than constituent-based groupings. Tracing its origins to Tesnière&s;s Éléments de syntaxe structurale, the chapter outlines key concepts such as nucleus, actants, and circonstants, and emphasizes the model&s;s capacity to represent hierarchical relations through dependency trees. Different approaches within dependency grammar are presented, including Word Grammar, Link Grammar, Meaning-Text Theory, the Prague Dependency Treebank, the Universal Dependencies project, Lexical-Functional Grammar, and transition-based models. Each framework is discussed with regard to its theoretical premises, scope of analysis, and contributions to linguistic research and natural language processing. Particular attention is paid to the advantages of dependency-based approaches for languages with flexible word order, such as Turkish, and to their practical applications in fields like machine translation, automatic parsing, sentiment analysis, and information extraction. While acknowledging criticisms related to argument structure and the treatment of morphology, the chapter underscores the enduring relevance of dependency grammar as a bridge between theoretical linguistics and computational applications.
The article is devoted to the theoretical substantiation of the essence of the grammatical aspect of foreign language speech as a key component of foreign language communicative competence. The relevance of the study is determined by the need to understand the structural content of the grammatical aspect of speech in the context of the requirements of modern educational standards. The paper presents a comparative analysis of the approaches of foreign and Russian researchers to understanding the grammatical aspect of speech. D. Larsen-Freeman's three-dimensional model, which includes form, meaning, and use of grammatical phenomena, is examined and illustrated with examples. The position of S. Thornbury, who defines grammar through morphology and syntax and emphasizes its meaning-making potential realized in representational and interpersonal functions, is analyzed. Attention is paid to M. Lewis's lexical approach, which assigns a secondary role to grammar. The article presents the views of Russian methodologists (N.D. Galskova, N.I. Gez, E.N. Solovova), who consider grammar as a fundamental component of speech activity that ensures practical language proficiency for solving communicative tasks. Based on the analysis conducted, the author formulates a definition of the grammatical aspect of foreign language speech as a complex of automated actions for selecting, combining, and using grammatical structures in accordance with communicative intention and language norms. It is concluded that an insufficient level of mastery of the grammatical aspect of speech leads to difficulties in the formation of foreign language communicative competence as a whole.
The Canvas Model's equality processor operates in two complementary feed-modes. Feed-backwards (Steering) corrects errors via gradient descent: d\mathcal{E}/d\tau = -\kappa \nabla_{\mathcal{E}} \mathbb{E}. Feed-forward (Driving) anticipates goals via gradient ascent: d\mathcal{E}/d\tau = +\kappa \nabla_{\mathcal{E}} \mathbb{E}. The first four papers in this series explored Feed-backwards—training, convergence, regularization, and modularity. This paper explores Feed-forward. What this paper provides: · A formalization of anticipatory dynamics. The positive-sign dynamics enable the system to project forward, predict optimal trajectories, and act preemptively to maximize anticipated alignment with future goals. The update uses gradient ascent on a projected future state, not gradient descent on the current state.· Three experimental validations across domains: 1. Anticipatory continuous control (point mass positioning). Feed-forward Driving achieves integrated error of 0.72 \pm 0.11 vs 1.00 \pm 0.15 for reactive PD control and 0.78 \pm 0.12 for Model Predictive Control (MPC). Overshoot is reduced from 23% (PD) to 8% (Driving), beating MPC (12%). Driving matches MPC performance without a learned dynamics model or horizon optimization. 2. Lookahead for discrete sequence generation (character-level language modeling on Penn Treebank). Feed-forward Driving reduces test perplexity from 78.2 (autoregressive baseline) to 72.8 \pm 1.3. The lookahead projection anticipates grammatical constraints, avoiding locally probable but globally incoherent choices. Qualitative inspection confirms fewer repeated characters and more plausible word formations. 3. Adaptive guidance for image generation (classifier-guided diffusion on CIFAR-10). Standard classifier guidance uses a fixed scale w. Driving introduces adaptive per-step guidance: w_t = w_0 / (1 + p(y \mid x_t)). When the classifier is uncertain, guidance is stronger; when confident, guidance relaxes. Adaptive Driving achieves Fréchet Inception Distance (FID) of 4.05 \pm 0.14 vs 4.31 \pm 0.18 for standard guidance, and class accuracy of 95.1% vs 94.2% — better image quality and higher fidelity.· Connection to existing methods. Driving is not a new algorithm. It is a unifying principle that explains why Model Predictive Control, classifier-guided diffusion, lookahead optimizers, and RLHF work. All are manifestations of the Feed-forward mode, distinguished only by the choice of projection mechanism and the spectral energy being ascended.· The combined architecture. The full processor operates in both modes simultaneously: d\mathcal{E}/d\tau = -\kappa_b \nabla \mathbb{E}[\mathcal{E}_\tau] + \kappa_f \nabla \mathbb{E}[\mathcal{E}_\tau^{\text{proj}}]. The negative term corrects errors. The positive term anticipates goals. This is the mathematical framework for cognition itself: a system that both learns from its errors and acts on its anticipations.· Limitations acknowledged. Driving's effectiveness depends on the quality of the projection. If the projected future state is inaccurate, Driving can move toward the wrong target. The combined architecture (\kappa_b, \kappa_f > 0) mitigates this by allowing the system to both anticipate and correct. Why this matters: Steering corrects. Driving aims. The same processor, the same meta-time, the same spectral energy. Only the sign changes. The Canvas Model reveals that the fundamental dichotomy in AI—reactive vs. anticipatory, error-correcting vs. goal-seeking, introverted vs. extroverted—is not a philosophical distinction. It is a mathematical one: the sign of the update in meta-time. Keywords: Driving, Feed-forward dynamics, anticipatory control, lookahead generation, classifier guidance, diffusion models, gradient ascent, Model Predictive Control, RLHF, Canvas Model, equality processor, meta-time, Steering, Feed-backwards
Arabic WordNet 4.0 is a comprehensive lexical database for Modern Standard Arabic, derived from the Open English WordNet 2024 using the expand approach. Features:120,630 synsets (full OEWN 2024 parity)136,041 lexical entries184,238 senses297,150 synset relations (0 skipped — exact parity with OEWN 2024)97.3% ILI coverage for cross-linguistic linkingFull WN-LMF 1.4 XML format complianceAll synsets include Arabic definitions with full tashkeel (diacritical marks) on lemmas What's new in v4.1.0:+10,720 satellite adjectives (pos=s) — completing full OEWN 2024 adjective coverage+9 missing hub verbs (act/move, change, travel, make, communicate, and others)+78 upper-ontology noun synsets completing the noun hierarchyAll 8 validation checks pass against OEWN 2024 Methodology:Initial 109,823 synsets (nouns, verbs, adjectives, adverbs) were generated using AI-assisted translation (Google Gemini 3 Pro Preview). The remaining 10,807 synsets (satellite adjectives, hub verbs, upper-ontology nouns) were translated using Anthropic Claude via an automated Docker pipeline. Attribution:Derived from Open English WordNet 2024 (https://en-word.net/) and Princeton WordNet 3.0 (https://wordnet.princeton.edu/), both licensed under CC BY 4.0.
This study is devoted to the analysis of the complex and multifaceted relationship between dialects and the literary language in the history of the Uzbek language based on a historical-linguistic approach. The study considers dialects as an important source in the formation and development of the Uzbek literary language, and the influence of their phonetic, lexical and grammatical features on the norms of the literary language is consistently highlighted. The work provides a comparative analysis of written sources of the ancient Turkic period, samples of the old Uzbek literary language and modern dialect materials, paying special attention to the issues of historical continuity and coherence. Also, the role of regional dialects in the formation of literary language norms, the processes of their selection and assimilation into the national language are studied in connection with sociolinguistic factors. The results of the study show that the interaction between dialects and the literary language is not a one-sided, but a dynamic and complex process. This scientific work serves as a theoretical and practical basis for a deeper understanding of the laws of the historical development of the Uzbek language, as well as for improving literary language norms and the effective use of dialect materials.
The review is devoted to the analysis of the textbook by O. Mykytiuk and I. Farion «Language and Linguists: The Establishment of the Norm», which corresponds to the curriculum of the course «Ukrainian Language for Professional Purposes», currently studied by students of all specialties. It is demonstrated that, by virtue of its content, systematically implemented through the general didactic principle of scientific rigour, this publication meets the standards of a scholarly educational edition. Particular attention is paid to the main object of the scholarly and didactic exposition – the phenomenon of the language norm in its multifunctional representation. The textbook justifiably prioritises a multidirectional interpretation of orthographic norms through a diachronic-synchronic lens. Emphasis is placed on the conceptual dominants of fifteen thematic units and their specific informational content, which enables the tracing of key stages of Ukrainian glottogenesis – from ancient times to the present – as well as significant milestones in lexicographic and terminological studies. The authors also reveal the lexical, phraseological, and word-formation richness of the Ukrainian language, highlight the specificity of its phonetic and grammatical structure, and outline the development of its stylistic system. The originality of the work is further determined by its linguo-personalised component, represented by narratives about precedent linguistic (linguistic-scholarly) personalities, including P. Berynda, M. Smotrytskyi, O. Potebnia, P. Zhytetskyi, B. Hrinchenko, O. Syniavskyi, A. Krymskyi, O. Kurylo, I. Ohiienko (Metropolitan Bishop Ilarion), B. Antonenko-Davydovych, O. Tykhyi, O. Horbach, S. Karavanskyi, O. Ponomariv and V. Nimchuk. High praise is also due to the linguodidactic support materials, through which the authors – employing both traditional and innovative educational methods – promote the development of life and professional competences of future specialists, as well as the cultivation of an intellectually mature, spiritually rich, educated, linguistically cultured, patriotic, and Ukraine-centred personality.
This article investigates the use and expression of euphemisms in newspaper texts, focusing on their role as strategic linguistic tools in media discourse. Euphemisms, which replace direct or potentially offensive terms with milder or socially acceptable alternatives, are widely employed in political, economic, and social reporting to manage sensitive topics, maintain editorial neutrality, and influence reader perception. The study examines the linguistic and stylistic strategies newspapers use to construct euphemisms, including lexical substitution, metaphorical phrasing, nominalization, and circumlocution. By analyzing the distribution, frequency, and function of euphemistic expressions, the research highlights their importance in shaping tone, framing information, and reflecting cultural and ideological norms. The findings contribute to a better understanding of media language, discourse strategies, and the socio-pragmatic mechanisms underlying the presentation of delicate or controversial subjects in contemporary journalism.
Emotional memories persist within individuals and over generations. Past research has examined functions of parent-child memory sharing, but little work has assessed how the emotional qualities of memories are transmitted from parent to child and whether emotion transmission relates to memory content transmission. An understanding of how emotional memories are transmitted may be particularly important during adolescence, a developmental period marked by heightened sensitivity to emotional information, increasing independence from caregivers, and the emergence of mental health symptoms. The current study investigated the intergenerational transmission of parents’ autobiographical emotional memories to their teen offspring in healthy dyads using behavioral measures and natural language processing tools. We found that emotional valence and arousal linked to parents’ memories are transmitted from parent to teen. Greater parent-teen agreement in valence ratings was associated with greater parent-teen overlap in both subjective vividness ratings and objective memory content of individual memories. Subjective memory transmission was modestly related to lower mental health symptoms in teens after accounting for parent symptoms. These findings demonstrate that parents’ emotional memories are transmitted to their teens and provide preliminary evidence that autobiographical emotional memory transmission from parents could be a protective factor for mental health in adolescence.
Artificial intelligence (AI)-personalized learning is often promoted as a way to match instruction to learner differences, yet its educational value should not be judged only by efficiency or test performance. This narrative review uses algorithmic belonging as a conceptual lens to examine how AI-personalized learning may shape students’ inclusion, identity, and participation in diverse classrooms. Across the reviewed literature, personalized and AI-supported systems are typically associated with adaptation, feedback, recommendation, and learner support, while belonging research emphasizes acceptance, recognition, cultural affirmation, and meaningful participation. Direct studies that explicitly connect AI personalization with belonging remain limited; most available evidence must therefore be synthesized from adjacent work on AI in education, personalized learning, school belonging, culturally relevant pedagogy, and diverse-classroom participation. The literature suggests that AI personalization may support inclusion when it increases access, responsiveness, and student agency, but it may undermine belonging when it stereotypes learners, obscures decision rules, or reproduces cultural and linguistic norms that marginalize some groups. The strongest implication is that personalization should be treated as a social and ethical design problem, not merely a technical one. AI systems are most likely to support belonging when they are explainable, fair, culturally sustaining, and embedded in teacher–student relationships that affirm identity and widen participation. Keywords: artificial intelligence; personalized learning; belonging; inclusion; identity; participation; diverse classrooms; culturally relevant pedagogy; algorithmic bias; educational equity
In the context of globalization, digital communication, and the growing dominance of standardized language forms, dialectal lexical units are increasingly marginalized in everyday communication. However, dialects continue to function as vital carriers of national culture, historical memory, and social identity. This study aims to investigate the role of dialect-specific lexical units in reflecting national identity through a comparative analysis of Uzbek and English dialects. The research is grounded in linguocultural, sociolinguistic, and functional-semantic approaches. Empirical data were collected from regional Uzbek dialects and English dialectal sources, including spoken discourse, literary texts, and previous scholarly studies. More than fifty dialectal lexical units were selected based on their cultural specificity, emotional-evaluative potential, and limited equivalence in standard language. The data were analyzed using descriptive, comparative, and interpretative methods to reveal their semantic structure and communicative functions. The findings demonstrate that in Uzbek dialects, national identity is predominantly reflected through lexical units related to kinship (qudachilik), neighborhood relations (mahalla), and traditional labor practices. These units embody collective values such as social solidarity, mutual responsibility, and respect for tradition, and they possess strong emotional and evaluative connotations. In contrast, English dialects primarily reflect national identity through rural life vocabulary, occupational terminology, and class-based linguistic markers. Dialectal expressions in English often function as indicators of social class, occupational background, and individual identity, highlighting the role of social stratification in linguistic variation. The comparative analysis reveals that while both languages employ dialectal vocabulary as a means of cultural identification, Uzbek dialects emphasize collective and community-oriented identity, whereas English dialects foreground individual and class-based identity. The study concludes that dialect-specific lexical units should not be viewed as deviations from linguistic norms, but as essential components of national linguistic heritage. These findings have practical implications for linguocultural studies, language education, translation practice, and the preservation of intangible cultural heritage.
This article investigates the comparative linguistic features of artificial intelligence (AI)-based translation systems, focusing on how machine translation models render linguistic structures across typologically different languages, particularly Uzbek and English. The study analyzes syntactic, lexical, semantic, and pragmatic shifts occurring in AI-generated translations and compares them with human translation norms. Special attention is given to neural machine translation (NMT) systems and their ability to handle idiomatic expressions, polysemy, and word order variations. The research highlights both the strengths and limitations of AI translation tools, emphasizing issues such as loss of cultural nuance, structural simplification, and contextual misinterpretation. The findings suggest that although AI translation systems have significantly improved in fluency and accuracy, they still require linguistic refinement to fully capture deep structural and cultural meanings across languages.
This article analyzes the problems of compliance with literary language norms in students' speech from linguistic and pedagogical perspectives. The literary language norm is interpreted as an important and stable element of the language system, and its role in shaping speech culture is highlighted. During the research process, lexical, phonetic, orthoepic, grammatical, orthographic, and punctuation errors occurring in students' speech are analyzed, and the causes of their emergence are identified. In particular, the influence of dialects and vernaculars, the growing role of mass media and internet speech, as well as methodological shortcomings in the educational process are indicated as main factors. The article scientifically highlights the effectiveness of the communicative approach, text-based work, corrective and analytical exercises, and creative tasks in improving compliance with literary language norms. Additionally, the exemplarity of teacher speech and the importance of shaping language culture in the school environment are substantiated. The research results demonstrate the necessity of a conscious and systematic approach to literary language norms in developing students' speech literacy and communicative competence.
Compilation of the achieved results on the paper Simplifying Administrative Texts for Plain Language using LLM: a Comparative Analysis.On our paper, we evaluated if LLMs could perform text simplification following plain language guides, we compared the results on statistical readability indexes, such as Flesch Reading Ease and Gunning Fog Index, together with morphosyntactic metrics from the NILC-Metrix project (https://doi.org/10.1007/s10579-023-09693-w). The following models were evaluated:gemini-2.5-flash-preview-04-17gemini-2.5-pro-preview-05-06phi4:14.7bphi3:3.8bcow/gemma2_tools:2bllama3.2:3.2bgemma3:4bqwen2.5:14bdeepseek-r1:14bgranite3-dense:2bgranite3-dense:8b<br>The used code is available at:https://github.com/Joao-Pedro-P-Holanda/text-simplification/tree/SBSI-2026Data CleaningPre-Generation CleaningTo ensure that only the text would take focus, we performed the following processing steps before sending the text content to the LLMS:1. Extracted the Markdown text from the original PDF with the prompt ”Convert this document to markdown” on Google Gemini 2.5-pro;2. Stripped date and document numeration from page headers;3. Edited section headings to match exactly the PDF metadata;4. Edited paragraphs and line breaks to match the PDF structure;5. Added images that could not be translated to text effectively as a reference in Markdown format;6. Removed or added bullet points in order to match the original PDF structure, points using icons were replaced by simple points;7. Removed digital signatures, but preserving the author names and positions;Cleaning for PDF GenerationAfter saving the LLM responses in Markdown files, combining all chunks in a single document, we only removed the Deepseek tag from the texts and then proceeded to generate PDF files for the original and AI Generated versions.The exact command to generate a file was:<b>pandoc -o -Vgeometry:margin=1in --pdfengine=lualatex -V header-includes="\usepackage{fontspec}" -V mainfont=Times New Roman</b>Cleaning for Metric CalculationAfter generating the pdf files, we made 2 extra steps on the markdown files to ensure the adequate computation of readability indexes and morphosyntatic metrics:Removal of Markdown tables, extra whitespace and enumeration line starts (I., a), etc.)Addition of new lines between dot (.) separated sentencesArtifacts<br>The file prompt_simplify_document.txt contains the prompt given to all models during our text simplification process, the anonymized_readability_metrics_results.csv and anonymized_morphosyntactic_results.csv files have the raw results for all files in each model.The confidence_intervals csv files describe the measure confidence of the mean value across all documents for each model in each metric. NILC-Metrix metrics that had a proportional relationship to complexity were separated from the inversely proportional ones (in this work, only personal pronoun ratio). The same occurred for the statistical indexes, where grade level metrics were put separated from metrics in the 0-100 scale.foreign_word_ratio was not described in the NILC-Metrix paper, we propose its usage on this same paper, using the "Foreign" feat from Universal Dependencies.For the readability indexes, we split syllables on the words using the Pyphen (https://doc.courtbouillon.org/pyphen/stable/), and considered all words outside the 5000 most common Linguateca's words (https://www.linguateca.pt/acesso/tokens/formas.todos.txt) as complex. All indexes had their values adapted to Brazilian Portuguese, following the paper ALT: UM SOFTWARE PARA ANÁLISE DE LEGIBILIDADE DE TEXTOS EM LÍNGUA PORTUGUESA (https://revistas.ufrj.br/index.php/policromias/article/view/54352)We used UDPipe to perform Tokenization, PoS tagging, Lemmatization and Dependency Parsing with the treebank portuguese-porttinari-ud-2.15-241121.All morphosyntactic metrics used the CoNLL-U result files from UDPipe, we excluded the decontracted tokens and considered only the contracted word, e.g. "do" was decontracted to "de" preposition and "o" definite article, but only the word "do" was used in our implementation.
ABSTRACT As demand for dining out rises with income, reviews on platforms like Yelp.com have surged, making it harder for consumers to select suitable restaurants. This study uses the Elaboration Likelihood Model to analyze factors influencing Helpful and Funny Votes in reviews and explores the moderating effect of restaurant type (ethnic vs. non‐ethnic). Analyzing Yelp reviews from New York City, it examines central cues (cognitive effort, emotional experience, multi‐sensory experience) and peripheral cues (review length, images, rating). Results show that both types of cues significantly affect hedonic and utilitarian evaluations, with taste and tactile cues being more influential in hedonic evaluations for ethnic restaurants. The findings highlight the importance for managers to address negative reviews and promote positive ones to improve customer satisfaction. Identifying these influential factors can be instrumental in enhancing customer satisfaction with restaurant reviews.
Abstract This article examines a new direction in the digitalization of Kazakh linguistics–the development of a six-language parallel subcorpus based on literary texts. The research focuses on a multilingual database composed of texts in Kazakh, English, Turkish, Uzbek, Uyghur, and Azerbaijani. Its structural organization, metadata annotation system, and principles of paragraph-level alignment are described. The research material is based on a multilingual corpus compiled from Volume I of M. Auezov’s novel-epic The Path of Abai. The study provides a comparative analysis of syntactic, semantic, and pragmatic equivalence across translations in different languages. Particular attention is given to complex sentences, dialogic structures, forms of address and endearment, ethnomarked units, literary and poetic vocabulary, figurative expressions and idioms, as well as interjections. These elements are systematically classified, and their cross-linguistic equivalence levels are identified. The results demonstrate a high level of structural and semantic equivalence among Turkic languages, while English translations tend to prioritize functional and pragmatic adaptation. Furthermore, the six-language parallel subcorpus of the Kazakh language is shown to be a valuable resource for applied research in corpus linguistics, translation studies, comparative linguistics, and digital humanities. The scientific novelty of the study lies in the first-ever proposal of a six-language parallel subcorpus model in Kazakh linguistics based on literary texts, along with a comprehensive description of its metadata annotation, alignment, and analytical principles. The findings can be applied in corpus linguistics, translation studies, comparative linguistics, and the development of linguistic databases for artificial intelligence systems.
This study made use of distributional semantics to explore the semantic structure of two-character compounds in Mandarin Chinese. A total of 2,843 compounds, extracted from the Chinese National Corpus, informed a series of exploratory analyses. For larger-scale evaluation, we extracted all 29,376 two-character compounds from the Chinese Lexical Database (which is based on SUBTLEX-CH and Leiden Weibo Corpus) for which embeddings are available in the Tencent AI Lab resource. We observed that the compound families of mono\-morph\-emic monosyllabic Chinese words (pivots) cluster in the embedding space. The quality of these clusters does not depend on the ontological class of the compounds and is as good, and often better, than the quality of the clusters of the embeddings of derived words sharing the same suffix. The compound families of noun pivots are represented more prominently in the early principal components of the PCA-orthogonalized embedding space compared to the compound families of verb and adjective pivots. Compound families also cluster in the space of shift vectors, indicating that individual pivots contribute a core meaning to the compounds in which they occur. Approximately half of all pivots show no positional preference in the compound, and for 83.6\% of Mandarin two-character compounds, the position of the pivot cannot be predicted from the compounds' meanings. Therefore, pivot position cannot explain the clustering in semantic space of pivot families. Furthermore, pivots cannot be classified as prefixes, suffixes, interfixes, or circumfixes. Not being bound to a specific position for their interpretation, pivots are best characterized as ``floatfixes'', and instantiate morphological non-configurationality.
The article presents a corpus-based empirical analysis of colloquial units in contemporary English, treating colloquial vocabulary as a dynamic, multifunctional, and internally heterogeneous subsystem of the lexical system. Colloquial units are examined not merely as markers of informal speech, but as linguistically significant elements reflecting ongoing transformations in communicative practices, discourse conventions, and stylistic norms. Drawing on data from large-scale, register-diverse English language corpora, the study investigates the frequency, dispersion, contextual variability, and pragmatic functions of selected colloquial units across spoken and written registers. Special attention is paid to processes of stylistic diffusion, pragmatic refunctionalization, and partial desemanticization.
This paper examines the phenomenon of translingual writing and code-switching in contemporary South Asian Anglophone literature, analyzing how writers from India, Pakistan, Sri Lanka, and Bangladesh deploy multilingual textual strategies to represent the linguistic realities of postcolonial societies. Through close readings of works by Salman Rushdie, Arundhati Roy, Mohsin Hamid, and Shehan Karunatilaka, the study identifies distinct modes of translingual practice, including lexical borrowing, syntactic calquing, script-switching, and strategic untranslatability. The paper argues that these practices constitute a politics of language that challenges the monolingual norms of Anglophone literary culture and asserts the legitimacy of multilingual consciousness as both a literary subject and a mode of literary expression.
While word embeddings derive meaning from co-occurrence patterns, human language understanding is grounded in sensory and motor experience. We present $\text{SENSE}$ $(\textbf{S}\text{ensorimotor }$ $\textbf{E}\text{mbedding }$ $\textbf{N}\text{orm }$ $\textbf{S}\text{coring }$ $\textbf{E}\text{ngine})$, a learned projection model that predicts Lancaster sensorimotor norms from word lexical embeddings. We also conducted a behavioral study where 281 participants selected which among candidate nonce words evoked specific sensorimotor associations, finding statistically significant correlations between human selection rates and $\text{SENSE}$ ratings across 6 of the 11 modalities. Sublexical analysis of these nonce words selection rates revealed systematic phonosthemic patterns for the interoceptive norm, suggesting a path towards computationally proposing candidate phonosthemes from text data.
The Canvas Model's equality processor operates in two complementary feed-modes. Feed-backwards (Steering) is anticipatory and introverted: the processor compares the current state against stored experience, projects forward, and adjusts before error materializes. Feed-forward (Driving) is reactive and extroverted: the processor engages directly with the input, acting on what is present without consulting stored experience. The first four papers in this series explored the Feed-backwards mode—baseline subtraction as regularization, the S-invariant attractor for convergence, energy separation for modular architectures, and meta-time continuous training dynamics. All four operate in the anticipatory, introverted mode: learning from accumulated experience, correcting toward equilibrium. This paper explores the Feed-forward mode. What happens when the processor does not consult the past? It acts. It drives. What this paper provides: · A formal definition of the two feed-modes. Steering (Feed-backwards): d\mathcal{E}/d\tau = -\kappa \nabla_{\mathcal{E}} \mathbb{E} — anticipatory, introverted, consults stored experience. Driving (Feed-forward): d\mathcal{E}/d\tau = +\kappa \nabla_{\mathcal{E}} \mathbb{E} — reactive, extroverted, engages with immediate input. The same processor, the same meta-time \tau, the same spectral energy \mathbb{E}. Only the sign and the temporal reference differ.· A cognitive interpretation. Steering corresponds to introverted functions (Ni, Si, Ti, Fi): consulting internal memory, comparing against stored patterns, pausing before acting. Driving corresponds to extroverted functions (Ne, Se, Te, Fe): engaging with the external world, reacting to present stimuli, acting without hesitation.· Driving in neural networks. The reactive update does not compare against stored targets. It acts on the immediate input: W_{t+1} = W_t + \eta \nabla_W \mathcal{J}(W_t, x_t), where \mathcal{J} is an objective evaluated on the current input alone. This is necessary when there is no stored experience, when speed matters, when the environment is the teacher, and when generation is the goal.· Experiment 1: Reactive control (point mass). A Driving PD controller outperforms a Feed-backwards model-predictive controller on a simple positioning task (integrated error 0.72 vs 0.78, overshoot 8% vs 12%). When the environment is simple and predictable, consulting a model adds overhead without improving action.· Experiment 2: Sequence generation with lookahead. On character-level language modeling (Penn Treebank), Driving reduces perplexity from 78.2 (autoregressive) to 72.8 by reacting to incoherence as it emerges, rather than relying solely on patterns learned during training.· Experiment 3: Classifier-guided image generation. On CIFAR-10 diffusion models, Driving (adaptive guidance) achieves FID 4.05 and class accuracy 95.1%, outperforming fixed guidance (FID 4.31, accuracy 94.2%) by reacting to uncertainty as it arises during generation.· The combined architecture. The full processor operates in both modes simultaneously: d\mathcal{E}/d\tau = -\kappa_b \nabla_{\mathcal{E}} \mathbb{E}[\mathcal{E}_\tau] + \kappa_f \nabla_{\mathcal{E}} \mathbb{E}[\mathcal{E}_\tau^{\text{now}}]. The negative term anticipates (consults stored experience). The positive term reacts (engages with the present). The coupling constants \kappa_b and \kappa_f determine the balance. This is the architecture of cognition. Why this matters: The same processor, the same meta-time, only the direction changes. Steering anticipates. Driving acts. Both modes are necessary. Neither is sufficient alone. This is the mathematical framework for cognition itself: introversion consults the past; extroversion engages the present. The processor does both. Keywords: Driving, Feed-forward dynamics, Steering, Feed-backwards, equality processor, reactive control, sequence generation, diffusion guidance, introversion, extroversion, cognitive functions, Canvas Model, meta-time
Social media platforms, particularly TikTok, have become primary arenas for linguistic experimentation among adolescents, yet systematic analyses of how platform-specific affordances shape lexical and semantic innovation remains limited. This study investigated lexical and semantic variations in adolescent digital communication on TikTok, addressing three research questions concerning the types of lexical innovations, processes of semantic change, and the role of platform affordances in shaping language evolution. Methods: A mixed-methods design integrated quantitative corpus linguistics with qualitative discourse analysis. A corpus of 2,848 TikTok comments was compiled across four major trends (September–December 2024). Lexical analysis identified neologisms, graphical variations, and acronyms; semantic analysis documented broadening, narrowing, metaphoric extension, and pejoration/amelioration; platform affordances analysis examined meme-driven language and intertextual policing. Analysis revealed 15 lexical innovations with 63 occurrences across semantic categories. Neologisms (fr, bestie, delulu) and graphical variations (tryna, cuz, ion) served dual functions of efficiency and identity performance. Semantic shifts included ameliorative broadening (slay, fire), pejoration (basic, cringe), metaphoric extension (era, main character), and reclamatory usage (ghetto). Platform analysis identified 11 meme-driven phrases generating 2,848 occurrences with near-neutral sentiment, and 347 policing instances (12.2%) concentrated during rising and peak trend phases, demonstrating active semantic negotiation through definition, debate, and correction. TikTok functions as an accelerated laboratory for language change where adolescents deploy multiple mechanisms of linguistic innovation simultaneously. Platform affordances fundamentally reshape traditional sociolinguistic processes, with intertextual policing serving as the mechanism by which communities enforce emerging semantic norms. The findings extend communities of practice frameworks to algorithmically-mediated digital environments. Educators should recognize digital language as systematic innovation; lexicographers should develop protocols for documenting ephemeral platform-specific terms; platform designers should account for in-group reclamation practices; and researchers should prioritize cross-platform longitudinal studies to track whether observed innovations represent enduring change or age-graded phenomena.
This article provides a thorough examination of the critical role that intercultural pragmatic competence plays in contemporary English language instruction. This sophisticated construct extends beyond traditional linguistic knowledge to encompass the nuanced understanding of how language functions within diverse cultural frameworks to convey meaning, intent, and social relationships. Contemporary English Language Teaching (ELT) methodologies have undergone a significant paradigmatic transformation, characterized by growing acknowledgment of the complex interdependence between linguistic structures, communicative intentions, and the sociocultural contexts that shape their interpretation. This comprehensive perspective deliberately moves beyond conventional pedagogical approaches that prioritized grammatical accuracy and lexical acquisition in relative isolation. Rather, it actively promotes a more profound comprehension of target cultures, recognizing that successful communication depends substantially on understanding culturally conditioned expectations regarding appropriateness, politeness, and discourse organization. Central to this evolving pedagogical framework is the systematic integration of communicative language teaching principles. This approach provides substantial theoretical foundations for investigating how cultural norms, social conventions, and contextual factors fundamentally influence language learners’ interpretation and production of meaning in authentic communicative situations. Ultimately, the findings presented herein compellingly demonstrate the imperative of equipping language learners not merely with structural accuracy and lexical diversity, but fundamentally with the pragmatic awareness essential for genuinely effective, contextually appropriate, and mutually comprehensible cross-cultural communication, thereby enabling them to navigate the complexities of international discourse with competence and cultural sensitivity.
The aim of the article is to uncover the cultural specificity of the meaning of lexical units “kola” and “yam” in Nigerian fiction. Methodology. To solve the tasks at hand, a number of both general scientific and specific methods of investigation were used. Systematic and cluster sampling methods were employed in selecting the linguistic material. The need to describe and analyse the semantic structure of the lexical units, as well as the material of the research, made it possible to resort to the methods of contextual and component analysis. The research deals with modern texts by Igbo authors (1990s – 2010s), as well as the classical works of Chinua Achebe. Results. This article identifies the culturally specific semantic properties of the lexemes “kola” and “yam”. The use of these lexical units is thoroughly analysed, especially in regards to units that can be considered as lexical neologisms as compared to the referent norm. These neologisms are a part of the lexical field “traditions and customs”. It is concluded that the culturally significant lexeme “yam” has gender markings and symbolizes the masculine principle, which is reflected both in the early and modern stages of the development of English-language Igbo literature. Research implications. The article is of theoretical and practical value to philologists and specialists that work with various variants of West African English. It provides recommendations as to the translation of phrases containing the kola unit.