Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Large language models (LLMs) exhibit cultural bias from overrepresented viewpoints in training data, yet cultural alignment remains a challenge due to limited cultural knowledge and a lack of exploration into effective learning approaches. We introduce a cost-efficient and cognitively grounded method: fine-tuning LLMs on native speakers' word-association norms, leveraging cognitive psychology findings that such associations capture cultural knowledge. Using word association datasets from native speakers in the US (English) and China (Mandarin), we train Llama-3.1-8B and Qwen-2.5-7B via supervised fine-tuning and preference optimization. We evaluate models' cultural alignment through a two-tier evaluation framework that spans lexical associations and cultural value alignment using the World Values Survey. Results show significant improvements in lexical alignment (16-20% English, 43-165% Mandarin on Precision@5) and high-level cultural value shifts. On a subset of 50 questions where US and Chinese respondents diverge most, fine-tuned Qwen nearly doubles its response alignment with Chinese values (13 to 25). Remarkably, our trained 7-8B models match or exceed vanilla 70B baselines, demonstrating that a few million of culture-grounded associations achieve value alignment without expensive retraining. Our work highlights both the promise and the need for future research grounded in human cognition in improving cultural alignment in AI models.
This article looks at the semiotics of humorous self-censorship. To this end, selected examples from a corpus of YouTube commentary videos and their respective comment sections are presented and discussed. On the one hand, the analysis focuses on structural aspects of humorous self-censorship signs from various modes (written, spoken, images, emojis, etc.) in their multimodal interplay. On the other hand, we analyze socio-semiotic aspects of our examples: their anchoring in specific speech communities, marked by background knowledge and shared communicative practices. The analysis shows that (1) structural manipulations of the spelling and/or phonetic shape of lexical items and of images etc. serve to secure both the understanding of the censored item as well as plausible deniability, while generating potentially humorous incongruities, (2) various positions in the participation process are exploited to participate in this process, including trigger warnings and other metacommunicative actions and levels, (3) humorous self-censorship serves a number of social-semiotic functions, such as the negotiation of group norms of sayability, the expression of group solidarity, and – importantly – entertainment.
This article is to research of the onomasticon of the Kazakh heroic epic “Kobylandy Batyr”, translated into the Sakha (Yakut) language by A. N. Zhirkov, T. S. Kirillin, G. G. Torotoev and published in 2025 as part of the international book series “Epic Monuments of the Peoples of the World”. The author of the article analyzes and arranges the core of onomastic units of this epic from the point of view of translation transformation and etymologization. The main research methods are the comparative method, the method of lexico-semantic analysis, the descriptive method, and statistical analysis. Comparatively, 71 onomastic units were analyzed, as a result, 6 groups of proper names were identified: 38 anthroponyms, 11 toponyms (6 oronyms, 5 astyonyms), 11 hydronyms (8 limnonyms, 3 potamonyms), 7 ethnonyms, 2 hagionyms, 2 zoonyms. Per semasiology, all onomastic units introduced into epic circulation represent a national and cultural value, and in one way or another contribute to the interpretation of the cultural code of the people. Kirillin, Zhirkov and Torotoev all use different translation methods in their translations, such as equivalent substitution, transcription, descriptive translation, addition, omission, out-of-text commentary, etc. The author of this article raises the problem of adequate translation of proper names from Kazakh to Yakut. The translation of onomastic units is a complex process, since they belong to the number of realities. Relative to the onomastic units used as comparison material, transcription is the dominant translation method (Kirillin – 63.5%, Zhirkov – 80.3%, Torotoev – 71.8%); the second position is occupied by equivalent substitution (Kirillin – 5.6%, Zhirkov – 12.6%, Torotoev – 28.2%). This proves the initiative of translators to follow the law of the Sakha language's synharmonicity and actualize the lexical potential of the Sakha language, while preserving the centuries-old customary norms of alliterative versification. This article is relevant because it provides an optimal solution to the problem of difficult translation situations, where epic translators face difficulties in transferring onomastics from the original language to the translation language.
How does the bilingual experience affect online processing? The distribution of lexical items shared between monolinguals and bilinguals can differ greatly. One critical difference is how code-switching allows more variability in the relative co-occurrence of words. The current study uses a visual world paradigm to test whether the relative distribution between Spanish gender-marked determiners ("el," "la") and the non-marked English determiner ("the") predict the Spanish-English bilingual's ability to predict and/or integrate an incoming noun. While we replicate a previously observed asymmetry among Spanish-English bilinguals between the masculine "el" and feminine "la," our cluster permutation test results reveal differences in how bilinguals predict and integrate nouns when preceded by "el" versus "la" or "the." Comparing our results to existing corpus data, we argue that bilinguals rely on the distributional norms they experience across both single-language and code-switched contexts to facilitate online processing.
This article investigates the pragmatic functions of first-person pronouns (I/we) in political speeches and formal writing across multiple language systems. By analyzing the discourse of prominent political leaders such as Shavkat Mirziyoyev, Joe Biden, and Emmanuel Macron, the study explores how first-person pronouns function to construct authority, inclusiveness, and responsibility. The research highlights variation across languages in terms of politeness, formality, and rhetorical strategy. Methodologically, the study employs comparative discourse analysis and pragmatic interpretation of political and academic texts. The findings demonstrate that the usage of "I" versus "we" reflects not only linguistic norms but also culturally embedded leadership styles.
Decision-making in economic and moral contexts involves complex affective processes that shape judgments of fairness, responsibility, and conflict resolution. While previous studies have primarily examined behavioral choices in economic games and moral dilemmas, less is known about the underlying affective structure of these decisions. This study investigated how individuals emotionally represent economic (ultimatum game) and moral (trolley dilemma) decision-making scenarios using multidimensional scaling (MDS) and classification. Participants rated their emotional responses, including positive (pleased, calm, happy, peaceful) and negative (irritated, angry, gloomy, sad, fearful, anxious) affective states, to 16 scenarios varying by game type, the presence or absence of conflict, and intensity. MDS revealed two primary affective dimensions of distinguishing conflict from no-conflict and economic from moral scenarios. No-conflict-economic scenarios were strongly associated with positive affective responses, while the no-conflict-moral scenarios elicited heightened fear and anxiety rather than positive emotions. Increasing unfairness in the ultimatum game affected affective representation, while variations in the number of lives at stake in the trolley dilemma did not. Cross-participant classification analyses demonstrated that game type and conflict conditions could be reliably predicted from affective ratings, indicating systematic and shared emotional representations across participants. These findings suggest that economic and moral decisions evoke distinct affective structures, with fairness modulating conflict perception in economic contexts, while moral decisions remain affectively stable despite changes in intensity.
Pre-trained Language Models (PLM) have enabled a cost-effective approach to handling various downstream applications via Parameter-Efficient-Fine-Tuning (PEFT) techniques. In this context, service providers have introduced a popular fine-tuning-based product service known as Model-as-a-Service (MaaS). This service offers users access to extensive PLMs and training resources. With MaaS, users can fine-tune, deploy, and utilize their customized models seamlessly, leveraging a one-stop platform that allows them to work with their private datasets efficiently. However, this service paradigm has recently been exposed to the possibility of leaking user private data. To this end, we identify the data privacy leakage risks in MaaS-based PEFT and propose a Split-and-Privatize (SAP) framework, mitigating the privacy leakage by integrating split learning and differential privacy into MaaS PEFT. Furthermore, we propose Contributing-Token-Identification (CTI), a novel method to balance model utility degradation and privacy leakage. As a result, the proposed framework is comprehensively evaluated, demonstrating a 65% improvement in empirical privacy with only a 1% degradation in model performance on the Stanford Sentiment Treebank dataset, outperforming existing state-of-the-art baselines.
We introduce the first dependency treebank containing Universal Dependencies (UD) annotations for Spanish learner writing from the UC Davis COWSL2H corpus.Our annotations include lemmatization, POS tagging, and syntactic dependencies.We adapt the existing UD framework for Spanish L1 to account for learner-specific features such as code-switching and non-canonical syntax.A suite of parsing evaluation experiments shows that parsers trained on learner data together with moderate sizes of Spanish L1 data can yield reasonable performance.Our annotations are openly accessible to motivate future development of learner-oriented language technologies.
This article explores the critical engagement of two academics who confront lived experiences with the institutional and tangible dimensions of linguistic barriers and discrimination in the Portuguese district of Faro. Centred on the challenges posed at the border, the study fits into the wider framework of mobilities between North Africa and southern Europe, and attempts to demonstrate how language, as a form of social practice, impacts access to employment, education and society at large. Portuguese emerges as a dual entity: an institutional barrier, a form of socio-spatial control that reinforces exclusion, illustrating exclusionary processes within hierarchies and structural violence; a transnational bridge that fosters belonging and, simultaneously, a battleground where identity and self-determination face the constraints imposed by the economic, social, and political order that impact Moroccan migration to southern Portugal. The research highlights the imbrications of the undervaluation of migrants' cultural knowledge, the ambivalence of linguistic identity within a globalized world, exacerbating social exclusion, and systemic discrimination of non-privileged migrants in a polarized region shaped by social and economic asymmetries and increasingly representative nationalisms. From a systemic justice perspective, this devaluation reinforces structural inequalities, marginalizing those who do not conform to dominant linguistic norms. The national languages’ role as linguistic and cultural gatekeeper exemplifies the intersection of identity construction and socio-political hierarchies in the context of mobilities. This study, grounded in a collaborative project blending autobiography and biographical research, employs qualitative methods, including biographic interviews, ethnographic observation, and critical human rights studies.
The obligatory use of third-person honorifics is a distinctive feature of several South Asian languages, encoding nuanced socio-pragmatic cues such as power, age, gender, fame, and social distance. In this work, (i) We present the first large-scale study of third-person honorific pronoun and verb usage across 10,000 Hindi and Bengali Wikipedia articles with annotations linked to key socio-demographic attributes of the subjects, including gender, age group, fame, and cultural origin. (ii) Our analysis uncovers systematic intra-language regularities but notable cross-linguistic differences: honorifics are more prevalent in Bengali than in Hindi, while non-honorifics dominate while referring to infamous, juvenile, and culturally exotic entities. Notably, in both languages, and more prominently in Hindi, men are more frequently addressed with honorifics than women. (iii) To examine whether large language models (LLMs) internalize similar socio-pragmatic norms, we probe six LLMs using controlled generation and translation tasks over 1,000 culturally balanced entities. We find that LLMs diverge from Wikipedia usage, exhibiting alternative preferences in honorific selection across tasks, languages, and socio-demographic attributes. These discrepancies highlight gaps in the socio-cultural alignment of LLMs and open new directions for studying how LLMs acquire, adapt, or distort social-linguistic norms. Our code and data are publicly available at https://github.com/souro/honorific-wiki-llm
V prispevku je predstavljen postopek izdelave korpusa CVET, ki vsebuje besedila patra Hijacinta Repiča, objavljena v verski reviji Cvetje z vertov sv. Frančiška v obdobju 1881–1916. Korpus je uporabljen kot podlaga za jezikovno in stilistično analizo, opravljeno z orodjem noSketch Engine. Z analizo frekvenčnosti izbranih spremenljivk sta opisana besedišče patra Repiča in njegov pripovedni slog. Nazadnje je na primeru besed tipa bralec/bravec opazovan sinhrono-diahroni in normativni vidik starejšega slovenskega jezika besedil v korpusu.
This study characterize mood and emotional regulation in women with premenstrual syndrome (PMS) using near-infrared spectroscopy (NIRS) and mood assessments. Hemodynamic responses in the prefrontal cortex (PFC) were measured while participants viewed emotion-inducing images during both the follicular and luteal phases of the menstrual cycle. Emotional valence and arousal ratings for each image were obtained immediately after the task. In addition, mood states were evaluated prior to the task using the Profile of Mood States Second Edition (POMS2). Forty-six women completed the procedures across both menstrual cycle phases. After excluding five participants diagnosed with premenstrual dysphoric disorder (PMDD), data from 41 participants (non-PMS: n = 25, PMS: n = 16) were analyzed. During the follicular and luteal phases, the PMS group showed significantly lower integrated oxyhemoglobin (oxy-Hb) responses to positive emotional stimuli in the right prefrontal region compared to the non-PMS group (follicular phase: p =.03, r = 0.35; luteal phase: p <.001, r =.0.52). These r values represent effect sizes (rank-biserial correlations) corresponding to the Mann–Whitney U tests. No significant group differences were found in the left prefrontal region and response to negative stimuli. Based on POMS2 scores, the PMS group showed significantly higher scores on six negative mood scales and a lower score on one positive mood scale during the follicular phase. During the luteal phase, only one negative mood subscale score was significantly higher in the PMS group. Group differences in subjective emotional valence and arousal ratings were evident only during the follicular phase. These findings suggest that women with PMS exhibit attenuated neural responses to positive emotional stimuli and altered mood states even during the follicular phase when symptoms are often not consciously recognized. This may highlight the possibility that early-phase emotional dysregulation may precede overt symptom manifestation in PMS, providing insights into its temporal dynamics.
This study characterize mood and emotional regulation in women with premenstrual syndrome (PMS) using near-infrared spectroscopy (NIRS) and mood assessments. Hemodynamic responses in the prefrontal cortex (PFC) were measured while participants viewed emotion-inducing images during both the follicular and luteal phases of the menstrual cycle. Emotional valence and arousal ratings for each image were obtained immediately after the task. In addition, mood states were evaluated prior to the task using the Profile of Mood States Second Edition (POMS2). Forty-six women completed the procedures across both menstrual cycle phases. After excluding five participants diagnosed with premenstrual dysphoric disorder (PMDD), data from 41 participants (non-PMS: n = 25, PMS: n = 16) were analyzed. During the follicular and luteal phases, the PMS group showed significantly lower integrated oxyhemoglobin (oxy-Hb) responses to positive emotional stimuli in the right prefrontal region compared to the non-PMS group (follicular phase: <i>p</i> =.03, r = 0.35; luteal phase: <i>p</i> <.001, r =.0.52). These r values represent effect sizes (rank-biserial correlations) corresponding to the Mann–Whitney U tests. No significant group differences were found in the left prefrontal region and response to negative stimuli. Based on POMS2 scores, the PMS group showed significantly higher scores on six negative mood scales and a lower score on one positive mood scale during the follicular phase. During the luteal phase, only one negative mood subscale score was significantly higher in the PMS group. Group differences in subjective emotional valence and arousal ratings were evident only during the follicular phase. These findings suggest that women with PMS exhibit attenuated neural responses to positive emotional stimuli and altered mood states even during the follicular phase when symptoms are often not consciously recognized. This may highlight the possibility that early-phase emotional dysregulation may precede overt symptom manifestation in PMS, providing insights into its temporal dynamics.
The article is devoted to studying socially determined linguistic processes, which are traditionally associated with the broad problem of social variation of communication (discourse) and linguistic variability of the German language. It presents the results of a study conducted within the framework of cognitive sociolinguistics - the linguistics of social meanings. The author observes continuity in the development of scientific thought, explores the problem of linguistic variability in two modes - theoretical and applied. The subject of the research is the terminological apparatus and specific linguistic facts that are based on the cognitive, communicative and social functions of language, that is, to be the reality of the thought of individual and/or collective knowledge of representatives of a certain society and the means of their communication. The purpose of the article is to analyze theoretical propositions, linguistic terms and existing specific language forms that convey social meanings, marked by a social feature in their terminological interpretations with a focus on the linguistic picture of the Germany new lands. The novelty of the research lies in cognitive-semantic analysis, systematization and modeling this phenomenon in the social aspect. As a result, the author comes to her own conclusions, which lead to understanding and rethinking the old views on the problems of dialectology in the context of modern realities of the German language society and the data of modern linguistics. The methodology of the research and the description of its results are determined by the principle of interdependence of the three most important didactic and linguistic strata - the study of modern language from the standpoint of linguistic norms, the study of linguistic variability and the analysis of language change in the framework of its historical development. General scientific methods and special methods of cognitive linguistics are used to analyze theoretical material and linguistic facts, including explanatory description, interpretation, cognitive modeling, cognitive dominance, and focusing.
The article explores irony as a complex linguistic and pragmatic phenomenon that plays a significant role in creating satirical effect within the sketch genre. The object of analysis is the sketch «Les Flics» by French comedian Coluche, in which irony functions as a tool of social critique, particularly through the ridicule of flaws within the institution of police. The study identifies the main linguistic, stylistic, and pragmatic means used to construct irony. It demonstrates that the central communicative strategy shaping the critical attitude toward the police system is the deliberate provocation of the audience and the emphasis on the absurdity of the depicted social conditions. The comic effect arises from the contrast between socially expected norms of behavior (including verbal conduct) and the reality represented by the narrator-character. Stylistic devices such as antiphrasis, metaphor, metonymy, grotesque, and sarcasm, along with linguistic features such as colloquial register, slang, and violations of linguistic norms, are shown to contribute to the portrayal of deep linguistic and social deviation in the police character. The article draws on contemporary linguistic theories, including pragmatics, speech act theory, polyphony theory, and relevance theory. It concludes that irony in the sketch is both staged and situational, with the comic effect emerging from the conflict between expectation and reality. Special attention is given to implicature, polyphonic structure of the utterance, and the satirical function of the narrator. The study shows that irony not only enhances the aesthetic dimension of the text but also performs a socially critical function by shaping a discourse of resistance. Future research directions include the analysis of other Coluche sketches in the context of linguistic critique of society.
Passive voice remains a key grammatical structure for English learners, particularly in academic writing, yet many students struggle to use it accurately. This study analyzes the types of passive voice errors made by 19 fifth-semester students in the English Education Study Program at Tadulako University. Specifically, it addresses two questions: (1) How do classroom interaction patterns such as teacher-centered grammar instruction, limited student negotiation of meaning, or feedback practices shape students’ understanding and use of passive voice, and to what extent might these dynamics contribute to the dominance of developmental errors? (2) In what ways do students’ sociocultural backgrounds, prior educational experiences, and exposure to English outside the classroom influence their difficulties with auxiliary verbs and tense agreement, and how do these factors mediate tensions between Indonesian linguistic norms and English academic writing conventions? A quantitative design was employed, with a test focusing on passive constructions in present continuous, past continuous, and past perfect tenses. Students’ responses were categorized using Dulay et al.'s (1982) comparative taxonomy of developmental and interlingual errors. Results revealed developmental errors as the most prevalent (89.9%), mainly involving incorrect auxiliary verbs (is, am, are, being, been), past participle formation, and tense agreement. These findings highlight the need for targeted grammar instruction on auxiliary patterns and participles, alongside enhanced practice, corrective feedback, and adjustments to classroom interactions and sociocultural considerations to boost accuracy.
Mood, an individual’s emotional state, fundamentally shapes how the brain interprets sensory input by providing a continuous affective context for prediction and evaluation. In language processing, mood may bias the interpretation of emotionally valenced words, amplifying or dampening their perceived affect. Yet, the temporal dynamics of these mood-valence interactions remain poorly understood. To clarify inconsistent evidence on the timing and nature of mood-valence interactions, we examined how induced mood influences early stages of emotional word processing using EEG. Participants performed a valence-rating task for positive, negative, and neutral words in a baseline condition and following positive or negative mood induction. Event-related potentials were analysed across early processing windows (N1, P2, EPN) using cluster-based permutation statistics. Positive mood selectively attenuated N1 amplitudes for highly valenced words, consistent with reduced prediction error under mood-congruent expectations. Later components (P2, EPN) showed decreased amplitudes for both high and neutral valence, suggesting reduced model updating under mood-congruent expectations. Negative mood, in contrast, produced weaker and temporally delayed modulations. Behaviourally, participants responded more quickly to valenced words under induced mood conditions, supporting the neural findings. Interpreted within a predictive coding framework, these results support the theoretical view that mood functions as a hyperprior, tuning the precision of predictive models during language comprehension. Positive mood appears to enhance predictive flexibility and facilitate the processing of affectively congruent words, whereas induced negative mood reduces positive affect. Taken together, the findings highlight how affective states dynamically modulate early predictive mechanisms in emotional language processing.
Background People with aphasia (PWA) often struggle to determine who/what pronouns refer to. However, spontaneous speech studies on non-fluent PWA have revealed a disparity: one cluster of PWA tends to omit pronouns, while another cluster tends to overuse them. This study improves our current knowledge of pronoun processing by presenting how aphasia impacts pronoun production in Turkish.Aims This study addresses three questions: (i) whether Turkish-speaking PWA show impairments in producing pronouns in spontaneous speech, either through omission or overproduction; (ii) whether their production of personal pronouns is selectively affected; and (iii) whether theyexhibit difficulties in producing deictic and/or non-personal pronominal elements, including demonstratives, indefinites, and possessives.Methods & Procedures Spontaneous speech samples from 10 PWA speaking Turkish and 10 matched healthy controls were analysed. Three groups of pronoun variables were quantified: (i) general characteristics of pronoun uses including total number of pronominal elements, and pronoun-to-noun ratios, (ii) total and null personal pronouns in subject and object positions, and (iii) other types of pronouns including demonstrative, indefinite, and possessive pronouns.Outcomes & Results The Turkish-speaking PWA were above the control norms in total number of pronominal elements, and pronoun-to-noun and pronoun-to-word ratios, but they showed reduced lexical diversity in noun usage. They produced more first-person (deictic) than third person (anaphoric) pronouns while this difference was not significant in the controls. The PWA exhibit overuses in null subjects and objects, demonstratives and indefinite pronouns.Conclusions Turkish-speaking PWA overproduce pronouns; however, this overuse does not represent a uniform pattern. Overuses caused mostly by deictic elements such as demonstratives and indefinite pronouns, in the context of reduced lexical diversity in nouns. While the overuse of pronouns likely reflects a communicative strategy, the extent and nature of this strategy seem to vary across individuals.
Large language models (LLMs) have reignited debate about whether machines without minds or intentions can genuinely participate in linguistic practice. Critics portray them as ‘stochastic parrots’ that manipulate form without meaning, whereas defenders emphasize their impressive functional capacities. This paper argues that these disputes conflate distinct dimensions of meaning and agency. I extend Huw Price’s distinction between i-representation and e-representation (roughly, inferential versus environment-tracking types of representation) by differentiating physical e-representation—such as a fuel gauge, grounded in causal coupling—from symbolic e-representation, exemplified in language and mediated by agents. This refinement clarifies what is at issue: LLMs clearly display i-representational competence through their participation in inferentially structured discourse. Whether their outputs possess symbolic e-representational content, however, is contested and framework-relative. It depends on whether agent-mediated uptake is taken to suffice, or whether additional grounding conditions—such as intentions, causal connections, or proper functions—are required. I further distinguish norm-sensitivity—the capacity to track and adapt to linguistic norms, which grounds their i-representational competence—from norm-responsibility, the reflexive capacity to own commitments and bear accountability. Technical analysis of LLM architectures shows that they exhibit advanced norm-sensitivity through statistical learning but entirely lack norm-responsibility. LLMs thus occupy a distinctive position: they are genuine functional participants in linguistic practices, yet fall short of the reflexive agency characteristic of responsible speakers.
<span lang="EN-US">The illicit act of appropriating programming code has long been an appealing notion due to the immediate time and effort savings it affords perpetrators. However, it is universally acknowledged that concerted efforts are imperative to identify and rectify such transgressions. This is particularly crucial as academic institutions, including universities, may inadvertently confer degrees for work tainted by this form of plagiarism. Consequently, the primary objective of this research is to scrutinize the feasibility of identifying plagiarism within pairs of Verilog algorithms and texts. this study aims to detect plagiarism in textual content and Verilog code by leveraging diverse linguistic characteristics from the WordNet lexical database. The primary objective is to achieve optimal accuracy in identifying instances of plagiarism, incorporating features such as modifications to text structure, synonym substitution, and simultaneous application of these strategies. The system's architecture is intricately designed to unveil instances of plagiarism in both textual content and Verilog code by extracting nuanced characteristics. The systematic process includes preprocessing, detailed analysis, and post-processing, supported by a feature-rich database. Each entry in the database represents a distinctive similarity case, contributing to a thorough and comprehensive approach to plagiarism detection.</span>
This study reviews the English language test of Singapore’s Primary School Leaving Examination, a high-stakes national assessment taken annually by nearly all primary six students for secondary school placement. Given the test’s importance in shaping students’ academic pathways and recent format changes, it is crucial to evaluate its validity, specifically its ability to provide accurate and fair assessments of students’ English language proficiency and academic readiness. The review outlines the test’s educational and policy context, followed by a description of the latest formats for both the English language and foundation English language versions. The analysis focuses on core dimensions of test validity, including content representativeness, construct validity, criterion-related validity (concurrent and predictive), and reliability (inter-rater reliability and internal consistency). Drawing on official documents and limited empirical studies, the review finds moderate improvements in content representativeness and construct validity. However, both longstanding and emerging concerns (e.g., the exclusion of local linguistic norms and genre scope) indicate that key limitations remain. While predictive validity, inter-rater reliability, and internal consistency appear supported, empirical research remains sparse across all reviewed test qualities, particularly in concurrent validity. The review integrates identified research gaps and proposes inquiry directions to inform future test development and policy adaptation. Strengthening the evidence base is essential for ensuring a valid, reliable, and equitable assessment system in Singapore’s primary education landscape.
Word Sense Disambiguation (WSD) is a fundamental task in Natural Language Processing (NLP), addressing the challenge of identifying correct word meanings in context. This task is particularly complex for morphologically rich and resource-limited languages like Hindi, which exhibit significant lexical ambiguity compounded by limited availability of annotated corpora. To address these challenges, we propose a supervised approach combining the multilingual BERT model (mBERT) with Hindi WordNet as a structured lexical resource. Using few-shot learning, we fine-tune mBERT on a dataset constructed from Hindi WordNet to disambiguate contextually ambiguous words across four parts of speech (POS): nouns, verbs, adjectives, and adverbs. Experiments on standard Hindi WSD benchmarks demonstrate that our method significantly outperforms traditional rule-based and embedding-based approaches, achieving 96.48% accuracy—an approximate 3% improvement over the strongest baseline. These results validate the effectiveness of integrating contextualized embeddings from pre-trained language models with structured lexical databases, highlighting the promise of hybrid techniques for advancing WSD in low-resource languages and providing a framework applicable to other morphologically complex languages with similar resource constraints.
The abbreviations by superscript letters in 18th century Portuguese are analyzed, with the aim of finding consistent uses by the copyist of abbreviation strategies by superscription. The hypothesis tested is that there are patterns when abbreviating a lexical item, and that the choices made by the copyist regarding the letters that are selected to be overwritten, as well as their quantity, are not random. The corpus used is the Manuscript PBA-749 from the National Library of Portugal, which title is Primeiro copiador das respostas dos senhores governadores desta capitania [minas gerais] às ordens de s[u]a mag[esta]de, e contas que lhe dãoão que principia no governo do sen[h]or Antonio de Albuquerque Coelho de Carvalho. Based on an exhaustive vocabulary of the abbreviations in the document, we selected the 3,793 occurrences of suspension abbreviations. Then, using the theoretical-methodological framework of René Pellén (2005), we examined all the occurrences in depth, looking for rules that could be generalized, thus helping us to better understand not only the brachygraphic system of the time, but also the 18th century Portuguese language. We conclude that abbreviations by superscript letters are not done arbitrarily, and that there are recurring norms and formulas that can be quantified and given generalizing qualitative analyses.
This article provides a comprehensive analysis of the features of conveying national color through lexical connotations in Tölögön Kasymbekov’s novel “The Broken Sword”. Lexical connotation is a linguistic phenomenon that adds emotional, cultural, social, or historical shades of meaning to the basic (denotative) meaning of a word. The aim of the study is to reveal the inner world of the characters, the historical consciousness of the people, their spiritual values, and social norms through the linguistically symbolic and culturally loaded expressions found in the novel. Lexical connotations in the novel are classified into seven main categories: history and mythology, nomadic lifestyle, natural and regional names (toponyms), traditions and customs, social values, patriotism, and kinship relations. Each category is analyzed in depth with examples from the novel, and the emotional, social, and cultural meanings of the linguistic units are explained. Idiomatic expressions such as "Let’s finish it from the horse’s mane", "We burned for the people", "The important paper", "Let it be a blessing", and "Secret chest" are shown to reflect the Kyrgyz people's national worldview, social relationship system, and spiritual values in an artistic manner. The scientific significance of the study lies in demonstrating the Kyrgyz language’s artistic and expressive potential and in revealing the function of connotation in conveying national culture through linguistic methods. The practical significance of the article is to contribute to the strengthening of national identity, the preservation of linguistic heritage, and the transmission of cultural values to future generations. Moreover, this research can serve as a valuable source for the teaching of the Kyrgyz language, for the interpretation of literary works, and for studies exploring the relationship between language and culture. The novel is regarded as an artistic and linguo-cultural resource that reflects the spiritual world, historical memory, and cultural characteristics of the Kyrgyz people.
This paper examines how social media discourse affects the learning of the English language on the tertiary level in Lahore and how informal online communication influences the academic English of students in Lahore. The qualitative design was employed to gather data on the basis of semi-structured interviews with English language learners and teachers. Braun and Clarke’s (2006) framework was used to conduct thematic analysis on the transcribed data. The results indicate that though social media provides appropriate exposure to vocabulary, pronunciation, and language use in real life, the casualness of linguistic norms has a very strong impact on the students’ academic writing and their communicative accuracy. The participants claimed to use short forms, abbreviations, slang and mixed-language texting on a regular basis, and these transferred to essays and presentations. Other themes were distraction and a lack of studying discipline, the inability to stick to the formal register, and the misunderstanding of words acquired in unconfirmed online situations. Teachers also reported on the same lines, as they observed poor writing standards in academic writing and increased dependence on social-media-driven language patterns. Pedagogical mechanisms to counteract these effects were also determined in the study with the focus being on register awareness, purposeful digital task integration, and curriculum modernization. On the whole, the study has determined that the power of social media is twofold, both positive when moderated and negative when uncontrolled and recommends that informed teaching and learning methods are needed to help students balance between informal online communication and formal academic language.
This study aimed to explore the problems and strategies of translating paratexts in digital media through content analysis, comparing the English and Arabic paratextual elements of six media digital sources, including the BBC News website, Al Arabiya blog, and Reuters forum. The study identified several translation problems, including culturally specific terms, cultural references, conflicting verbal and paratextual components, and reframing the narrative functions of paratexts. The emergent themes revealed that 97% of the problems were attributed to terminology and lexical choices, whereas 83% were related to the integration of visual elements. Themes of cultural adaptation for governing the content of paratextual elements and reframing the narrative functions of paratexts each reached 80%. The emergent themes also revealed the strategies that translators use to address the translation of paratexts. These strategies include employing the translation adaptation approach (95%), applying functional translation (90%), using thematic analysis for intralinguistic subtitles (87%), using multimodal transcription (85%), and using transposition (82%). The study recommends replacing cultural references, idioms, or symbols with equivalents that resonate with the target audience. Functional equivalence ensures that translated paratexts serve the same purpose in the target language, adhering to cultural norms. The study highlights the distinction between traditional criteria for identifying paratexts and the functionality of translation in identifying digital paratexts. Understanding functional equivalence dynamics can help translators produce high-quality translations for the digital era. Thus, the study contributes to the fields of translation studies, translation and culture studies, and translation and information technology.
Фонетический аспект обучения иностранному языку играет важную роль в постановке произношения, развитии лексических и речевых умений, формировании навыка восприятия живой речи, в том числе в профессиональном контексте. Актуальность статьи заключается в необходимости анализа фонетических ошибок и создания перечня слов для модернизации содержания профессионально-ориентированной иноязычной подготовки в медицинском вузе. Материалом исследования послужили 100 наиболее частотных слов, выявленных в процессе долгосрочного педагогического наблюдения. Были применены методы контрастивного анализа, интроспекции, анализа фонетических ошибок, методического и статистического анализа. В качестве модели выбрано нормативное британское произношение (RP), являющееся стандартом в учебных материалах. Выявлено, что наибольшие затруднения для студентов представляют различия в фонетическом строе английского и русского языков (в частности, гласные звуки, их долгота и краткость), фонетическая интерференция русского и латинского языков, произношение определенных суффиксов, непроизносимые буквы, словесное ударение. В результате проведенного анализа и классификации слова были представлены в таблицах, которые могут быть использованы как самостоятельный учебно-методический материал и как база для разработки системы фонетических упражнений для работы над произношением на основе актуальной лексики. The phonetic aspect of a foreign language training exerts a pivotal influence on the pronunciation mastery, vocabulary and speaking skills, speech perception and spoken word recognition. Specifically, it is important for attaining the communication goal in a professional context. The relevance of the study lies in analyzing phonetic errors and compiling a list of words aiming to update the content of professionally oriented foreign language training at a medical university. The research is based on 100 most frequent words identified and selected in the long-term pedagogical observation. Methods of contrastive analysis, introspection, phonetic error analysis and methodological analysis were applied. The standard British pronunciation (RP), which is the set norm in educational materials, was chosen as the model. It is revealed that the greatest difficulties for students are differences in the phonetic structure of English and Russian languages (in particular, vowel sounds, their length variation), phonetic interference of both Russian and Latin languages, pronunciation of certain suffixes and endings of the verbs, silent letters, and lexical stress. Following the word classification carried out, they were systematized into tables, which can be used as independent educational and methodological visual aids and as the source for developing a system of phonetic exercises for teaching and practicing pronunciation targeted at the specific vocabulary.
Adverbs play a central role in structuring discourse, conveying speaker stance, and modifying propositional content. -Ly adverbs constitute up to 55% of common adverbs and are frequently used in academic prose. Attaining a nuanced grasp of adverbial usage in learner English, and of how closely Turkish learners’ patterns align with native-speaker norms, is crucial. In this regard, this paper examines the use of -ly adverbs by Turkish EFL learners of English in comparison to native speakers. The investigation relies on two corpora of novice academic English: The Turkish International Corpus of Learner English (TICLE) and the Louvain Corpus of Native English Essays (LOCNESS), and one corpus of expert academic English: BNC (British National Corpus), representing learner and native speaker writing. The frequencies, lexical choices, and the distribution of -ly adverbs were analyzed across three corpora. In addition, syntactic functions of the identified -ly adverbs were classified according to Quirk et al. (1985) and Hasselgård’s (2015) classifications. The analysis reveals that Turkish EFL learners rely on a narrower range of -ly adverbs, frequently using those associated with spoken rather than academic discourse, whereas native expert writers demonstrate a more varied and academically appropriate adverbial repertoire. Additionally, Turkish learners underuse most -ly adverb categories, particularly adjuncts and disjuncts, while overusing conjuncts and intensifiers. These findings highlight the gap between native and non-native academic writing, emphasizing the need for explicit instruction in the use of adverbials to develop a more advanced academic style.
Pragmatics, a fundamental branch of linguistics, focuses on the implicit rules that govern human interaction. It delves into how meaning is constructed, conveyed, and interpreted in everyday communication beyond literal language. This study explores the mechanisms of pragmatics in daily life, particularly how context, shared knowledge, and social conventions influence interactions. Key areas examined include speech acts, implicature, and politeness strategies, which highlight the dynamic nature of meaning-making in various social contexts. Speech acts, such as requests, apologies, and promises, illustrate how utterances perform actions beyond their lexical content. Implicature examines how individuals infer meaning that is not explicitly stated, relying on context and shared assumptions. Politeness strategies, which vary across cultures, play a crucial role in managing interpersonal relationships and ensuring smooth communication by balancing social expectations and individual intentions. Understanding these unspoken rules is essential for effective communication, as they often dictate the success of social interactions. Misinterpretations in pragmatics can lead to misunderstandings, social faux pas, or conflicts, emphasizing the importance of cultural and contextual awareness. Pragmatic competence is crucial in multilingual and multicultural settings, where differing conventions and norms can create unique challenges. By synthesizing theoretical insights from existing literature, this study highlights the pervasive influence of pragmatics in shaping human interaction. It underscores the necessity of pragmatic awareness in fostering interpersonal understanding, enhancing communication skills, and navigating the complexities of social life effectively.
Abstract This paper presents a novel framework for modeling role and task allocation in cooperative wheeled soccer robot systems by leveraging latent knowledge extracted from past collaborative interactions. Inspired by recent advances in heterogeneous multi-robot collaboration, the proposed method encodes a soccer team as a set of Multidimensional Relational Structures (MDRSs), capturing both temporal and spatial relations among robot roles, actions, and stimuli. A structured dataset, termed the Soccer Robot Collaboration Treebank (SRCT), is introduced to represent play-by-play histories of robot behaviors, parsed through a formal grammar to support structured learning. Probabilistic modeling and Non-Negative Tensor Decomposition (NTD) are applied to the resulting tensors, enabling robust inference and latent knowledge estimation even in scenarios with sparse data or communication loss. Simulated experiments using a team of wheeled soccer robots in the Webots environment demonstrate the system’s ability to dynamically reassign roles, reason over incomplete histories, and predict collaborative behaviors such as passing, defending, or role-switching. The results show that the proposed framework enhances both strategic flexibility and robustness, providing a foundation for real-time decision-making in robotic soccer under uncertainty.
The deployment of large language models (LLMs) across heterogeneous environments requires format-specific conversion, precision tuning, and consistent evaluation-tasks that are often fragmented across multiple tools. This work presents SOLO-Export, a unified command-line interface (CLI) framework for multi-format export and post-export benchmarking of causal LLMs. Having precision options for FP16 and INT8 where appropriate, the system supports the ONNX, TorchScript, Hugging Face, TensorFlow Lite, and TensorRT backends. Device-aware exports for both CPU and CUDA targets are made possible by a configuration-driven workflow that generates artifacts in a uniform directory structure. Each exported model is benchmarked using the Penn Treebank dataset by the integrated evaluation harness, which reports inference latency, token-level accuracy, and perplexity. According to experimental results, FP16 exports on GPU-oriented backends like TensorRT achieved up to 3.2 times lower latency than baseline FP32 models. On the top of that, with minimal impact on perplexity, storage size was reduced by more than 60% thanks to INT8 quantization. The combined approach reduces manual configuration overhead, speeds up deployment preparation, and ensures consistent performance insights across formats. This study demonstrates that a single, scalable pipeline can be effective.
Ce travail vise à explorer l’effet facilitateur d’une langue romane (français ou espagnol) et de l’anglais dans l’apprentissage lexical d’une autre langue romane (espagnol ou français) dans un contexte universitaire. D’une part, l’appartenance à la même famille des langues romanes facilite largement l’acquisition des connaissances linguistiques. D’autre part, un étudiant ayant réussi l’examen d’entrée à l’université est censé avoir atteint au moins le niveau indépendant (B2) en anglais, tandis qu’un étudiant spécialisé en langues étrangères peut atteindre un niveau autonome (C1), selon les normes du programme d’anglais dans le cadre de la scolarité obligatoire publiés par le ministère de l’Éducation de la République populaire de Chine (2018). Un tel niveau de compétence en anglais joue nécessairement un rôle non négligeable dans l’apprentissage d’une autre langue indo-européenne. L’étude repose sur une classification des catégories d’unités transparentes du point de vue d’apprenants sinophones, ainsi que sur la comparaison de trois lexiques: un lexique espagnol élaboré à partir des manuels scolaires destinés aux étudiants chinois; un lexique français construit selon la même logique; et un lexique espagnol extrait d’un corpus naturel issu du projet iRead4Skills, destiné aux apprenants natifs. Ces lexiques, différenciés par la langue cible et le public visé, offrent une vue d’ensemble pertinente pour analyser la transparence lexicale dans l’acquisition du vocabulaire roman chez les apprenants sinophones.
The object of consideration in this article is the peculiarities of functioning of texts in the state language of the Russian Federation in their official-legal variety in legal-linguistic coverage, the subject is the conflict potential of these texts created on the basis of the official style of the modern Russian language. The purpose of the study is to identify and typologize, based on the generalization of data from linguo-expert practice, the most conflict-provoking genres of official communication, the reasons for the communicative failures of developers of legislative and creators of administrative texts, the conflict potential of which, laid down at the development stage, becomes an obstacle to their effective functioning in society, as well as ways to minimize them in real law enforcement practice. The material of the study is more than 20 controversial texts of the official sphere of different levels that have undergone the procedure of linguistic research in non-governmental expert organizations of Altai Krai. The text material was analyzed based on a comprehensive linguistic analysis, including the method of semantic and structural-logical analysis of the statement, techniques for interpreting the meaning and construal of the meanings of the elements of the statement, as well as the method of linguistic analysis that helps to determine the contextual meaning of language units used in the statement, their lexical-semantic, grammatical and syntactic properties and functions. Reference to the regional and national Russian practice of linguistic research of legislative and administrative-managerial texts, as well as reviews of judicial practice on issues arising during the consideration of cases related to the interpretation of a controversial fragment of the text of a legislative act, made it possible to identify genre varieties of documents that have realized conflict potential during their application (85% ‒ texts of administrative genres, 5% ‒ legislative; 10% ‒ "semi-official" texts created by private individuals and lacking official status).), and to determine the reasons that reduce the effectiveness of legislative texts. Among the most pressing is the failure of developers to focus on the addressee, ignoring the ambiguity of possible interpretation of bulky constructions of legislative texts, admissible alternatives for securing the norm of the law in the process of its implementation, which leads to interpreting the document in everyday practice.
BACKGROUND: Public health crises are governed not only through policies but also through talk. Government press conferences are ritualized arenas where authorities construct meaning, claim competence, and manage domestic and international legitimacy. China's abrupt transition from "zero‑COVID" to a strategy of coexistence provides a critical case for examining how transnational pressures-from the World Health Organization, diplomatic partners, markets, and global media-shape official communication over time. MAIN BODY: This study analyzes 154 central government press‑conference transcripts (February 2020-February 2023) using a mixed‑methods design that combines topic modeling with qualitative frame analysis and process tracing of international pressure events. We segment the period into four phases-International Scrutiny, Global Cooperation, International Isolation, and Global Alignment-and identify seven recurring frames spanning health‑system capacity, epidemiological standards, vaccine diplomacy, economic-health trade‑offs, supply‑chain interdependence, and policy adaptation. Event‑timing analysis shows a consistent lag of roughly 7-21 days between major international cues and subsequent adjustments in domestic frames, with the economic-health and policy‑adaptation frames most responsive. A micro‑level discourse analysis demonstrates "semantic governance": lexical substitutions ("optimization," "new phase") and contextual recoding that converted a substantive policy reversal into a narrative of adaptive improvement. We argue that authorities achieved discursive alignment with evolving global norms without immediate policy convergence, illustrating how sovereignty sensitivities are managed communicatively. The findings also reveal equity‑relevant mechanisms: semantic smoothing that stabilizes compliance can under‑specify risks for vulnerable groups during transition windows, and generic references to "key populations" can displace time‑bound commitments to protection and access. Building on these insights, we propose two practical tools for global health governance: (1) an equity checkpoint for each policy pivot (plain‑language risk summaries, service guarantees, and a short equity note), and (2) a discursive alignment dashboard that tracks lead-lag to international guidance, domain‑specific alignment, and semantic markers of convergence or divergence. CONCLUSIONS: Pandemic communication in China followed a cyclical frame‑reinforcement pattern rather than a linear arc, and relied on semantic governance to manage rapid policy change under transnational pressure. Recognizing and monitoring these communicative mechanisms can strengthen global health governance and reduce equity risks during future protracted emergencies.
The article critically analyses the methodological practices of studying the lexicographic stratification of economic terms in English-Ukrainian translation dictionaries. Particular attention is paid to the analysis of the effectiveness of the key methods and techniques used in these methodological practices, in particular, the structural method and its technique, namely component analysis, as well as the functional method. The importance of quantitative and statistical methods is emphasized in such studies to determine the frequency of use of terms, which allows establishing the regularities of their functioning in texts of professional languages. The study also underscores the expediency of using the methods of unification, normalisation and cluster analysis, which contribute to the process of standardisation of industry terminology. The role of translation dictionaries for terminological practices is explored, with an emphasis on the importance of reflecting in them the ways of adapting authentic terms to the norms of the target language and the peculiarities of sectoral terminology, especially economic terminology. Preliminary assumptions suggest that the existing methodological practices do not have clear criteria for stratifying the terms of professional languages, which impacts the quality of the compiled lexicographic resources of the translation type. The proposed methodology includes four stages. The first stage involves the systematisation of terms according to the criteria of abstraction, scale, research object and sectoral differentiation. This makes it possible to classify terms depending on their level of generality and scope. At the second stage, the structural method is used to analyse the internal organisation of terms, including their phonetic, morphological, syntactic and lexical aspects. The third stage focuses on the semantic analysis of foreign language terms using distributional analysis to account for contextual differences. The fourth stage involves quantitative and statistical analysis using Zipf’s law to determine the frequency of terms. This helps to create frequency dictionaries that optimise the translation and adaptation of terminology. The developed methodology has the potential for further development of terminological translation lexicography and optimisation of interlingual professional communication processes.
Мақалада түркі тілдерінің синтаксистік құрылымын формалды грамматика тұрғысынан және заманауи аннотациялық модельдер негізінде сипаттаудың тәжірибесі қарастырылады. Синтаксистік аннотация тілдің грамматикалық жүйесін формалды түрде сипаттайтын және оны автоматты өңдеуге мүмкіндік беретін маңызды құрал ретінде танылады. Зерттеу барысында «Universal Dependencies» (UD), «MaTT» (Multilingual Aligned Treebank of Turkic) және «Kazakh Dependency Treebank» (KazDT) сияқты жобаларға сүйеніп, түркі тілдеріне тән морфологиялық және синтаксистік ерекшеліктер сипатталды. Синтаксистік белгіленім модельдері: «құрамдық», «аралас», «басыңқы-бағыныңқылық грамматикасы» т.б. тәсілдердің сипаты, ерекшеліктері, түркі тілдері үшін ұтымды тұстары мен кемшіліктері сараланды. Нәтижесінде басыңқы-бағыныңқы қатынастар грамматикасы негізінде жасалған синтаксистік аннотация моделі түркі тілінің құрылымын тиімді сипаттауға мүмкіндік беретіні дәлелденді. Басыңқы-бағыныңқы грамматикасының (басыңқы-бағыныңқы қатынастар) теориялық негіздері, синтаксистік аннотацияның форматы мен стандарттары сараланды. Түркі тілдерінің жалғамалы табиғаты мен еркін сөз тәртібінің «UD» сияқты әмбебап жобаларға бейімделуі талдауға түсті. Сонымен қатар, қазақ тілінің аннотацияланған корпустарын жетілдіру, автоматты парсинг, тілдік білім беру жүйесіне енгізу секілді болашақтағы бағыттары көрсетілді. Мақала түркі тілдерінің синтаксистік белгіленім тәжірибесі негізінде қазақ тілін цифрлық кеңістікке енгізудің маңызды қадамдарының бірі ретінде синтаксистік аннотацияны ғылыми тұрғыда негіздеуді мақсат етті. Түйін сөздер: түркі тілдері, синтаксистік аннотация, басыңқы-бағыныңқы грамматикасы, «UD», KazDT, формалды модельдер, парсинг.
We present a family of encodings for sequence labeling dependency parsing, based on the concept of hierarchical bracketing. We prove that the existing 4-bit projective encoding belongs to this family, but it is suboptimal in the number of labels used to encode a tree. We derive an optimal hierarchical bracketing, which minimizes the number of symbols used and encodes projective trees using only 12 distinct labels (vs. 16 for the 4-bit encoding). We also extend optimal hierarchical bracketing to support arbitrary non-projectivity in a more compact way than previous encodings. Our new encodings yield competitive accuracy on a diverse set of treebanks.
Background: Color plays a pivotal role in visual perception, shaping emotions, attention, and cognition, particularly in art-related contexts. However, the influence of artistic training on color perception and neural processing remains poorly understood.Methods: This study examined differences in color perception between art and non-art groups using behavioral ratings and EEG data. Forty-four participants (22 art majors: 21.82 ±1.56 years old; 22 non-art majors: 20.73 ± 1.67 years old) with an equal gender ratio were recruited. Participants completed color perception tasks involving cool, warm, and neutral hues while EEG data were recorded with a 65-electrode system. Behavioral ratings and ERP components (P2 and P3) were analyzed, supplemented by decoding analysis to uncover neural processing patterns.Results: Behavioral data indicated that warm hues elicited higher emotional valence ratings than cool and neutral hues for both groups. EEG analysis revealed that warm and cool hues evoked larger P3 amplitudes compared to neutral hues. A group-hue interaction was observed in the P2 component, with the non-art group showing greater variability in P2 amplitudes across hues. Decoding analysis provided further evidence of distinct neural processing differences between the two groups.Conclusion: These findings demonstrate that color perception differs between art and non-art groups, particularly in the neural processing of the P2 component. Warm and cool hues elicit stronger emotional and attentional responses, highlighting distinct cognitive mechanisms influenced by artistic expertise.Keywords: ERP; color perception; P2; P3; artistic training
The linguistic features of the Uzbek language - complex agglutinative morphology, free word order, and limited resources - necessitate a specialized approach and thorough research in the application of morphological and syntactic methods. Within the framework of the study, morphological analysis methods and syntactic analysis methods are reviewed based on scientific sources. Each section presents the existing advantages and disadvantages, experience of their use in the Uzbek language, as well as a comparative analysis with foreign languages. Rule-based methods, statistical models (HMM, CRF, etc.), Neural network-based approaches (BiLSTM-CRF, seq2seq) of morphological analysis in the Uzbek language are discussed, and the results are given in examples and percentages. It is shown that syntactic parsing is implemented using dependency and constituency parsing analysis methods. The issue of building a UD treebank for the Uzbek language with SOV order is considered. The impact of complex morphological structure and free word order in sentences on the construction of parsers is highlighted. As a result of the studied approaches, the issue of building hybrid parsers, integrating them with morphological analysis and assigning grammatical categories of words to the parser is raised. Also, the development of neural constituency parsers based on neural networks and the effectiveness of the results obtained from them are analyzed.
The article examines the evolution of Judaeo-Arabic translations, ranging from early biblical renditions to modern adaptations of European literature. Early pre-Saadian translations, written in a phonetic transcription, preceded Saʿadya Gaon’s Tafsīr (Bible translation), which adhered to ‘classical’ linguistic norms and became the authoritative translation for centuries. However, later translations introduced local dialectal features to meet the needs of diverse Jewish communities. The theoretical framework of ‘centre versus periphery’ is employed to analyse the dynamics of translation traditions, highlighting the interaction between cultural centres like Meknes, Morocco, and Constantine, Algeria, and their peripheries. By the nineteenth and twentieth centuries, Judaeo-Arabic translations extended to Haskala novels and French and English classics such as Robinson Crusoe, demonstrating the influence of global cultural trends. The study emphasises the dual role of these translations in preserving Jewish identity and adapting to contemporary linguistic and cultural shifts.
Deep neural networks employ specialized architectures for vision, sequential and language tasks, yet this proliferation obscures their underlying commonalities. We introduce a unified matrix-order framework that casts convolutional, recurrent and self-attention operations as sparse matrix multiplications. Convolution is realized via an upper-triangular weight matrix performing first-order transformations; recurrence emerges from a lower-triangular matrix encoding stepwise updates; attention arises naturally as a third-order tensor factorization. We prove algebraic isomorphism with standard CNN, RNN and Transformer layers under mild assumptions. Empirical evaluations on image classification (MNIST, CIFAR-10/100, Tiny ImageNet), time-series forecasting (ETTh1, Electricity Load Diagrams) and language modeling/classification (AG News, WikiText-2, Penn Treebank) confirm that sparse-matrix formulations match or exceed native model performance while converging in comparable or fewer epochs. By reducing architecture design to sparse pattern selection, our matrix perspective aligns with GPU parallelism and leverages mature algebraic optimization tools. This work establishes a mathematically rigorous substrate for diverse neural architectures and opens avenues for principled, hardware-aware network design.
Classical deviance theories about metaphor argue that the metaphorical sense of a word or expression, w, deviates from the sense of the word or expression interpreted literally. Developments in lexical pragmatics challenge these theories by claiming that deviance pervades (nearly) all aspects of linguistic communication. If deviance is the norm, then classical explanans offer little to no insight. In fact, many theorists have abandoned the idea of the literal-metaphorical distinction. This move carries significant consequences for theories of language and communication. We argue against this move and in favour of a linguistically robust literal-metaphorical distinction. We have three goals: The first is to argue that the literal-metaphorical distinction is important for theories of language and communication. The second is to assess Allott and Textor’s Non-Conformity View of deviance which gives up the idea that deviance is marked by a departure from conventional word meaning. Pace Allott and Textor, we claim that the literal-metaphorical distinction must ultimately be couched in some account of lexicalised meaning. Our third goal is to develop a form of qualified deviance that avoids what we take to be shortcomings of the Non-Conformity View.
Интервью продолжает тему перевода богослужения на современные языки (Вестник Свято-Филаретовского института. 2020. Вып. 36. С. 100–128). В предлагаемой публикации в первую очередь затрагиваются вопросы перевода на русский язык. Сравниваются возможные подходы к решению проблемы понимания смысла богослужения: перевод богослужения, подстрочный перевод, комментирование текста, изучение церковнославянского языка. В интервью приводятся примеры, показывающие, что трудности в понимании богослужения связаны не только с незнанием церковнославянского языка, но и с особенностями перевода оригинального греческого богослужения на церковнославянский язык, с библейской образностью богослужебных текстов. Перевод в наибольшей степени способствует действенности молитвы. Обсуждается наиболее подходящее выражение для обозначения языка, на который совершается перевод (современный русский язык, русский литургический язык, церковнорусский язык) и влияние языка богослужебных переводов на современную языковую ситуацию. Уделяется внимание таким проблемам, как принципы перевода, выбор источников для перевода, воз- можность лексической, синтаксической, прагматической, художественной эквивалентности перевода оригинальному тексту, сохранение в переводе привычных славянизмов, влияние богословского смысла текста на выбор того или иного переводческого решения и др. Особо говорится о влиянии перевода богослужебных текстов на устроение миссии и катехизации, жизнь церковных общин, вхождение новых людей в традицию Церкви. Священник Георгий Кочетков начал переводческую деятельность с 1970-х гг. в контексте взрослой катехизации. К. А. Мозгов и П. С. Озерский — филологи, участники работы по переводу православного богослужения, ведущейся в Свято-Филаретовском институте. На сегодняшний день вышел перевод всего корпуса неизменяемых богослужебных текстов, а также канона прп. Андрея Критского, избранных песнопений Октоиха и Постной Триоди, готовится перевод Цветной Триоди. Игумен Силуан (Туманов) занимается переводческой деятельностью с 2003 г., выпустил ряд книг с переводом древних литургий и отдельных богослужебных текстов. Протоиерей Георгий Иоффе — автор поэтического перевода Псалтири и других богослужебных текстов на русский язык. Священник Максим Плякин входит в рабочую группу Издательского совета РПЦ по кодификации акафистов и выработке норм акафистного творчества. The interview continues the theme of translating of Church worship services into modern languages (The Quarterly Journal of St. Philaret’s Institute, 2020, Iss. 36, pp. 100–128). The proposed publication primarily addresses the issues of translation into Russian. It compares the possible approaches to solving the problem of understanding the meaning of the divine service: translation of the divine service, literal translation, commentary on the text, and study of the Church Slavonic language. The interview provides examples showing that the difficulties in understanding the divine service are related not only to ignorance of the Church Slavonic language, but also to the peculiarities of the translation of the original Greek divine service into Church Slavonic, and to the biblical imagery of the divine service texts. Translation contributes most to the effectiveness of prayer. The interview discusses the most appropriate expression for the language into which the translation is made (modern Russian language, Russian liturgical language, Church- Russian language) and the influence of the language of liturgical translations on the contemporary linguistic situation. Attention is paid to such problems as the principles of translation, the choice of sources for translation, the possibility of lexical, syntactic, pragmatic, artistic equivalence of the translation to the original text, the preservation of familiar Slavicisms in translation, the influence of the theological meaning of the text on the choice of a particular translation solution, and others. Special mention is made of the influence of the translation of liturgical texts on the organisation of mission and catechesis, the life of church communities, and the entry of new people into the tradition of the Church. Priest Georgy Kochetkov began translating Church worship services in the 1970s in the context of adult catechesis. К. A. Mozgov and P. S. Ozersky, professional philologists, were members of his translation group in different years. To date, the translation of the entire corpus of unchanging liturgical texts has been prepared, as well as the Canon of St. Andrew of Crete, selected hymns of the Octoechos, and the translation of the Lenten and Colored Triodion is being prepared. Hegumen Siluan (Tumanov) has been engaged in translation work since 2003, and has published a two-volume book of translations of ancient liturgies and individual liturgical texts. Archpriest George Ioffe is the author of a poetic translation of the Psalms into Russian. Priest Maxim Pliakin is a member of the working group of the Publishing Council of the Russian Orthodox Church on the codification of akathists and the development of norms for akathist creativity.
We investigate the performance of state-ofthe-art (SotA) neural grammar induction (GI) models on a morphemically tokenised English dataset based on the CHILDES treebank (Pearl and Sprouse, 2013).Using implementations from Yang et al. (2021b), we train models and evaluate them with the standard F1 score.We introduce novel evaluation metrics-depth-ofmorpheme and sibling-of-morpheme-which measure phenomena around bound morpheme attachment.Our results reveal that models with the highest F1 scores do not necessarily induce linguistically plausible structures for bound morpheme attachment, highlighting a key challenge for cognitively plausible GI.
Emojis are widely used in digital communication to convey emotional cues alongside text, yet their impact on word-level reading within sentence contexts remains unclear. We conducted an eye-tracking experiment to examine how positive (e.g., 🤩) versus neutral (e.g., 🧑🦳) face emojis embedded mid-sentence in otherwise neutral sentences affect the processing of the preceding and following words (e.g., positive “Did you change your hair 🤩 something is different” vs. neutral “Did you change your hair 🧑🦳 something is different”). We observed robust parafoveal-on-foveal (PoF) effects on the n–1 word, with longer fixations in first-fixation, gaze duration, and single-fixation measures when the parafoveal emoji was positive rather than neutral. This valence effect persisted even after accounting for mislocated fixations, suggesting that positive emotional content genuinely modulates foveal word processing. In contrast, the n+1 word showed no valence-based facilitation, implying that the influence of a positive mid-sentence emoji does not extend to subsequent words in continuous reading. At the sentence level, positive emojis were associated with faster overall reading times and higher valence ratings, although dashed (no-emoji) sentences in the pre-test were rated more positively than emojified versions in the experiment. These findings reinforce models of eye movement control that allow parallel processing of foveal and parafoveal information, highlighting how affective face emojis can shape real-time reading dynamics.
This paper presents the structure and principal components of the linguistic resources required for sentiment analysis in the Uzbek language. The research aims to identify and develop effective approaches for constructing a linguistic database - referred to as SentiUzNet - and to establish a foundational sentiment lexicon tailored specifically to the characteristics of the Uzbek language. In particular, the paper discusses key principles for annotating words with sentiment polarity and subjectivity scores, as well as methodological foundations for building a lexicographic database to support automated emotional analysis of texts. A significant part of the research focuses on experimenting with large-scale user-generated content, specifically social media comments written in Uzbek. These datasets were used to train and evaluate sentiment analysis models, thereby allowing an assessment of their performance and practical applicability. The results of this research represent one of the first comprehensive attempts to facilitate automatic sentiment detection in the Uzbek language and are expected to contribute substantially to the advancement of natural language processing technologies in under-resourced linguistic settings.
The purpose of this study is to evaluate the English language proficiency of Sayed Jamaluddin Afghani University students enrolled in the English Department. It primarily looks at the grammatical mistakes that impede kids' language development and how they affect their ability to communicate. A structured questionnaire was used in a quantitative survey with a sample of 120 students in order to collect data. The results show that improper application of linguistic norms, infrequent practice, and the effect of students' native language, which frequently interferes with English usage, are the primary causes of grammatical errors. The efficacy and clarity of students' speech are greatly impacted by these errors. According to the study's findings, English professors should implement specialized courses that emphasize communicative grammar instruction. It also recommends giving students regular chances to practice active language usage, which can improve their overall communication skills and grammatical precision, ultimately increasing their English proficiency.