Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Deep neural networks employ specialized architectures for vision, sequential and language tasks, yet this proliferation obscures their underlying commonalities. We introduce a unified matrix-order framework that casts convolutional, recurrent and self-attention operations as sparse matrix multiplications. Convolution is realized via an upper-triangular weight matrix performing first-order transformations; recurrence emerges from a lower-triangular matrix encoding stepwise updates; attention arises naturally as a third-order tensor factorization. We prove algebraic isomorphism with standard CNN, RNN and Transformer layers under mild assumptions. Empirical evaluations on image classification (MNIST, CIFAR-10/100, Tiny ImageNet), time-series forecasting (ETTh1, Electricity Load Diagrams) and language modeling/classification (AG News, WikiText-2, Penn Treebank) confirm that sparse-matrix formulations match or exceed native model performance while converging in comparable or fewer epochs. By reducing architecture design to sparse pattern selection, our matrix perspective aligns with GPU parallelism and leverages mature algebraic optimization tools. This work establishes a mathematically rigorous substrate for diverse neural architectures and opens avenues for principled, hardware-aware network design.
The linguistic features of the Uzbek language - complex agglutinative morphology, free word order, and limited resources - necessitate a specialized approach and thorough research in the application of morphological and syntactic methods. Within the framework of the study, morphological analysis methods and syntactic analysis methods are reviewed based on scientific sources. Each section presents the existing advantages and disadvantages, experience of their use in the Uzbek language, as well as a comparative analysis with foreign languages. Rule-based methods, statistical models (HMM, CRF, etc.), Neural network-based approaches (BiLSTM-CRF, seq2seq) of morphological analysis in the Uzbek language are discussed, and the results are given in examples and percentages. It is shown that syntactic parsing is implemented using dependency and constituency parsing analysis methods. The issue of building a UD treebank for the Uzbek language with SOV order is considered. The impact of complex morphological structure and free word order in sentences on the construction of parsers is highlighted. As a result of the studied approaches, the issue of building hybrid parsers, integrating them with morphological analysis and assigning grammatical categories of words to the parser is raised. Also, the development of neural constituency parsers based on neural networks and the effectiveness of the results obtained from them are analyzed.
Emojis are widely used in digital communication to convey emotional cues alongside text, yet their impact on word-level reading within sentence contexts remains unclear. We conducted an eye-tracking experiment to examine how positive (e.g., 🤩) versus neutral (e.g., 🧑🦳) face emojis embedded mid-sentence in otherwise neutral sentences affect the processing of the preceding and following words (e.g., positive “Did you change your hair 🤩 something is different” vs. neutral “Did you change your hair 🧑🦳 something is different”). We observed robust parafoveal-on-foveal (PoF) effects on the n–1 word, with longer fixations in first-fixation, gaze duration, and single-fixation measures when the parafoveal emoji was positive rather than neutral. This valence effect persisted even after accounting for mislocated fixations, suggesting that positive emotional content genuinely modulates foveal word processing. In contrast, the n+1 word showed no valence-based facilitation, implying that the influence of a positive mid-sentence emoji does not extend to subsequent words in continuous reading. At the sentence level, positive emojis were associated with faster overall reading times and higher valence ratings, although dashed (no-emoji) sentences in the pre-test were rated more positively than emojified versions in the experiment. These findings reinforce models of eye movement control that allow parallel processing of foveal and parafoveal information, highlighting how affective face emojis can shape real-time reading dynamics.
Multilingual Large Language Models (LLMs) have shown remarkable performance across various languages; however, they often include significantly less data for low-resource languages such as Urdu compared to high-resource languages like English. To assess the linguistic knowledge of LLMs in Urdu, we present the Urdu Benchmark of Linguistic Minimal Pairs (UrBLiMP) i.e. pairs of minimally different sentences that contrast in grammatical acceptability. UrBLiMP comprises 5,696 minimal pairs targeting ten core syntactic phenomena, carefully curated using the Urdu Treebank and diverse Urdu text corpora. A human evaluation of UrBLiMP annotations yielded a 96.10% inter-annotator agreement, confirming the reliability of the dataset. We evaluate twenty multilingual LLMs on UrBLiMP, revealing significant variation in performance across linguistic phenomena. While LLaMA-3-70B achieves the highest average accuracy (94.73%), its performance is statistically comparable to other top models such as Gemma-3-27B-PT. These findings highlight both the potential and the limitations of current multilingual LLMs in capturing fine-grained syntactic knowledge in low-resource languages.
This repository contains data accompanying the publication "Auditory localization and subjective assessment of autonomous cleaning robot sounds: A VR experiment on speed, operating mode and alerting signals", submitted for review to the Acta Acustica. The dataset contains: Audio and video material Stimuli consisting of robot recordings under all evaluated conditions (0.3 m/s and 0.8 m/s speed, with and without cleaning, with and without noise AVAS or multi-tone AVAS, both with and without added amplitude modulation). All sounds were exported as 32-bit float wav files; i.e., reading the files into Matlab with audioread results in calibrated Pa values. The files uploaded here were used as source signals in the auralization, assuming a distance of 1 m. The final binaural stimuli were rendered by TASCAR and include an attenuation corresponding to the simulated 7 m distance. 30cms_cleaning_noAVAS.wav 30cms_noCleaning_multiTone.wav 30cms_noCleaning_multiToneAM.wav 30cms_noCleaning_noAVAS.wav 30cms_noCleaning_noise.wav 30cms_noCleaning_noiseAM.wav 80cms_cleaning_noAVAS.wav 80cms_noCleaning_multiTone.wav 80cms_noCleaning_multiToneAM.wav 80cms_noCleaning_noAVAS.wav 80cms_noCleaning_noise.wav 80cms_noCleaning_noiseAM.wav ambienceNoise.wav Excerpt of background noise played back during the experiment. localizationTaskDemo.mp4 Participant POV recording of localization task. This recording was done with a fixed head position, in the actual experiment participants were turning their heads freely. Experiment results and analysis localizationData.csv Table containing the mean and standard deviation of absolute localization error, aggregated for each participant and stimulus. subjectiveData.csv Table containing mean and z-scored annoyance, arousal, trust, and valence ratings for each participant and stimulus. stimuliAnalysis.csv Table containing results of level, loudness, sharpness, roughness, tonality, fluctuation strength, and impulsiveness analysis for all stimuli.
This study explores the syntactic network characteristics of English e-commerce live-streaming discourse by employing a syntactic treebank and syntactic complex network analysis. The main findings are: (1) The syntactic network of English e-commerce live-streaming discourse exhibits small-world and scale-free properties, which are hallmark traits of complex networks. (2) The central nodes of the network are be, I, and the, with be serving as the most central node, while I and the act as local central nodes. (3) The central node be demonstrates both strong centrifugal and centripetal forces. Its centrifugal force is most frequently associated with subject relations and adjective complements, while its centripetal force is characterized by auxiliary and clausal complements. These findings indicate that the syntactic structure of English e-commerce live-streaming discourse is highly robust. This robustness underscores the discourse’s functional purpose: to convey information clearly while engaging users through personalization and specificity. Furthermore, the study highlights the critical role of be in attributive and descriptive constructions. Overall, this research provides insights into the syntactic organization of e-commerce discourse and demonstrates the effectiveness of complex network analysis in linguistic studies.
BACKGROUND AND OBJECTIVES: It is well documented that the fear of specific stimuli and situations can be acquired through the social observation of the actions of another person. In contrast, it is still a matter of debate, whether processes related to fear attenuation, extinction, and extinction-retrieval can equally be achieved through social observation after de novo fear conditioning. METHODS: Here, we used a differential fear conditioning procedure and investigated whether the variation of the context of video-based vicarious extinction learning (VEL) will affect subsequent extinction learning and extinction-retrieval. Conditioned fear acquisition, extinction, and extinction-retrieval was measured using psychophysiological (skin conductance responses) and subjective measures (CS-UCS contingency ratings and CS-valence ratings). RESULTS: Participants showed enhanced fear extinction learning after VEL as compared to controls. VEL improved extinction learning relative to controls but appeared to be highly context-dependent. The beneficial effect of VEL on subsequent extinction learning was abolished when the context in which the model was performing in the video was different from the context in which the observer performed all stages of the experiment. LIMITATIONS: Data were obtained in a non-clinical sample which does not permit the extrapolation of findings to clinical populations. CONCLUSION: Our results suggests that safety information derived from VEL promotes fear extinction when model and observer perform the experiment in the same context. Given that fear extinction is considered as an experimental proxy of exposure therapy, our findings might be instructive for the development of novel clinical interventions to promote exposure treatment efficacy.
Abstract The age-related positivity bias refers to the finding that older adults recount events more positively (or less negatively) as compared to younger adults (i.e., a main effect of age on memory valence). This bias is closely related to the positivity effect, which reflects an interaction between age and valence of information to be remembered. We examined the age-related positivity bias and positivity effect using a one-year longitudinal design with a sample that spanned adulthood (N = 374; age range 19-90; M= 47.41; SD= 16.75). Participants answered questions regarding their memories of learning about the outcome of the 2020 U.S. presidential election. Analyses examined the association between age and valence ratings (positive, negative) and ratings of feelings (happy, elated, upset, and shaken) at Time 1, as well as the association with age between change scores for each of those variables, while controlling for who the participant voted for in the election. Results indicate that increased age was associated with reporting feeling less negative at the time of the event, and also remembering feeling more positive (elated and happy) when reconstructing the event one year later, thereby providing evidence of the positivity bias. There was no evidence of an age by valence interaction in a 2 (Valence) x 3 (Age) mixed ANCOVA on the positive and negative change scores, indicating there was not a positivity effect. Depressive symptoms partially mediated the relationship between age and valence variables, indicating that depressive symptoms may be one mechanism for explaining the age-related positivity bias.
Background: Many previous studies highlighting a relationship between depression and emotional face recognition have relied on measures of classification accuracy to determine recognition deficits. However, the perception of emotions is also related arousal levels and valence, and more research is needed to determine how depression impacts these dimensions.Aims: To compare performance on both an objective forced choice emotional recognition task and subjective emotional face valence rating task in participants with self-reported high depression.Methods: Based on screening using the depression sub-scale of the DASS-42, 46 participants (23 males, 23 female) were in the high depression group (mean DASS-42 34±5) and 50 participants in the control groups (25 males, 25 females) with DASS-42 scores of either 0 or 1. All participants completed both a performance-based task (objective) as well as a rating task (subjective) of emotional facial expressions. Results: The data indicate that difference in performance exist in classification accuracy between the groups, with depressed participants demonstrating reduced accuracy for anger, sadness and neutral facial expressions. Additionally differences in subjective ratings exist in the depressed group, but with the important caveat that these only relate to faces display positive emotional expressions.Discussion: The limitations of relying solely on objective tasks where recognition accuracy is the main outcome measure are discussed as well as the data quantitatively demonstrating a reduced response in the depression group to positive stimuli. This study justifies the need for future studies using both objective and subjective measures to assess emotion classification deficits in depression.
Emojis are widely used in digital communication to convey emotional cues alongside text, yet their impact on word-level reading within sentence contexts remains unclear. We conducted an eye-tracking experiment to examine how positive (e.g., 🤩) versus neutral (e.g., 🧑🦳) face emojis embedded mid-sentence in otherwise neutral sentences affect the processing of the preceding and following words (e.g., positive “Did you change your hair 🤩 something is different” vs. neutral “Did you change your hair 🧑🦳 something is different”). We observed robust parafoveal-on-foveal (PoF) effects on the n–1 word, with longer fixations in first-fixation, gaze duration, and single-fixation measures when the parafoveal emoji was positive rather than neutral. This valence effect persisted even after accounting for mislocated fixations, suggesting that positive emotional content genuinely modulates foveal word processing. In contrast, the n+1 word showed no valence-based facilitation, implying that the influence of a positive mid-sentence emoji does not extend to subsequent words in continuous reading. At the sentence level, positive emojis were associated with faster overall reading times and higher valence ratings, although dashed (no-emoji) sentences in the pre-test were rated more positively than emojified versions in the experiment. These findings reinforce models of eye movement control that allow parallel processing of foveal and parafoveal information, highlighting how affective face emojis can shape real-time reading dynamics.
This article attempts to investigate the complicated relationship of form of language, narrative vagueness, and cultural interpretation within Frank Stockton’s “The Lady, or the Tiger?” Applying text linguistic methods to the analysis of this classic short story, the researcher tackles two research questions, specifically, how does the linguistic structure of the narrative reinforce the uncertain ending of “The Lady, or the Tiger?” What are the effects of cultural and language influences on how the reader interprets the important characters and themes of “The Lady, or the Tiger?” Grounded in theoretical frameworks from Halliday and Hasan’s cohesion and coherence, and more recent cognitive linguistics, the paper addresses how the deployment of ambiguity within language functions as a reader engagement and theme exploration strategy. The research concludes that the narrative’s linguistic construction—defined by its deliberate plot of unresolvable conflict, indirect characterization, and sophisticated temporalities—is a significant factor in the development of the legendary vagueness of the story. The research further contends that cultural contexts and linguistic norms governing the reader’s interpretation heavily influence the apparently moral, motivational, and destined lives of characters and thus the broader themes of justice, choice, and human nature. Finally, the paper concludes that “The Lady, or the Tiger?” illustrates the outstanding power of language in creating narrative doubt and highlights the importance of cultural spectacles in the reception and interpretation of literature. In this sense, the story illustrates the general implications of text linguistics in the understanding of narrative ambiguity and cultural impact.
The article examines speech culture as a key component of language competence among higher education students. The author emphasizes that mastering the norms of the literary language, adhering to ethical and stylistic standards in communication, is an indicator not only of a person’s general education but also of their readiness for professional and social interaction. The main components of speech culture are analyzed, including accuracy, clarity, logic, appropriateness, purity, expressiveness, and aesthetic quality of speech. Particular attention is paid to common violations of linguistic norms observed in the student environment: the use of colloquial, slang, and foreign words without necessity, unjustified calques, bureaucratic expressions, as well as syntactic and orthoepic errors. The article outlines the main causes of linguistic carelessness, such as low reading culture, the influence of social media, and the decline in linguistic standards in everyday and educational communication. The author proposes a number of pedagogical and methodological strategies aimed at cultivating a high level of speech culture among students. These include the integration of communicative training into the educational process, regular involvement of students in stylistic text analysis, and the activation of creative language practices. Examples of typical speech situations are provided to demonstrate the contrast between cultured and uncultured language use, highlighting the importance of speaker selfreflection in improving overall language competence. The relevance of the study is due to the growing importance of speech culture in the modern educational environment, where effective communication is a key component of a specialist’s professional training.
The nouns of our language refer to either concrete entities (like a table) or abstract concepts (like justice or love), and cognitive psychology has established that concreteness influences how words are processed. Accordingly, understanding how concreteness is represented in our mind and brain is a central question in psychology, neuroscience, and computational linguistics. While the advent of powerful language models has allowed for quantitative inquiries into the nature of semantic representations, it remains largely underexplored how they represent concreteness. Here, we used behavioral judgments to estimate semantic distances implicitly used by humans, for a set of carefully selected abstract and concrete nouns. Using Representational Similarity Analysis, we find that the implicit representational space of participants and the semantic representations of language models are significantly aligned. We also find that both representational spaces are implicitly aligned to an explicit representation of concreteness, which was obtained from our participants using an additional concreteness rating task. Importantly, using ablation experiments, we demonstrate that the human-to-model alignment is substantially driven by concreteness, but not by other important word characteristics established in psycholinguistics. These results indicate that humans and language models converge on the concreteness dimension, but not on other dimensions.
This essay deals with two colour-related adjectives, badius and baietus, in the Medieval Latin documentary sources of Catalonia studied by the Glossarium Mediae Latinitatis Cataloniae (GMLC). The collection of documental testimonies has been conducted through the lexical database Corpus Documentale Latinum Cataloniae (CODOLCAT), a digital corpus of the Latin texts of this linguistic domain. These sources of notarial and juridical nature contain a considerable amount of colour adjectives, which usually serves to identify and differentiate lexical elements within the same referential class. One of the most attested colour adjectives is badius “bay, brown”, frequently documented with its variant baius and always referring to equines. There are also few occurrences of baietus “brownish”, a lexical innovation derived from badius. To refine their precise definitions, this study explores the forms and uses of both adjectives, taking into consideration the contexts in which they appear. Additionally, this study highlights the importance of the integration of digital tools into lexicographical research and emphasizes the need to incorporate insights from other linguistic domains to achieve a more comprehensive understanding of the words under analysis.
The article explores orthographic interference in the process of improving the normative tools of national writing. Orthographic interference is a linguistic phenomenon that arises from the interaction of different language systems, manifesting as deviations or violations of writing norms. The study analyzes changes and adaptations of linguistic norms and their impact on the writing skills of language learners. It focuses on identifying linguistic mechanisms to improve the national writing system by preventing and reducing orthographic interference. The study employs content analysis, comparative methods, and qualitative analysis. Over 60 students’ written works were examined to determine the frequency, typology, and causes of orthographic errors. The analysis identified common types of interlingual and intralingual interference. Interlingual interference results from the influence of Russian and English graphic, phonological, and morphological features, while intralingual interference arises from inadequate understanding of the phonetic and phonological foundations of the language. Content analysis identified the most frequent orthographic errors, with primary causes including insufficient mastery of spelling rules, differences between native and target language writing systems, and teaching method shortcomings. These errors serve as indicators of students’writing experience and language proficiency. The identified types of interference contribute to improving normative tools, assessing the effectiveness of educational programs and methodologies, standardizing orthographic norms, and enhancing writing culture in multilingual societies. The findings support the codification of Kazakh orthographic norms and the development of a scientific and methodological foundation for reducing linguistic interference.
Despite the growing use of NLP in second language (L2) research, model accuracy in L2 settings remains underexplored. This study addresses this gap by evaluating and fine-tuning a Korean language model to extract morphosyntactic features (i.e., morpheme tokenization/tagging and dependency parsing) from L2-Korean texts. We begin by evaluating a domain-general Korean language model on a gold-annotated L2-Korean treebank. We then fine-tune the model on L2-Korean data and quantify the resulting gains across diverse L1- and L2- datasets. Finally, we examine how model reliability varies with learner proficiency scores. Three key findings emerge: while the domain-general model excels at morpheme tokenization, it underperforms on morpheme tagging and dependency parsing; fine-tuning substantially improves adaptability to L2 morphosyntax; and proficiency has minimal effect on morpheme-level tasks but significantly affects dependency-parsing reliability. These results highlight the importance of incorporating L2 training data to improve morphosyntactic analysis in L2 settings and caution against uncritical reliance on automated dependency annotations, especially when performance varies across proficiency levels.
This paper examines the phonetic value of the semivowel in Gavrilo Stefanović Venclović’s Služabna knjiga. The work was written between 1711 and 1716 in the Old Church Slavonic language using old Cyrillic script. The script is of the semi-uncial type with elements of cursive writing. The analysis is based on photographs of the manuscript (РГБ, Собрание Н. П. Румянцева Ф. 256 № 401). The text was transcribed using the Transkribus software platform, while most of the excerpted examples were processed using the AntConc software. The main findings of the research can be summarized as follows: (a) semivowel signs and the apostrophe at the end of a word have a purely orthographic function (e.g. даръ, хотѣщимь, готов); (b) as in the vernacular, in the Serbian Slavonic language epoch, the semivowel in a strong position generates the reflex a (e.g. вѣнацъ, диванъ, четвртакъ); (c) the secondary semivowel is vocalized as a (e.g. оганъ, петарь, жизан), or not recorder at all (e.g. жизнъ, огнъ, помыслъ); (d) the semivowel in a weak position, according to the rules of the Serbian Slavonic linguistic norm, is pronounced as a within the prepositions въ (e.g. въ бꙑтїе), съ (e.g. съ нами), and къ (e.g. къ г҃ꙋ), prefixes въ- (e.g. въходиⷮ), въз- (e.g. въⷥдиханїе), and съ- (e.g. съдѣлавъ), and the root въс- (e.g. въсакаа), as evidenced by examples where the letter а appears in place of the former weak position semivowel (e.g. васака, множаствѣ, саблюденїе)
The study examines how children, parents and staff in a kindergarten create a social space in the kindergarten’s cloakroom through (linguistic) actions and language choices. Based on one year of ethnographic fieldwork, which includes participant observation, documentation of the kindergarten’s semiotic landscape, field conversations and interviews with staff, the study shows how the cloakroom becomes a site for multilingual practices, while the kindergarten otherwise is dominated by a Norwegian language norm. The study demonstrates how children and adults take ownership of this space and create a multilingual environment through their actions. The analysis is grounded in Lefebvre’s theory of the production of social space – through spatial practice, representations of space and lived space – Gadamer’s perspectives on play and Pascual-de-Sans’ concept of idiotopy. The findings reveal that language choices and the construction of the cloakroom as a social space are influenced by complex processes related to place, time and agency. When the cloakroom is less in focus for the staff, it can become an important site for children’s play, where they negotiate both linguistic norms and other rules. The study argues for the significance of such social spaces as part of the kindergarten’s linguistic environment, where children can take ownership and make language choices that deviate from established majority language norms. At the same time, the study highlights the importance of time in research on language and place, both theoretically and methodologically, as well as the roles of researchers in gaining access to such spaces through invitations from the children.
This paper presents a real-time American Sign Language (ASL) recognition system utilizing a hybrid deep learning architecture combining 3D Convolutional Neural Networks (3D CNN) with Long Short-Term Memory (LSTM) networks. The system processes webcam video streams to recognize word-level ASL signs, addressing communication barriers for over 70 million deaf and hard-of-hearing individuals worldwide. Our architecture leverages 3D convolutions to capture spatial-temporal features from video frames, followed by LSTM layers that model sequential dependencies inherent in sign language gestures. Trained on the WLASL dataset (2,000 common words), ASL-LEX lexical database (~2,700 signs), and a curated set of 100 expert-annotated ASL signs, the system achieves F1-scores ranging from 0.71 to 0.99 across sign classes. The model is deployed on AWS infrastructure with edge deployment capability on OAK-D cameras for real-time inference. We discuss the architecture design, training methodology, evaluation metrics, and deployment considerations for practical accessibility applications.
Large language models (LLMs) have gained a lot of attention and achievements recently because of their significant comprehension and generative abilities. However, the large-scale parameters of LLMs require considerable computational resources in the training and inference process, which restricts their wide application. To overcome this challenge, we propose an efficient mixed precision weight quantization (EMWQ) method for LLMs in this article. Specifically, we introduce a new outlier detection method by analyzing the weight distribution instead of the conventional weight magnitude. Then, we propose a dual-quantization strategy that quantizes both the outlier critical columns and the residual matrices with different precision. Besides, we introduce two effective EMWQ-based application frameworks, the EMWQ-R and EMWQ-O in our study. Comprehensive experiments are conducted on the Penn Treebank (PTB), C4, ARC-Easy datasets, and MMLU benchmark across various tasks. The comparison results demonstrate that the proposed EMWQ achieves state-of-the-art performance in mixed precision quantization and further reduces computational memory cost. Besides, it has higher generalizability compared with conventional methods.
The article discusses multi-senses connectives and their annotation in text corpora.The author clarifies the concept, which is usually applied to three different phenomena: i) the uncertainty of annotators in choosing one of the meanings of a polysemic connector; ii) the possibility of establishing more than one relation, both explicit and implicit, between two fragments of text; iii) the "combination" of several meanings by a connective.The author then considers how these cases are annotated in the Penn Discourse Treebank and in the Supracorpora Database of Connectives, which use a multi-label annotation for discourse relations.The study also raises a number of theoretical questions: what information can we obtain from double labels, which relations are distributionally close, i.e. can be established in the same contexts, how to separate the contribution of the connective and of the context to the overall meaning of a sequence of sentences, how legitimate is it to talk about the combination of meanings by the connective, how to annotate the features of the context of the connectives so that this data can be used for Natural Language Processing?The solution to these questions is important for both theoretical cognitive and applied research (in particular, for machine learning and machine translation).
This research was conducted within the framework of gender linguistics, the main task of which is to study how grammatical categories related to biological sex are reflected in language and affect the perception of men and women in the minds of speakers of a particular language. The research paper focuses on the use of feminatives in Russian and Greek among bilinguals and monolinguals. The aim of the study was to identify trends in the use of feminine correlates of the names of professions and job positions in legal and political fields, as well as to determine the influence of the linguistic environment on the formation of linguistic norms regarding gender terminology. The authors employed an experimental approach based on a comparative analysis of the use of feminatives among Russian-Greek natural bilinguals aged 16-25 living in Cyprus and learning English as a foreign language, and artificial bilinguals, native speakers of Russian/Greek who also speak English. The experiment participants were given English sentences containing professional designations. The task set for them was to translate the sentences in such a way that the agentive subject or addressee was indicated as a female; after that a systematic analysis of the use of feminatives and their derivational models was conducted. This analysis revealed gender asymmetry in both linguistic contexts, reflecting the unequal representation of male and female genders in the linguistic consciousness of the speakers. The main conclusions emphasize the importance of the linguistic environment and cultural factors in shaping gender-specific lexicon used in everyday communication and media, and indicate the presence of interference in the speech of bilinguals.
Abstract Written culture in medieval Scandinavia featured the concurrent use of different languages–Latin and the local vernaculars–and two scripts–the Roman script and the runic script. Although Latin carried substantial cultural capital as the cosmopolitan language of religious and juridical authorities, diplomacy, and high literate culture, the vernacular runic written tradition had already been established for many centuries and continued to be used alongside Latin and the Roman script. This coexistence was particularly evident in the epigraphic landscape and resulted in several inscriptions that mixed both languages and scripts. This article explores the choices of language and script made by two medieval inscribers who, despite being primarily trained within a vernacular runic tradition, used Latin and/or the Roman script, particularly in their signatures. The study argues that these inscribers leveraged the symbolic value of Latin and the Roman script to project a distinctive professional identity in response to a changing linguistic market in which these linguistic resources were gaining value. The analysis also illustrates how these inscribers’ highly individual practices resulted from a negotiation of different linguistic norms and highlights a tension between their level of literacy and their efforts to display cultural capital. Finally, the article argues that the use of Latin and the Roman script was motivated more by their symbolic value than by purely communicative necessity, offering parallels to modern instances of language commodification.
a Real close relationship exists between transferring and public media, imposing a noticeable light on the importance of language especially translation in casting discourse media since translation goes on an essential role in representing the diversity between cultures and perspectives in media, however translation may distorting the culture disparities and context unless not managed cautiously,in this sense, translators have to create a balance between loyalty to the original texts and the target ones taking into consideration the diversity in linguistic and cultural norms, this balance can be molded by applying certain approaches and strategies, this article aims to present theoretical and analytical paths to explore the diverse techniques and mechanisms used in transferring media discourse which are important to be adopted in the process of translation given that translation in the media can play a really malicious role in spreading misinformation and tendentious narratives, faulty or biased translation effects the public perception dramatically and fabricates stereotypes of misconceptions, accordingly, this study deals with real world media discourse and examines the strategies adopted by the translators, the findings show that the process of transferring from one language to another comprises modifying the linguistic norms and cultural backgrounds,spotting the light on power and responsibility that emerge during the process of translation since societies become increasingly interconnected so it is an urgent matter for translators and journalists to maintain ethical, cultural and linguistic standards to preserve accuracy and solidity of translated content and to protect the public perception from distorted narratives and misinformation as much as possible.
This study examines the linguistic landscape (LL) of Chikan Old Street, a historic district in Zhanjiang, Guangdong, through the lens of the SPEAKING model. As one of the most well-preserved historical districts in southern China, Chikan Old Street embodies a rich maritime heritage and commercial traditions, making it a compelling site for LL research. The study investigates the interaction between official language policies, regional linguistic identity, and globalization, providing insights into how language hierarchies are constructed in heritage sites. A mixed-methods approach is employed, integrating quantitative corpus analysis, qualitative semiotic interpretation, and public perception surveys. The findings reveal a clear stratification of language use: Chinese dominates official signage, with pinyin and English in subordinate positions, reflecting state-imposed linguistic norms. Private signage, however, demonstrates greater linguistic flexibility, incorporating Cantonese expressions, traditional Chinese characters, and creative bilingual adaptations. By highlighting the negotiation between top-down language standardization and bottom-up linguistic agency, this study contributes to broader discussions on language policy, cultural heritage preservation, and multilingual accessibility in historical districts. The findings underscore the need for improved linguistic planning, standardized translation policies, and greater public engagement in signage design to ensure that linguistic landscapes in heritage sites are both culturally authentic and globally navigable.
In our contemporary world lots of social networks have emerged having an immense influence on the language development. Facebook, Twitter, Tiktok and others not only change the way we communicate, but also fundamentally transformed the nature of the language. Generally, the nature of the language development has always been a slow process. What once took decades or centuries now happens in months or years. New words, phrases, and expressions can spread globally within hours through viral posts, memes, and trending hashtags But social media significantly accelerates this process. The speed of spreading new words and terms is getting higher and higher. The dictionaries, academic institutions, and formal media—no longer control the pace of linguistic innovation. Linguistic innovations like Memes, Hashtags and viral phrases create new forms of language that don’t conform to traditional linguistic norms. Social media has made the trend toward informal language use much quicker and inevitable. Casual spelling, grammar, and vocabulary that were once characteristics of personal correspondence are now common in public discourse. Social media has indeed dismantled traditional boundaries between formal and informal language registers.As language evolvement is a gradual process it requires our great attention to follow the changes that can determine the true nature of the language.
Online media has become the primary source of information for modern society due to its quick access and ease of news presentation. However, this development also poses challenges, particularly in maintaining writing quality, such as diction accuracy. Appropriate diction plays a crucial role in delivering clear messages, avoiding misunderstandings, and preserving media credibility. This study aims to analyze diction errors in the daily online news jateng.akurat.co edition of November 12, 2024, including non-standard word usage and word mismatches. This research employs a qualitative descriptive method to identify the types of errors, their causes, and their impacts on readers. Data were collected using documentation techniques by observing, recording, and analyzing the news texts published in the selected edition. The analysis process involved comparing diction usage with applicable linguistic norms and relevant contexts. The findings reveal several common errors, such as the use of terms that do not align with formal language norms, inappropriate word choices, and the improper adaptation of foreign words. Diction errors can create reader confusion, diminish media credibility, and affect perceptions of the information conveyed. This study highlights the importance of accurate diction selection in journalism, especially for regionally based media, which must consider cultural contexts and local values. These findings are expected to serve as a guideline for enhancing linguistic accuracy in digital journalism, maintaining media credibility, and supporting journalism's role as a pillar of a healthy democracy.
In 2025, we held the fourth iteration of the DIS-RPT Shared Task (Discourse Relation Parsing and Treebanking) dedicated to discourse parsing across formalisms.Following the success of the 2019, 2021, and 2023 tasks on Elementary Discourse Unit Segmentation, Connective Detection, and Relation Classification, this iteration added 13 new datasets, including three new languages (Czech, Polish, Nigerian Pidgin) and two new frameworks: the ISO framework and Enhanced Rhetorical Structure Theory, in addition to the previously included frameworks: RST, SDRT, DEP, and PDTB.In this paper, we review the data included in DISRPT 2025, which covers 39 datasets across 16 languages, survey and compare submitted systems, and report on system performance on each task for both treebanked and plain-tokenized versions of the data.The best systems obtain a mean accuracy of 71.19% for relation classification, a mean F 1 of 91.57(Treebanked Track) and 87.38 (Plain Track) for segmentation, and a mean F 1 of 81.53 (Treebanked Track) and 79.92 (Plain Track) for connective detection.The data and trained models of several participants can be found at https://huggingface. co/multilingual-discourse-hub.
Using masculine forms for mixed-gender groups or individuals of unknown gender leads people to think of men. In grammatically gendered languages, using feminine and paired forms (gender-inclusive language, GIL) is a common and effective strategy to increase the visibility of women. Although GIL benefits women as a group, its adoption may encounter resistance, especially among employed women. They may refrain from using feminine forms to refer to themselves due to apprehension about potential backlash for deviating from professional and linguistic norms. Additionally, they may hesitate, so as not to evoke societal stereotypes that associate femininity with lower competence and status in professional settings. In two representative samples of Polish self-identified women, we examined the prevalence of GIL forms in professional self-reference. In both studies, approximately half of the participants used feminine forms. The tendency to use GIL was less pronounced among employed participants, with only a third of currently employed women using feminine job titles. Gender identification moderates this effect, with working women who strongly identify with other women being more likely to use feminine forms. These findings shed light on the potential social and professional factors influencing the adoption of GIL by documenting who uses these forms in what contexts. By identifying potential barriers to broader adoption this study underscores the need to address these challenges with professional and policy-oriented interventions.
Word order difference between source and target languages is a major obstacle to cross-lingual transfer, especially in the dependency parsing task. Current works are mostly based on order-agnostic models or word reordering to mitigate this problem. However, such methods either do not leverage grammatical information naturally contained in word order or are computationally expensive as the permutation space grows exponentially with the sentence length. Moreover, the reordered source sentence with an unnatural word order may be a form of noising that harms the model learning. To this end, we propose an Implicit Word Reordering framework with Knowledge Distillation (IWR-KD). This framework is inspired by that deep networks are good at learning feature linearization corresponding to meaningful data transformation, e.g. word reordering. To realize this idea, we introduce a knowledge distillation framework composed of a word-reordering teacher model and a dependency parsing student model. We verify our proposed method on Universal Dependency Treebanks across 31 different languages and show it outperforms a series of competitors, together with experimental analysis to illustrate how our method works towards training a robust parser.
This paper presents a real-time American Sign Language (ASL) recognition system utilizing a hybrid deep learning architecture combining 3D Convolutional Neural Networks (3D CNN) with Long Short-Term Memory (LSTM) networks. The system processes webcam video streams to recognize word-level ASL signs, addressing communication barriers for over 70 million deaf and hard-of-hearing individuals worldwide. Our architecture leverages 3D convolutions to capture spatial-temporal features from video frames, followed by LSTM layers that model sequential dependencies inherent in sign language gestures. Trained on the WLASL dataset (2,000 common words), ASL-LEX lexical database (~2,700 signs), and a curated set of 100 expert-annotated ASL signs, the system achieves F1-scores ranging from 0.71 to 0.99 across sign classes. The model is deployed on AWS infrastructure with edge deployment capability on OAK-D cameras for real-time inference. We discuss the architecture design, training methodology, evaluation metrics, and deployment considerations for practical accessibility applications.
We present a systematic empirical study of small language models under strict compute constraints, analyzing how architectural choices and training budget interact to determine performance. Starting from a linear next-token predictor, we progressively introduce nonlinearities, self-attention, and multi-layer transformer architectures, evaluating each on character-level modeling of Tiny Shakespeare and word-level modeling of Penn Treebank (PTB) and WikiText-2. We compare models using test negative log-likelihood (NLL), parameter count, and approximate training FLOPs to characterize accuracy-efficiency trade-offs. Our results show that attention-based models dominate MLPs in per-FLOP efficiency even at small scale, while increasing depth or context without sufficient optimization can degrade performance. We further examine rotary positional embeddings (RoPE), finding that architectural techniques successful in large language models do not necessarily transfer to small-model regimes.
This paper presents an enhanced Bidirectional Long Short-Term Memory (Bi-LSTM) network with attention mechanism for sentiment analysis in e-commerce user-generated content, integrating approaches from cognitive science, linguistics, and deep learning. Drawing on cognitive attention theories and linguistic frameworks for sentiment expression, our model implements a cognitively-inspired multi-head attention mechanism combined with domain-specific word embeddings. Experiments conducted on a combined dataset of 75,000 reviews from IMDb and Stanford Sentiment Treebank, supplemented with 10,000 Amazon product reviews, demonstrate the model’s effectiveness. The enhanced Bi-LSTM achieves 91.8% accuracy, showing improvements of 4.2% over vanilla Bi-LSTM and 2.7% over BERT-base models. Computational efficiency analysis reveals a 35% reduction in inference time while maintaining stable memory utilization at 2.8GB during peak operation. This interdisciplinary approach effectively bridges theoretical insights from cognitive science with practical applications in natural language processing, providing robust solutions for real-world sentiment analysis challenges.
This article is dedicated to the study of the use of derived nouns of the type Gegröle – Grölerei with the suffixes -ei/-erei and the prefix ge- in German online newspaper texts. The choice of material for the study is determined by the exceptional role of online media in modern society and their influence on public opinion, cultural, and linguistic norms. The author emphasizes the use of shorter text formats and concise headlines in online news portals to attract readers' attention, which can be achieved, among other ways, with the aforementioned derivatives. Such word formations arise from various derivational patterns and often express nuances in meaning and stylistic coloring. The main focus of the article is on the functional-semantic analysis of these parallel derivatives. For this purpose, their etymology, features of formation, and meanings based on lexicographic sources are first examined. Then, the denotative and connotative meanings of both forms in newspaper contexts are considered, and their role in expressing a sense of contempt is determined. The author analyzes how derivatives with Ge- and the suffix -erei create similar but slightly different meanings that can provoke differentiated perceptions in readers. For example, Gegröle is often used to denote an indefinite, annoying sound, while Grölerei adds a behavioral and more pejorative component. The analysis attempts, based on Harden's hypotheses, to establish authors' preferences in choosing one form or another capable of eliciting various emotional reactions from readers.
This study explores the impact of social media on the evolution of discourse structures and the linguistic identity of Generation Z. Using a qualitative, literature review methodology, the research analyzes existing studies to understand how English is adapted and reshaped within digital communication practices. Social media platforms, such as Twitter, Instagram, and TikTok, have led to a shift from traditional linguistic norms towards more informal, brief, and multimodal forms of communication. The review of scholarly works reveals that the structure of language has changed, with an increased use of abbreviations, emojis, and other non-textual elements, reflecting a more fluid and adaptive form of discourse. Additionally, the study highlights how social media serves as a platform for linguistic identity formation, where Generation Z constructs and negotiates their self-representation through language, often incorporating hybrid linguistic forms and code-switching. These changes, while facilitating new ways of expressing identity and fostering digital affiliation, may also pose challenges regarding the preservation of formal language skills in academic and professional contexts. The findings underscore the need for future research to explore the long-term effects of these shifts on language proficiency and their implications for education systems. This paper contributes to the growing understanding of the intersection between digital communication and linguistic evolution in the age of social media.
تقدم هذه الدراسة تحليلًا لغويًا قائمًا على المدونات لنص ديني كنص فعال، مثل النصوص السياسية أو الدينية. الغرض من استخدام التحليل الكمي القائم على المدونات هو تحديد الوزن الحقيقي للنص من حيث الميزات الجمالية. في هذه الدراسة، اخترت ترجمتين للسورة 56 من القرآن الكريم. هذه الترجمات هي من قبل غالي وعبد الحليم تركز اللغويات الحاسوبية والأسلوبية على العلاقة بين الشكل والمعنى. على الرغم من أن الأسلوبية تضع تركيزًا قويًا على الانحراف اللغوي، إلا أن التحليل الأسلوبي يأخذ أيضًا في الاعتبار القواعد اللغوية This study provides a corpus-based linguistic analysis of a religious text as an effective text, like political or religious ones.The purpose of using a corpus-based quantitative analysis is to determine the real weight of the text in terms of aesthetic features. In this study, I chose two translations of chapter 56 from the Noble Quran. These translations are by Ghali and Abdel Haleem. Corpus linguistics and stylistics focus on the relationship between form and meaning Although stylistics places a strong emphasis on linguistic deviation, stylistic analysis also takes into account linguistic norms.
It has often been assumed that receivers interpret the emotional tone of a text-based message differently and often more negatively than intended by the sender. It is unclear, however, whether this is true in everyday online conversations between non-strangers. We therefore tested this by comparing sender and receiver ratings of text messages exchanged in an informal context (Study 1) and emails exchanged in a work or educational context (Study 2). In both studies, we asked participants ( N Study1 = 347; N Study2 = 361) to rate the valence of a message they received and asked the sender of that message ( N Study1 = 171; N Study2 = 61) to report the intended valence of the message. We tested six possible moderators: (1) the length of the message, (2) the use of emoji, (3) the gender of the receiver, (4) the age of the receiver, (5) the social closeness between sender and receiver, and (6) the degree of neuroticism of the receiver. In both studies, we find no indication for misunderstanding as receivers' and senders' valence ratings align very well. We also find no evidence for moderation effects. This shows that, in the context of everyday text messages and emails, people are able to correctly interpret the emotional valence of a text-based message. This finding challenges the popular assumption of prevalent online misunderstanding and provides empirical support for the idea that people can and do successfully adapt their communication style to accurately convey the emotional tone in text-based messages.
This chapter examines the transformative yet contested role of digital language technologies in multilingual educational contexts. It explores how tools such as AI-powered language platforms, machine translation engines, and digital storytelling applications shape power dynamics, reinforce linguistic hierarchies, and affect the rights and representation of minoritized languages. Through case studies and comparative analysis of widely used digital platforms, the chapter reveals how algorithmic design can marginalize regional dialects and elevate dominant linguistic norms, thereby contributing to digital linguistic imperialism. Despite these serious challenges, the chapter highlights the empowering potential of language technologies to promote adaptive learning, cultural exchange, and intercultural communication, when they are applied through thoughtful, critical, and equitable pedagogical approaches. It further emphasizes the need to examine how learners engage with these technologies in informal contexts and how their experiences reshape language attitudes and agency. By proposing an integrative pedagogical framework grounded in critical language awareness and inclusive digital literacy, the chapter provides practical insights for educators seeking to navigate the ethical and pedagogical complexities of digital language education. Ultimately, this work contributes to a growing discourse on digital justice and linguistic human rights, advocating for multilingual digital futures that center diversity, inclusion, and empowerment within increasingly technologized learning environments.
Designers aim to create designs that resonate with users, the consumers. However, there is often a gap between the sensory and emotional “words” expressed by users and the “words” understood by designers. To address this issue, we propose a method called “Language-based engineering” to reflect the meanings users seek in designs accurately. Using the packaging design of golf gloves as an example, we conducted a survey on six sensory words (I want to pick up, novel, visible, cool, cute, luxurious) to gather user feedback. We analyzed the relationship between user ratings and design elements using Quantification Theory Type I and created a database of these relationships. Furthermore, we converted the sensory word ratings into aspiration levels using the satisficing trade-off method and proposed a method to determine the combination of design elements that satisfy these desired levels as a multi-objective optimization problem. We conducted trade-off analyses on all 30 patterns of the six sensory words to identify the trade-off relationships based on users’ impressions of the packaging design. Additionally, we presented design examples and guidelines that align with users’ sensibilities by combining design elements that meet the aspiration levels.
In this paper we present a sample treebank for Old English based on the UD Cairo sentences, collected and annotated as part of a classroom curriculum in Historical Linguistics. To collect the data, a sample of 20 sentences illustrating a range of syntactic constructions in the world's languages, we employ a combination of LLM prompting and searches in authentic Old English data. For annotation we assigned sentences to multiple students with limited prior exposure to UD, whose annotations we compare and adjudicate. Our results suggest that while current LLM outputs in Old English do not reflect authentic syntax, this can be mitigated by post-editing, and that although beginner annotators do not possess enough background to complete the task perfectly, taken together they can produce good results and learn from the experience. We also conduct preliminary parsing experiments using Modern English training data, and find that although performance on Old English is poor, parsing on annotated features (lemma, hyperlemma, gloss) leads to improved performance.
In the context of globalization and expanding intercultural contacts, the formation of a secondary language personality becomes particularly significant. For Chinese philology students learning Russian as a foreign language (RFL), mastering not only linguistic norms but also culture-specific elements such as color-related proverbs presents a crucial challenge. These linguistic units reflect the national worldview and often lack direct equivalents in Chinese linguoculture, creating additional learning difficulties. The article aims to identify and substantiate methodological principles for teaching Russian color proverbs to Chinese philology students, facilitating the development of their secondary language personality and intercultural competence. The study analyzes core problems related to intercultural differences in color symbolism perception and the challenges Chinese learners face in acquiring proverbs. Special attention is given to the need to account for cultural connotations in teaching; principles of non-translation and communicative approaches; the role of situational-thematic organization of learning materials. The authors conclude that effective teaching of Russian color proverbs requires an integrated approach combining linguistic and cultural aspects. The application of proposed methodological principles (non-translation, communicative approach, intercultural interaction, etc.) not only enhances students’ language competence but also facilitates their immersion into the Russian worldview which is a fundamental condition for forming a secondary language personality.
The emergence of digital technology has brought profound changes in the ways language is produced, shared, and interpreted. With the widespread use of social media platforms, instant messaging applications, blogs, and online discussion forums, communication has become faster, more interactive, and less bound by traditional linguistic norms. These digital environments encourage new linguistic practices that influence vocabulary, grammar, spelling, and discourse patterns. As a result, language in the digital age is constantly evolving and reshaping conventional forms of communication. This study investigates language change in the digital age by conducting a linguistic analysis of online communication. The research focuses on identifying key linguistic features commonly used in digital contexts, such as abbreviations, acronyms, emoji’s, creative spellings, grammatical simplification, and code-switching. A qualitative research approach is employed to analyze selected online texts in order to understand how technological and social factors contribute to linguistic variation and innovation. The findings of the study reveal that digital communication promotes linguistic creativity and flexibility rather than linguistic decay. Online language reflects users’ identities, social relationships, and communicative needs within digital spaces. The study concludes that language change driven by digital communication is a natural process of linguistic evolution and highlights the importance of recognizing online discourse as a significant and legitimate area of linguistic research.
International agreements are legally binding documents, which are characterized by highly structured and formalized language designed to ensure legal precision and enforceability. To translate them in line with legal and linguistic norms underlying legal translation, translators should be familiar with the characteristics of legal texts as well as be able to analyze grammatical structures and patterns in legal texts. This chapter aims to examine the translation of international agreements from English to Slovenian. It focuses on syntactic structures of simple and complex sentences at sentence and clause levels within the framework of systemic functional linguistics. The chapter thus aims to explore how these structures are maintained or adapted during translation. The focus is on identifying and assessing syntax-related translation shifts, including their frequency and the implications for the legal integrity of the target text. The chapter demonstrates the degree to which the syntactic characteristics of the source texts in English are preserved in Slovenian translations, given the strict requirements of legal language and the rules of the Slovenian language linguistic system. The key finding of the research presented in this chapter is that syntax-related translation shifts occur only when necessary to conform to the Slovenian language grammar system. The nature of the established shifts points to the translators’ adherence to the rule regarding the maintenance of the highest possible degrees of accuracy when translating normative texts while observing the grammatical characteristics of legal language.
The most common usage of the Greek particle οὖν/oûn is to create coherence within a large portion of discourse and convey inferential meaning. This chapter addresses its usage in documentary papyri of the Roman and Byzantine periods, aiming at providing a preliminary general exploration of the contextual usage of this particle and approaching it through pragmatic and discourse analysis approaches. It focuses on the linguistic elements that surround this particle and identifies different patterns that frequently occur within the corpus of documentary papyri. By using the PapyGreek treebank corpus and Trismegistos, statistics for these patterns are added in order to draw observations on the co-occurrence of the particle and other linguistic elements that precede it. By means of examples, the chapter discusses a selection of these patterns, their context of use, and their linguistic function. It analyses the position of some οὖν/oûn utterances, which for instance typify a request, a warning, or a piece of advice, looking at their sequential placement within the discourse. It therefore addresses the question of how inferences are constructed and conveyed. Finally, focusing on some cases of negative οὖν/oûn utterances, the chapter discusses some aspects of their meaning and functions by using a cognitive linguistic approach.