Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The article deals with the language description transfer from the encyclopaedic edition «Languages of the World» to the «Languages of the World» database web version. The feature base structure and dependencies within different sections are analysed, along with the possible ways of terminology unification. The article may be of interest to specialists in linguistic typology, data formalisation and development of linguistic databases.
Large language models (LLMs) have materially changed natural language processing (NLP).While LLMs have shifted focus from traditional semantic-based resources, structured linguistic databases such as WordNet remain essential for precise knowledge retrieval, decision making and aiding LLM development.WordNet organizes concepts through synonym sets (synsets) and semantic links but suffers from inconsistencies, including redundant or erroneous relations.This paper investigates an approach using LLMs to aid the refinement of structured language resources, specifically WordNet, by an automation for multiple hypernymy resolution, leveraging the LLMs semantic knowledge to produce tools for aiding and evaluating manual resource improvement.
Theory of Mind (ToM)—the ability to infer others’ mental states—is fundamental to social cognition. Social categorization, the grouping of individuals into in-group or out-group categories, shapes these inferences. These processes co-occur during facial perception, with recent research suggesting both shared neurocognitive mechanisms and modulation of ToM by social category cues. However, existing tools for studying the impact of racial diversity on social cognition are limited: some databases prioritize racial representation but restrict stimuli to the six basic emotions, while others broaden mental state diversity but lack diversity in social appearance.Here we introduce the McGill Diverse Face Database, a validated set of 1,286 images of 14 actors from socially perceived racial categories portraying 92 complex mental states. Validation included three experiments: (1) a four-alternative forced-choice task assessing recognition accuracy, (2) a “point-and-click” task rating valence and arousal in a two-dimensional affective space, and (3) a trait-rating task evaluating perceived actor characteristics. Participants accurately identified mental states across categories (77 % of stimuli). Mean valence–arousal ratings revealed a non-linear one-dimensional manifold structure that correlated with behavioral measures. An interactive online visualization allows users to explore this “Theory of Mind manifold” (https://hctor99.github.io/TheoryofMindManifold/). By integrating social-category diversity with complex emotional expression, this database provides a new resource for studying how socially perceived group membership shapes the perception and inference of mental states.
Understanding and expressing emotions is subjective, leading to varying levels of ambiguity in how emotions are perceived. Humans naturally handle this ambiguity, and it's crucial for machines to do the same for natural human-machine interaction. Efforts are being made to develop emotion recognition systems that can handle ambiguity by modelling emotions as distributions of arousal/valence labels. However, existing approaches often assume that an underlying distribution can be inferred from ground truth ratings. Yet, the underlying distribution is never observed and the inferred distribution from the ratings can accurately represent the true distribution only when enough raters are involved. Our study investigates how the number of raters impacts the uncertainty in inferring distributions from ground truth ratings in the context of modelling emotion ambiguity. We then explore how many raters may be sufficient to effectively approximate the ambiguity present in the emotion annotations. Using the Belief Mismatch Coefficient (BMC) for analysis, we compare distributions to different numbers of ground truth ratings leading to quantitative and interpretable inferences. Experimental analysis was conducted using both simulated ratings and the real emotion (both arousal and valence) ratings collected from MSP-Conversation dataset. Our results suggest that using 5 raters may be sufficient to approximate a stable distribution representing the underlying emotional ambiguity when time continuous emotion ratings are captured and annotated from speech. Our study provides insights and practical advice for collecting emotion labels. Additionally, the analyses methods presented in this paper are expected to be relevant in any field of study dealing with labelling of subjective quantities that can vary from person to person.
[... waveform poisoning initiated... ] This archive is the missing addendum to the manuscript "The Atlas Code: A Hitchhiker's Guide to No Man's Sky." It contains the complete Glyph Rosetta Dataset — a linguistic database of 4,820 words and their associated glyphs across the five languages of No Man's Sky: Atlas, Gek, Korvax, Vy'keen, and Autophage. This version (v1.1) has been rebuilt from the ground up and version locked to No Man's Sky Update 6.33 (April 15, 2026). It supersedes v1.0, which should no longer be cited. The rebuild accounts for new words added by Hello Games in Updates 5.6 and 6.33, corrections to internal string errors, and clerical errors identified during three independent manual verification passes of the complete dataset. The completed rebuild reduced the total dataset size by 150mb. The archive includes all in-game menu screenshots catalogued with a three-part inventory system for rapid cross-referencing, the complete dataset in both CSV and Excel formats, a First Pass Analysis worksheet documenting shared words, bigram assignments, and cross-language vocabulary overlaps, and a full Table of Contents for the original manuscript. This data is provided for the sole purpose of allowing the community to verify the findings of the original manuscript.
In this paper, we present a novel lemmatization method based on a sequence-to-sequence neural network architecture and morphosyntactic context representation. In the proposed method, our context-sensitive lemmatizer generates the lemma one character at a time based on the surface form characters and its morphosyntactic features obtained from a morphological tagger. We argue that a sliding window context representation suffers from sparseness, while in majority of cases the morphosyntactic features of a word bring enough information to resolve lemma ambiguities while keeping the context representation dense and more practical for machine learning systems. Additionally, we study two different data augmentation methods utilizing autoencoder training and morphological transducers especially beneficial for low-resource languages. We evaluate our lemmatizer on 52 different languages and 76 different treebanks, showing that our system outperforms all latest baseline systems. Compared to the best overall baseline, UDPipe Future, our system outperforms it on 62 out of 76 treebanks reducing errors on average by 19% relative. The lemmatizer together with all trained models is made available as a part of the Turku-neural-parsing-pipeline under the Apache 2.0 license.
In this 2-session functional MRI study, we investigated the neurocognitive mechanisms of counterconditioning (CC). We observed ventromedial prefrontal cortex (vmPFC) activation during regular extinction, but vmPFC deactivation and nucleus accumbens activation during CC (MRI). Physiological threat responses returned in the regular extinction but not the CC group (Pupil). Furthermore, threat-associated memories from the learning and CC phase were enhanced (Behavior). This collection contains the raw, pseudonomized data of all tasks. Processed data, the SPSS mask to replicate all but the fMRI analyses (Analyses.sav), and scripts to replicate the publication figures are available at https://osf.io/mvsq6. For all tasks, fMRI, skin conductance (SCR) and pupil dilation (Pupil) data are available. For the item memory test, behavioral data (Behavior) are available. In addition, valence and arousal rating data using self-assessment manikin scales are available.
Our study examines the integration of large language models (LLMs) on psycholinguistic research by prompting several LLMs to rate figurative expressions on familiarity, aptness, concreteness, metaphoricity, and constituency, with two manipulations: (1) stimuli presentation (in context versus in isolation) to examine whether LLMs benefit from the presence of context when rating familiarity—but not aptness—as observed in a recent study (Pissani and de Almeida, in press.) and (2) type of instructions (graded ratings versus categories) to examine whether LLMs perform better when assigning categories rather than numerical ratings, as observed in a recent study (Bavaresco et al., 2025). In addition, we will substitute LLM-produced ratings into existing studies of metaphor comprehension to determine whether they can predict the same effects observed with human ratings. Our study serves three purposes. First, we aim to examine whether LLM outputs replicate humans in rating tasks involving figurative language and whether they can serve for data augmentation when human ratings are unavailable. Second, we aim to understand whether LLMs process these ratings in the same way as humans (e.g., familiarity ratings may reflect both frequency of use and ease of understanding, while aptness ratings may reflect the quality of the metaphorical mapping, regardless of context). Third, we aim to provide best practices for obtaining reliable ratings from LLMs, whether using graded scales or categorical outcomes.
This paper introduces new models designed to improve the morpho-syntactic parsing of the five largest Latin treebanks in the Universal Dependencies (UD) framework. First, using two state-of-the-art parsers, Trankit and Stanza, along with our custom UD tagger, we train new models on the five treebanks both individually and by combining them into novel merged datasets. We also test the models on the CIRCSE test set. In an additional experiment, we evaluate whether this set can be accurately tagged using the novel LASLA corpus (https://github.com/CIRCSE/LASLA). Second, we aim to improve the results by combining the predictions of different models through an atomic morphological feature voting system. The results of our two main experiments demonstrate significant improvements, particularly for the smaller treebanks, with LAS scores increasing by 16.10 and 11.85%-points for UDante and Perseus, respectively (Gamba and Zeman, 2023a). Additionally, the voting system for morphological features (FEATS) brings improvements, especially for the smaller Latin treebanks: Perseus 3.15% and CIRCSE 2.47%-points. Tagging the CIRCSE set with our custom model using the LASLA model improves POS 6.71 and FEATS 11.04%-points respectively, compared to our best-performing UD PROIEL model. Our results show that larger datasets and ensemble predictions can significantly improve performance.
Sanskrit is traditionally described as a free-word-order language, and the indigenous grammatical tradition explains what constrains word combination through three notions: ākāṅkṣā (syntactic expectancy), yogyatā (semantic compatibility), and sannidhi (proximity or contiguity). This study asks whether sannidhi, read as a locality constraint, is empirically supported, and how it relates to the cross-linguistic principle of dependency-length minimization (DLM). Using the Universal Dependencies Treebank of Vedic Sanskrit (27,182 sentences; 206,440 tokens), I compared observed dependency lengths against a random-projective baseline and a minimal-arrangement heuristic, partitioned arcs by grammatical relation, quantified non-projectivity, and traced variation across the corpus's chronological layers. At the whole-sentence level, Vedic showed no dependency-length minimization beyond the projectivity constraint: observed mean length (1.996) was statistically indistinguishable from the random-projective baseline (2.002) and far above the minimal arrangement (1.512). This near-parity, however, masked a systematic relation-specific split. Core verb-argument relations—the kāraka-type expectancy relations—were placed reliably closer than their own random baseline (mean deviation −0.36 tokens, 95% CI [−0.38, −0.34]), whereas coordinate, appositional, and modifier relations were placed at or beyond chance distance (+0.16 tokens, 95% CI [+0.13, +0.18]). Non-projectivity was common (19.4% of sentences) but overwhelmingly mild and well-nested, and it declined sharply from the Ṛgvedic layer (35.9%) to the Sūtra layer (11.0%). An independent classical treebank reproduced both the aggregate parity and the non-projectivity rate. I argue that sannidhi is best understood not as global length optimization but as a selective, expectancy-scoped locality operating over ākāṅkṣā-linked pairs, and that Vedic word order becomes measurably more projective over time.
This randomized controlled study at University of Graz examines the impact of regular practice of an individual self-regulation training on psychophysiological well-being. Participants are randomly assigned to either intervention or control group. Individual self-regulation training is based on a self-regulation method used in NeuroDeescalation®. It combines elements of body movement or touch, breathing and self- encouragement. Participants in the intervention group get videos which support the acquisition of the individual self-regulation training, which should be practiced three times a day over a period of two weeks. In both groups HRV independent of metabolic demands (lmdHRV) is assessed the successive three days before and after intervention period. Furthermore, psychological variables are measured using items of the Positive and Negative Affect Schedule (PANAS), and the Mindful Attention Awareness Scale – State (MAAS – State). We expect that regular practice of an individual self-regulation training leads to higher increases in HRV independent of metabolic demands (lmdHRV), as well as higher increases in positive affect ratings, lower increases in negative affect ratings, and higher increases in state mindfulness from pre- to post- intervention compared to the control group. Statistical analyses will be conducted using mixed Analysis of Variance (ANOVA) to account for between-subjects factor (group) and within-subjects factor (measurement period).
This thesis aims to examine the impact of mind wandering on statistical language learning, a form of implicit learning that allows the detection of patterns and regularities from external input. Mind wandering refers to the shift of attention away from an external task toward internal thoughts, often occurring involuntarily. While mind wandering is traditionally associated with impaired performance in attention-demanding tasks such as learning, emerging research suggests that mind wandering may have a beneficial effect on certain instinctual cognitive processes, such as statistical learning and implicit learning (Vékony et al., 2025). Two experiments are conducted to investigate whether mind wandering influences performance and metacognitive confidence in artificial language learning tasks. In Experiment 1, participants listen to a continuous stream of trisyllabic pseudowords and complete a two-alternative forced choice (2AFC) task, providing confidence ratings for each decision. In Experiment 2, participants complete familiarity rating tasks after exposure to both paused and continuous speech streams with either front-vowels or back-vowels. In both experiments, mind wandering is assessed through a self-report questionnaire administered after each exposure. Results indicate no significant relationship between mind wandering and statistical language learning performance. However, no participants reported fully disengaging from the tasks, which likely affected the results. Due to the limited variability in mind wandering, as indicated by the mind wandering questionnaire scores that do not indicate high mind wandering, further research is needed. Future studies should explore whether more substantial or sustained mind wandering might enhance statistical language learning, as reported in emerging research. In addition, the study design could be applied to other cognitive domains.
The article aims to focus on the impact of digital technologies on language, paying particular attention to digital English discourse and its linguistic features in online platforms, reflecting current Internet communication trends.The role of digital platforms (social networks, messengers, forums) in forming new language models, lexical innovations, syntactic simplifications and multimodal communication is considered.The paper examines how IT technologies have transformed communication, language use and reshaped linguistic norms and social engagement.Typical lexical items, stylistic features and pragmatic functions of digital discourse are analysed.The author highlights the contributions of linguists in investigating key features of digital discourse such as informality, multimodality, immediacy and interactivity.A rapid development of new linguistic forms such as emojis, emoticons, abbreviations, shortenings, hashtags, and innovative lexical items is observed, reflecting both the expressiveness and efficiency demands of digital communication.The research emphasises the process of human-machine interactions with the role of chatbots and their ability to simulate empathy and involvement in informal, dynamic conversation.Ongoing transformation of human communication due to digital technologies, particularly the dynamic nature of digital discourse and its multimodal features, suggests the prospects for further research.Linguists must investigate the influence of emerging technologies such as artificial intelligence, augmented reality and virtual reality on language use and linguistic norms.The research material aims to contribute to a deeper understanding of the dynamics of the English language in digital contexts and the notion of digital discourse.
In the article, the issue of rhetoric as a factor in shaping the linguistic culture of Ukrainian society under conditions of globalization transformations is revealed.The theoretical foundations for interpreting language culture as a conscious aspiration for normativity, accuracy, and stylistic appropriateness of linguistic expression in various sociocultural contexts are highlighted.It is clarified that rhetoric occupies an integrative position between linguistics, logic, cultural studies, and pedagogy, ensuring the development of the ability to use linguistic resources persuasively, argumentatively, and ethically.It is emphasized that rhetorical competence serves as the foundation for the development of critical thinking, civic engagement, and linguistic subjectivity of an individual.It is underlined that rhetorical practice facilitates the transition from passive mastery of linguistic norms to their active application as an instrument of social interaction, influence, and communicative responsibility.It is stated that rhetoric ensures the conceptual integrity of an utterance, which relies on the logical consistency of thinking, argumentation, and the organization of the semantic structure of speech.It is outlined that the ability to logically structure, formulate, and justify the speaker's position is essential for shaping collective opinion, developing public argumentation, and preserving cultural heritage.It is specified that rhetorical expressions serve as a means of representing national identity, cultural memory, and historical-semantic values.It is established that rhetoric currently demonstrates an insufficient level of integration in the educational process, which limits the development of skills in argumentation, persuasion, and effective public communication.It is noted that the application of rhetorical methods in education activates the cognitive, cultural, and communicative potential of society.It is outlined that the combination of rhetorical, ethical, and logical-cultural components ensures the formation of high linguistic culture as an indicator of the nation's spiritual maturity.
The article describes intra- and extra-linguistic factors that influence the stability of the onymic space of the Ukrainian language. Intra-linguistic factors are the creation of a certain proper name according to its inherent derivational model; the correspondence of a particular onymic to the formed linguistic norm. Extra-linguistic factors are the level of linguistic, spiritual and political culture of society and its national consciousness. This influence is most pronounced on two classes of onymic vocabulary — anthroponyms and toponyms, as well as on toponymic derivatives — the names of the inhabitants of the corresponding settlement (katoikonyms) and derived adjectives (adjectonyms). It is specially noted of the three-component anthroponymic formula (first name, patronymic, surname) in modern Ukrainian official speech; the elimination of variation in the declension of Ukrainian surnames of the masculine gender with the possessive suffix -iB. It is shown that main means of creating katoikonyms in the Ukrainian language are the suffixes -u-i (plural), -eub, -K-a (singular). Possible functioning of parallel catoikonymic forms with suffixes -a-i (-eab, -K-a) h -aH-H / -hh-h (-aH-HH, -aH-K-a / -hh-hh, -HH-K-a). The paper deals with problem of establishing derivational models of adjectonyms and their normalization. The author considers that it is necessary to carefully study the historical patterns of the creation and use of adjectonyms in the Ukrainian language, as well as to take into account local traditions. The destructive impact on the Ukrainian oikonymic space of numerous unmotivated ideological renamings of the Soviet era and artificial formations with a Russian-language structure is analyzed. The author shows that it is necessary to cleanse the Ukrainian oikonomika and urbanonymika of such names.
Concreteness and valence ratings are frequently controlled for in psycholinguistic experiments. Associated valence and concreteness effects appear to be well-established phenomena within psycholinguistic research. However, recent studies indicated replication difficulties (e.g., Kousta et al., 2009; Yap & Seow, 2014; Brysbaert et al., 2016; Kousta et al., 2011). Considering that some approaches to valence and concreteness effects such as the Dual Coding Theory or the Affective Priming Approach would suggest that additional embodied information could be responsible for the processing advantage associated with concrete and positive lexical items, individual differences could partially explain why these effects could be found in some studies but not others: For instance, the temperament trait of Sensory Processing Sensitivity has been associated with a deeper and more detailed processing of sensory and affective information and could therefore influence how likely the effects could be found in a particular study considering the respective participant sample. Specifically, certain explanations at the lexico-semantic level such as the Dual Coding Theory, the LASS account, and the Affective Priming Approach would suggest stronger valence and concreteness effects in individuals with high Sensory Processing Sensitivity. Other lexico-semantic approaches that explain these effects in a disembodied manner, would predict no differences between sensitivity groups. Finally, the automatic vigilance approach as well as a general tendency to pause and check before engaging in a response would predict reduced concreteness and valence effects for individuals with high SPS. These predictions will be tested in a lexical decision task on four lists of 12 words each that differ in terms of their valence and concreteness (2 x 2) but are matched for number of syllables, number of letters, frequency, and arousal. In addition, distractor items and an equal number of nonwords will be included. After the lexical decision task, participants will complete the SPS measure.
Emotional states are fundamentally embodied, emerging from the dynamic interplay between central neural processing and peripheral physiological adjustments orchestrated by the autonomic nervous system (ANS). While ANS outputs like heart rate variability (HRV) and electrodermal activity (EDA) reflect emotional arousal and valence, understanding the precise temporal coordination between brain activity and these peripheral signals is crucial for elucidating brain-body interactions. This study investigates neural-autonomic phase synchrony during the conscious processing of distinct emotional states (positive, negative, neutral) by quantifying the temporal alignment between cortical and physiological rhythms. We employ a multimodal approach, simultaneously recording high-temporal-resolution electroencephalography (EEG), electrocardiography (ECG) for HRV analysis (specifically Root Mean Square of Successive Differences, RMSSD), EDA, and functional near-infrared spectroscopy (fNIRS) while participants view validated emotional video clips. Our primary analysis quantifies the Phase Locking Value (PLV) between frontal EEG oscillations (Alpha, Beta bands) and continuous signals derived from HRV (reflecting parasympathetic influence) and phasic EDA (reflecting sympathetic influence). EEG channel selection for PLV analysis was informed by task-related hemodynamic activity measured via fNIRS to focus on functionally relevant cortical areas. We hypothesize that PLV, indicating brain-body temporal integration, will be significantly modulated by emotional content compared to neutral conditions. We further expect synchrony strength to correlate with subjective arousal ratings. By examining the phase synchrony between brain signals and ANS-mediated physiological outputs, this research provides novel insights into the dynamic, embodied mechanisms underlying emotional experience. Understanding this temporal binding is critical for models of psychophysiological function and may inform assessments of cognitive load or stress regulation capacity, potentially impacting performance monitoring and optimization in demanding operational environments.
The article deals with the evolution of translation strategies, comparing traditional methods with contemporary machine translation technologies.The analysis encompasses a wide range of approaches, from direct and indirect translation to neural network-based models.The key theme is the examination of differences in decision-making processes during translation and the exploration of the advantages and limitations of each approach.The main challenge is the accuracy of content transmission and the consideration of context in translation.Human translators have the ability to adapt texts, taking into account not only grammatical and lexical norms, but also other factors such as emotional nuances, styles and tones, which is important for maintaining the appropriateness of the translation.At the same time, machine translation, even with the most advanced technologies, may struggle with texts that require a high level of cultural and stylistic context.For example, in the context of literary and fictional texts, machine translation systems have been observed to underperform with regard to capturing the nuances and depth of meaning that are characteristic of human translation.Moreover, the article discusses the potential for combining human and machine translation.It is important to note that for certain types of texts (e.g., technical or scientific ones), machine translation can be highly effective, especially when dealing with standardized terms and expressions.However, for more complex and creative texts, such as literary or marketing materials, human translation remains indispensable.In the future, the use of combined systems that integrate the advantages of both approaches may achieve high accuracy while preserving context in translation.Furthermore, the article focuses on the development of new machine translation methods aimed at improving accuracy and contextual understanding.At the same time, the limitations of existing technologies and the challenges faced by researchers in addressing issues related to context, style, and cultural adaptation of texts are discussed.One of the important topics is the role of cultural context in translation.Human translators are able to consider cultural differences and the specific characteristics of each language, which allows for more accurate and appropriate translations.Machine translation, on the other hand, often encounters difficulties in understanding cultural aspects and is not always able to adequately convey meaning when it has strong cultural ties.The article also emphasizes the importance of integrating cultural aspects into machine translation to ensure more accurate and relevant results.
In the modern Ukrainian communicative space, speech culture and rhetoric appear as interconnected components that determine the quality, effectiveness, and ethicality of public expression.The rhetorical tradition, integrated with the norms of speech culture, is a powerful tool for influencing the audience's consciousness and emotions, contributing to the effective achievement of communicative goals in various spheres of social life.The purpose of this study is to conduct a comprehensive examination of the relationship between speech culture and rhetorical practice, to define their common and distinctive features, to analyze the rhetorical aspects of speech culture, and to outline the problems and prospects for the development of rhetorical culture in contemporary Ukraine.An interdisciplinary approach combining descriptive, comparative, linguostylistic, and discourse-analytical methods was employed.The study established that speech culture forms the foundation of rhetoric, ensuring adherence to linguistic norms, logicality, structure, expressiveness, and communicative appropriateness of utterances.It was revealed that the rhetorical aspects of speech culture -harmonious combination of form and content, adaptation to the addressee and situation, and the ability to persuade through linguistic means -are key factors in communication effectiveness.The spheres of speech culture implementation were outlined: education, political discourse, journalism, blogging, social networks, legal, pedagogical, and business communication.Attention was drawn to the phenomena of simplification, democratization, and hybridization of styles in modern public speech.A number of problems were identified: the spread of surzhyk (mixed Ukrainian-Russian speech), vulgarization of language, communication standardization, and the decline of ethical standards in public discourse.The conclusions emphasize that the effective development of rhetorical culture in Ukraine is possible through the combination of systematic language education, preservation and adaptation of rhetorical tradition, cultivation of critical thinking, and enhancement of communication ethics.Speech culture, as a component of rhetoric, should be regarded not only as linguistic competence but also as a sociocultural tool influencing the democratization of society, national identity, and the quality of public discourse.
The aim of teaching any foreign language is to ensure that learners are able to communicate freely in that language.The position of English as a means of international communication is strengthening year by year.Creating interest and enthusiasm for learning this language from secondary school onwards can provide strong impetus for students' future academic activities as well as for their integration into the information society.The primary goal in foreign language instruction should be the acquisition of communicative competence.H. Douglas Brown emphasizes this issue as follows: "One of the most important aspects of language learning is not merely the acquisition of its grammatical and lexical forms, but the ability to use these forms to fulfill the communicative function of language.If the mastery of vocabulary and structures does not allow the learner to convey thoughts, express ideas and feelings, or comprehend what is expressed in that language, then the knowledge of language acquired does not fulfill its communicative function".[7, p.189] Vocabulary forms the foundation of a language.It occupies a central position because meaning in communication is conveyed precisely through words.To listen and understand, to read and comprehend, and to communicate both orally and in writing in a foreign language, one must possess an adequate vocabulary stock.Vocabulary plays an essential role in the realization of both receptive skills-reading and listening-and productive skills-speaking and writing.For productive speech activities such as speaking and writing, it is necessary to possess lexical-thematic associations, to combine newly learned words with previously acquired lexical units in order to form correct sentences, to select appropriate words from synonymic and antonymic groups, and to use words in accordance with linguistic norms and the conditions of communication.In receptive activities such as listening and reading, learners must be able to associate the sound and written form of a word with its semantic meaning, to distinguish homonymy, synonymy, and antonymy, to differentiate phonetically or orthographically similar words based on meaning, and to determine meaning through context.[6] In addition to these, learners should also acquire socio-cultural knowledge in the field of vocabulary.This includes the ability to communicate based on speech forms existing in the country of the target language.Moreover, learners should know Azərbaycan Milli Elmlər Akademiyası M. Füzuli adına Əlyazmalar İnstitutu,
У статті зазначено, що в сучасному глобалізованому світі, де міжкультурна комунікація набуває особливої актуальності, проблема адекватного перекладу ділових паперів постає як одна з ключових у сфері професійного перекладознавства. Зокрема, особливої уваги заслуговують процеси запозичення та калькування в міжмовному трансфері між українською та англійською мовами, оскільки вони суттєво впливають на якість і точність передачі офіційної, юридичної та комерційної інформації. Доведено, що ділова документація є складним і водночас чутливим жанром текстів, що вимагає від перекладача не лише високого рівня володіння мовами, а й глибокого розуміння правових, культурних і лінгвістичних норм обох країн. У цьому контексті особливої ваги набуває питання про доцільність використання запозичень та кальок як перекладацьких стратегій. Наголошено, що запозичення – це процес включення іншомовних елементів у текст без істотної зміни їхньої форми, який часто використовується для збереження термінологічної точності або коли у мові перекладу відсутній адекватний відповідник. Проте не завжди запозичення є виправданими. Іноді їхнє некритичне використання може призводити до плутанини, втрати змісту або стилістичних відхилень, особливо якщо існує усталений український еквівалент. Калькування, у свою чергу, передбачає буквальний або частковий переклад іншомовного вислову з збереженням структури або окремих його компонентів. У діловому перекладі кальки можуть виникати як результат прагнення зберегти формальну відповідність, однак часто вони порушують норми цільової мови. Крім того, перекладач повинен брати до уваги культурно-національні особливості: англомовний юридичний дискурс відзначається своїми лексико-синтаксичними моделями, які часто не мають прямого відповідника в українській мові, і навпаки. Саме тому іноді запозичення або калькування – це не просто перекладацький вибір, а необхідність, продиктована функціональною еквівалентністю. У підсумку вказано, що раціональне використання цих стратегій уможливлює створення точного, зрозумілого й правомірного перекладу, який відповідатиме як формальним вимогам, так і очікуванням реципієнта. The article notes that in today’s globalized world, where intercultural communication is becoming particularly relevant, the problem of adequate translation of business documents is one of the key issues in the fild of professional translation studies. In particular, the processes of borrowing and calquing in interlingual translation between Ukrainian and English deserve special attention, as they signifiantly affct the quality and accuracy of offial, legal and commercial information. It is proved that business documentation is a complex and at the same time sensitive genre of texts that requires not only a high level of language profiiency but also a deep understanding of the legal, cultural and linguistic norms of both countries. In this context, the question of the appropriateness of using borrowings and calques as translation strategies is of particular importance. It is emphasized that borrowing is the process of including foreign language elements into a text without signifiantly changing their form, which is often used to preserve terminological accuracy or when there is no adequate equivalent in the target language. However, borrowings are not always justifid. Sometimes their uncritical use can lead to confusion, loss of meaning, or stylistic deviations, especially if there is an established Ukrainian equivalent. Calculation, in turn, implies a literal or partial translation of a foreign language expression while preserving the structure or its individual components. In business translation, calques may arise as a result of the desire to maintain formal conformity, but they often violate the norms of the target language. In addition, the translator must take into account cultural and national peculiarities: Englishlanguage legal discourse is characterized by its lexical and syntactic models, which often have no direct equivalent in Ukrainian, and vice versa. That is why sometimes borrowing or calquing is not just a translation choice, but a necessity dictated by functional equivalence. The author concludes that the rational use of these strategies makes it possible to create an accurate, understandable and legitimate translation that meets both formal requirements and the expectations of the recipient.
Walking is known to be good for mental and physical health. However, its effects on neural activity remain unclear. To investigate the neural effects of walking, we utilized whole-brain fMRI-based predictive models (“emotion decoders”) for valence and arousal, trained with multivariate pattern analysis (MVPA). Decoder development comprised a discovery phase (model training and internal validation) and a validation phase (external cohort and confirmatory tests). Data collection for the discovery phase is complete; this preregistration specifies the validation-phase acquisition and analyses. In 2024, we completed the fMRI data collection for 32 healthy participants, forming the discovery cohort. In the discovery-phase experiment, participants completed two tasks: an emotional rating task and a one-back task. In the rating task, fMRI data were collected during picture viewing and the emotional rating process. The picture-viewing data period was used to train each valence and arousal neural decoder. Specifically, participants underwent an fMRI paradigm in which they were presented with four types of emotional pictures (pleasant-high arousal; pleasant-calm; unpleasant-high arousal; unpleasant-calm). Following each picture, participants were asked to rate the level of each valence and arousal on a continuous scale of 1 to 100 (from unpleasant [-50] to pleasant [50], and from calm [-50] to arousal [50]). For the brain data analysis, we performed pre-processing of the fMRI data using the SPM default pipeline. Specifically, functional images were realigned to the mean image of the series, slice-time corrected, motion corrected, co-registered to the structural image, normalized to MNI space, and spatially smoothed with a 6-mm FWHM Gaussian kernel. In addition, potential outlier scans were identified from the resulting subject-motion estimates and from BOLD signal indicators using default thresholds in the CONN toolbox preprocessing pipeline (5 standard deviations above the mean in the global BOLD signal change, or framewise displacement values above 0.9 mm). Then, the single-trial first-level fMRI analysis was performed to obtain beta images for each picture rating per participant. We used GLMsingle (Prince et al., 2022) to estimate single-trial beta maps per picture at the individual-subject level. After that, we applied whole-brain MVPA to obtain patterns that predict participants’ valence and arousal ratings from single-trial beta maps for each picture. The brain mask was restricted to the gray matter. For the decoding method, we employed the LASSO-PCR algorithm (Wager et al., 2013), based on individual beta maps for each picture, as features to predict participants’ emotional experience. We used a 10-fold participant cross-validation procedure to evaluate the decoder's performance. In the discovery cohort (N=32), the neural valence and arousal decoders predicted participants’ ratings with mean within-subject trial-wise Pearson correlations exceeding r = 0.45 for valence and r = 0.35 for arousal. Next, participants performed a one-back task on emotional pictures in the fMRI scanner after either walking or reading. We applied the valence and arousal decoders to fMRI data time-locked to picture viewing to obtain decoder-derived neural emotion scores. We then compared these neural emotion scores between the walking and the reading control conditions in a within-subject design. In the walk condition, participants walked on a treadmill at a self-selected comfortable pace. Whereas in the reading control condition, participants read a book while seated. After 20 minutes of walking or reading, participants entered the fMRI scanner. The order of conditions was counterbalanced across participants. During the one-back task, participants viewed a picture at a time and pressed a button when they identified a repetition of an image presented in a previous trial. The procedure ensures that their attention is sustained. To demonstrate the neural effects of walking, we fitted the neural emotion decoder on the fMRI data collected during the picture-viewing period. We obtained and compared the neural emotion score for each walking and reading condition. As a result, we found that decoder-derived neural valence scores for pleasant-high arousal pictures were higher after walking than after reading. A crucial next validation phase for pre-registration will (i) externally evaluate decoder performance and (ii) replicate and extend the findings that walking increases decoder-derived neural valence score relative to reading for pleasant-high arousal pictures. We will acquire fMRI and behavioral data from an independent validation cohort (planned N=30) under a similar experimental paradigm. The validation phase includes a one-back task and a picture-rating task identical to those in the discovery phase. In this phase, the one-back task will be performed initially. The rating task will be used to assess decoder generalization both after walking and after reading conditions (within-subject, counterbalanced). In the discovery cohort, we did not investigate whether subjective evaluations differed between walking and reading conditions, so in the validation cohort, we plan to have participants evaluate under both conditions. Furthermore, we will investigate whether the decoder's predictions hold up even when the same emotional images are shown multiple times. All preprocessing and first-level modeling will follow the analysis pipeline in the discovery phase, yielding single-trial beta maps. The valence and arousal decoders trained by the discovery dataset will be applied to the validation data. We expect that, if the neural emotion decoders in the discovery phase can be generalized to out-of-sample participants, statistically significant prediction-outcome correlations will be observed in the validation phase for individuals in the cohort. Moreover, decoder-derived neural valence scores will be compared between walking and reading conditions, with a focus on pleasant and high-arousal pictures. We expect higher neural pleasure scores and subjective pleasure ratings after walking than after reading, providing robust evidence for a neural effect of walking.
The present dissertation examines the role of context-specific simulations in influencing the complexity of affective experiences, drawing on a constructivist approach to emotion. To link literatures on mental simulation and emotion, in Chapter 1 a connection is made through the grounded theories of cognition. Chapter 2 describes the development of a novel dataset consisting of context-dependent stimuli (i.e., 1,381 picture-word cues derived from 320 pictures-only stimuli) validated through online experiments (NExp1= 1,934; NExp2 = 403). Hence, an investigation of how contextual information influences the affective experience is illustrated, revealing that context more often enhances affective complexity by widening, rather than narrowing, the variation in the between-subject valence ratings. Chapter 3 employs a set of stimuli selected from Chapter 2 in a lab-based experiment in which participants (N = 30) rated affect intensity, and reported which emotions and bodily sensations experienced in response to generating both mental images and verbal thoughts. Mental imagery was found to enhance emotion complexity as reflected in the richness of the reports provided, with affect intensity and autobiographical recall accounting for the effect. Finally, in Chapter 4 the influence of mental imagery on emotion complexity will be studied across the imagery spectrum. To this end, in an online experiment participants (N = 72) completed measures of imagery vividness, alexithymia and gave written reports on how they feel when experiencing emotions at varying levels of valence and arousal, to obtain indexes on the complexity of emotion conceptualization. In line with the predominant literature, the more vivid the visual mental imagery of participants, the less alexithymia was reported, i.e., the less impaired is the process of emotion conceptualization. As highlighted in the final chapter (Chapter 5), overall the present dissertation contributes to deepening the study of the relationship between mental imagery and emotion. Assuming variation as inherent to emotion, consistently through different experimental designs, methods, languages and indexes, it is shown how context-specific simulations enrich the emotional sphere, by enhancing the complexity of affective experiences.
ལ་ལས་བུ་རབས་ཚ་རྒྱུད་དང་བཅས་རང་སྐད་ཡལ་བར་དོར་ནས་གཞན་སྐད་ལ་བརྩོན་པ་ནི། མི་རིགས་རང་ལ་མཚོན་ན། རྟ་ཐོག་ནས་གཡག་ཐོག་ཏུ་ཞོན་པ་ལས་མ་འདས་་་་། Some people make their progeny abandon [their] own language (rang skad, རང་སྐད།) and strive for others’ language (gzhan skad, གཞན་སྐད།) instead. Thinking from the nation's perspective, [this is] no different from [jumping] off a [fast] horse and riding on a [slow] yak [instead]. In his provocative piece “Legacies of Bandung,” the postcolonial scholar Dipesh Chakrabarty (2005, pp. 4816–4817) compared two responses to linguistic colonialism: On the one hand, the Nigerian writer Chinua Achebe proposes to appropriate the colonial language (English) and reinvent it; on the other hand, the Kenyan writer Ngũgĩ wa Thiong'o advocates an unwavering defense of one's own mother tongue. In this classic postcolonial tension between decolonizing the dominant colonial language vis-à-vis adhering to one's mother tongue, the late tenth Panchen Lama recruits an Indigenous Tibetan metaphor. For the Buddhist master, the question of linguistic choice for any postcolonial subject should be as simple as choosing between riding the tenacious yak or mounting the tractable horse—a no-brainer for anyone with a hint of commonsense in Tibetan pastoral lifestyle. This strong position to defend the Tibetan language continues to resonate with many Tibetan intellectuals today. More recently, in a special issue of the Tibetan humanities journal Yeshe, Ngũgĩ wa Thiong'o was explicitly cited by several Indigenous Tibetan scholars in their decolonial advocacy toward “center[ing] the Tibetan language in Tibetan Studies” and valuing “the richness of Tibetan language” (Gyal 2024, p. 7). Practices of defending Tibetan language sovereignty are ever more resolute nowadays in the censored Tibetan public—as Tibetan language is shrinking in the education sector1 while rigid boarding school policies are separating Tibetan children from their parents.2 Writing between these tensions, Gerald Roche and Shannon Ward, in their respective books, on the one hand, examine the wake of Tibetan linguistic resistance under ongoing Chinese colonialism in Tibet and, on the other hand, introduce a humanistic lens through which to question hierarchies within and beside “the Tibetan language.” I would like to briefly introduce the context that concerns both authors: Like many other Indigenous areas in the world, Tibet is linguistically diverse. As characterized in Tibetan idioms, each different village has its own distinctive “language” or “vernacular” (skad). The Tibetan Empire (AD 618–842) and the later Tibetan Buddhist Ganden Phodrang government (AD 1642–1959) undertook several language standardization projects at different scales, primarily centered around written Tibetan. In contemporary Tibet, three dominant “dialects,” Lhasa, Amdo, and Kham Tibetan, are recognized and institutionalized, supported by broadcasting and educational resources from both the Chinese state and the diasporic government in India (the Central Tibetan Administration). Both books deal with the region where Amdo Tibetan serves as a lingua franca. Following both authors, I use the terms “minority Tibetan speakers” or “minority Tibetans” to refer to those whose mother tongue does not fit into the categories of Amdo, Kham, or Lhasa Tibetan. In his The Politics of Language Oppression in Tibet, Roche focuses on Manegacha, a minoritized language spoken by roughly 8000 speakers3 in the Tibetan valley of Rebgong. Rich in Tibetan and Chinese loan words, Manegacha is recognized by neither the Chinese state nor the majority of Tibetan speakers as a distinct language worthy of preservation and transmission. Critiquing a one-language-one-people assumption—which is implied in both China's nationality framework (Tb: Mirik, Ch: Minzu)4 as well as many Tibetans’ anticolonial resistance toward it—Roche's book witnesses the biopolitical suppression and linguistic erasure of a marginal population (Manegacha speakers) among an already marginalized people (Tibetans). Diverging from a prominent bottom-up call to defend the Tibetan language within and outside of Tibet, Roche demonstrates how languages with smaller social domains, such as Manegacha, are unintentionally marginalized and eradicated through a grassroots effort to defend the Tibetan language against China's colonial policies. Following a similar stance from the periphery of the periphery, Ward's Amdo Lullaby traces the language use of Tibetan children and their caretakers from Tsachen Village. This village speaks “farmer talk,” and the children in question migrated to the multiethnic city of Xining.5 Her fine-grained, microscale research provides a detailed snapshot of intergenerational language choices, contemplations, and challenges among ordinary Tibetans amid uneven state policies and rapid developmentalism. Sympathetic to parents’ frustration with their children's mixed language use, Ward nevertheless emphasizes children's agency, treating their mixed expressions as a generative effort to sustain the vitality of Tibetan language(s). Here, when linguistic concerns in Tibetan margins are foregrounded, the tenth Panchen Lama's deictically anchored categories of “[one's] own language” (rang skad) and “others’ language” (gzhan skad) start to entail much more beyond named languages like “Tibetan” and “Chinese.” Classic research on language shift has viewed the death of a minoritized language like Manegacha with scientific reserve, conducting meticulous analyses while remaining non-interventionist (Fishman 1964). Siding with younger generations who often pioneer new language choices, linguists and linguistic anthropologists often resist their own intuitive impulse to wish for a language's preservation (see Gal 1979; Kulick 1992). Roche, however, takes the diminishing of a minoritized language as a symptom of linguistic oppression, a general manifestation of intergenerational violence, and a “part of broader patterns of oppression and violence” (Roche 2024, p. 6). Adopting a stance that calls for linguistic rights and “a minority Tibetan standpoint” (Roche 2024, p. 8), Roche critiques both state policies that initiated violence and colonized communities who unintentionally compound state violence onto recursively marginalized others. Similarly, Ward critiques the popular demand for standardized uses of the Tibetan language as “mirroring the state's emphasis on language standardization as a marker of national identity,” which further marginalizes Tibetan children living in urban areas (Ward 2024, pp. 131–132). To be clear, Roche and Ward respectively address two distinct kinds of language shift in their books—both imbricated within a larger linguistic ecology. Roche is concerned with the shift from Manegacha and Tibetan bilingualism to Amdo Tibetan monolingualism (and possibly bilingualism with Putonghua). At the same time, Ward focuses on the shift from Amdo Tibetan monolingualism to Putonghua monolingualism (as well as Tibetan-Putonghua bilingualism in effect). A common backdrop for both studies is China's political, economic, and linguistic colonization of Tibet, together with strong linguistic resistance in heavily monitored Tibetan public spaces. Against this backdrop, Putonghua is state-sponsored and explicitly promoted, while a standardized Tibetan language is both the language of direct state control, as well as the unified language for articulating anticolonial resistance. Here, languages and vernaculars of smaller social domains like Manegacha or any singular farmer talk (rong skad) are situated at a doubly marginalized position—readily disposable by a colonial governance project that seeks large-scale population control, on the one hand; while being easily rejected in light of publicly admirable anticolonial resistance, on the other. Overall, both books critique what can be described as the colonial installation and public pluralization of “recursive monolingualism” in Tibet. Inspired by both books, I use the term “recursive monolingualism” to refer to the dialectical process between the Chinese state's mandate to impose Putonghua monolingualism on Tibetan subjects and Tibetan communities’ equally monolingual anxiety over the loss of the Tibetan language. Here, monolingualism informs both the state's coercive practices and anticolonial groups’ counterstrikes. However, putting the seemingly unassailable structural analysis aside, we must ask: How do Tibetan communities themselves think about Tibetans who speak otherwise (or whom both authors call “minority Tibetans”)? What are young Tibetans’ own critiques of what they call the Tibetan “language police”? Here, despite their insights, both books also fall short in seeing the minority and majority Tibetan subjects’ own reflexive critiques of the Tibetan linguistic hierarchy. Also, neither author addresses emerging Tibetan voices that advocate for Tibetan multilingualism and diverse Tibetan representation in recent years. As I will demonstrate in this review essay, the latter trend includes not only traditional Tibetan linguistic scholarship but also emerging domains of Tibetan women's literature, popular music, and other audio-visual social media. Both books, in conjunction with recent language-related scholarship on Tibet, inspire a further question: What would a Tibetan sociolinguistics look like? By “Tibetan sociolinguistics,” I mean an assemblage of metalinguistic analyses, language ideologies, as well as metapragmatic norms and stances concerning Tibetan subjects’ own analyses and practices regarding the languages around them. Here, Tibetan sociolinguistics might be seen as situated in a series of thick multilayered backdrops: with historical sediments of diverse Indigenous regional languages/dialects (Roche 2014) as well as multiethnic language contact (Vasantkumar 2014); linguistic colonialism and resistance since China's Cultural Revolution (Shakya 1994; Willock 2021); vibrant practices of honorifics (Agha 1998), humilifics (Samdrup and Suzuki 2019), and other oral and literary traditions (Jabb 2019; Thurston 2024); contemporary developmental colonialism that threatens a Tibetan lingua franca (Schutte Ke 2024); histories of Buddhist multilingual translation and traditions of book and 2024); diasporic that an emerging public that of and this scholarship to ask: What and subjects living on the Tibetan over the several toward living through as well as linguistic by with colonial linguistic as a population with a is one of the on the and colonial violence, one often to on marginalized others. from of postcolonial of violence against the that violence is to the who this has that violence is within “the structural of and colonial and p. of colonial are seen to and the violence of to the in contemporary Tibet, Chinese state violence is in social not in is against the backdrop of in what can be “recursive I Gerald book on the of in Tibet. The of The Politics of Language Oppression in Tibet biopolitical and state as the of linguistic and marginalized Tibetans to a of Roche the of Tibetans who speak a minoritized language like Manegacha, as to not only through the broader biopolitical of oppression within but also a that continues to and minority languages in the policies in the of to the of Roche that the is not one as Roche the population to a of one for each of the with “Tibetan” being one of the and being the (Roche 2024, p. on of the biopolitical Roche how the state is and This is by the state's of that sustain a singular Tibetan language and its while languages in Tibet like Manegacha, At the same time, the biopolitical state is also and in the that minoritized languages be in for the state to its (Roche 2024, p. Roche, a prominent of studies has the biopolitical to how the Chinese since while in its governance such studies how biopolitical in minority nor how language use more Roche provides a that the of Tibetan language and Tibetan since the a biopolitical in the Tibetan language while other languages in the region Here, Tibetan scholars in as of the traditional Tibetan language and of Tibetan However, an question for scholars on the of Tibetan scholars within Chinese are already of a biopolitical colonial what would the general call for research and mean in this of the What are the and in a decolonial or a that to with the while with the The of the book a provocative that the Tibetan communities’ resistance to state continues to state violence and further speakers of other minoritized languages in Tibet. Writing about anticolonial resistance within the Roche how minoritized languages like Manegacha, as a of the of Tibetan As Roche the Tibetan language is to be by and a into language. In this from of from a and and through (Roche 2024, p. Here, in light of over Putonghua the majority Tibetan to a Manegacha speakers to the language shift that Roche from bilingualism to Tibetan Roche critiques the Tibetan outside of Tibet as also from state from and to India and the of whom do not minoritized Here, research on linguistic in Tibet and has these anticolonial and Thurston Roche is in how minoritized linguistic fall to such linguistic In this Roche also linguistic for different compared with who with generations and subjects Roche the loss of like Manegacha, in However, as a I not with to “language” as an social I that Roche has an to the of and reflexive language well as the between and the languages they Manegacha an or or Here, a or analyses would to the and for the and of a language like of Manegacha compared with Tibetan and Chinese the that the book to The of the book the process of Manegacha erasure in and with in the scholarship on language Roche demonstrates that Manegacha speakers themselves are to the choice of and the Tibetan language. Roche the that such can school and violence” that Manegacha speakers are to by other Tibetan speakers (Roche 2024, p. The Manegacha language was described as language” or at by Tibetan speakers in the Manegacha with of in public Roche Tibetan subjects who direct such toward Manegacha speakers as who to and those who do not fit the As Roche at the of the these are of violence between and within (Roche 2024, p. toward the of the Roche Tibetan violence against Manegacha this analysis with while the violence that ordinary Tibetans from government as well as policies. what in are the many subjects’ equally toward Tibetan Overall, The Politics of Language Oppression in Tibet a provocative of an minoritized languages in Tibet beyond Tibetan. The book the of monolingualism and a that violence continues to communities under colonial However, from a minority Tibetan different from minority Tibetans’ own As Roche many Manegacha of whom are Tibetan Buddhist to the Tibetan of Manegacha and not make in (Roche 2024, p. By Roche Manegacha own toward the language as and (Roche 2024, p. As a I would further ask: which language the the and by which these or by multilingual Manegacha To what are these from or of state By these as under colonization or “the of (Roche 2024, p. to their patterns of an to his toward on as well as in a singular of the and for or a on the of subjects under To the and of the colonized subjects is not to for under colonial is to these subjects’ and with more while the colonial with that resistance, and advocacy against state both within and outside of Tibet are and Roche a new of and Tibetans in Tibet are at the Manegacha speakers at the of the and outside of are at the This further Tibetans and minority Tibetans under Chinese while subjects as As Roche violence and “the of are and in colonized Tibet and in at (Roche 2024, p. Here, we also that who the same state and oppression while conducting research in Tibet, can be from the and violence of for a book with a that to language oppression in to (Roche 2024, p. language shift is not an of many studies in Amdo Lullaby provides a living of under an ongoing colonial like and colonial are the diverse Tibetan In this to is also to a for and for a to As scholarship on is the as to p. subjects the uneven of education and Here, Tibetan children's linguistic and the language a from which to and of structural and on of in Amdo Tibet, Ward an from a village that speaks “farmer talk,” as they between and urban In Ward a and analysis of Tibetan children's language amid a general trend of language shift from Amdo Tibetan monolingualism to Putonghua As the book Ward's with Amdo Tibet also from of Tibetan languages and Amdo as well as with the diasporic Tibetan communities in Here, an for the book is also a between Tibetan in the multiethnic city of in and the multiethnic of Ward that language or the loss of one's mother tongue, should not be to be of (Ward 2024, p. the Ward demonstrates that Tibetan children not their mother as Tibetan communities and Tibet often these children new urban and language into traditional Tibetan between and (Ward 2024, p. emerging linguistic among a young of Ward are also Tibetan Ward Amdo Lullaby with an of a that children with one at the the in their a Tibetan village to the of or Ward how children the they and well as how they to as from these This seemingly children's to be in social when also their to their language children's deictically of these Ward an Tibetan metalinguistic or are among Amdo Tibetan and often the Indigenous region where the speakers or their As Ward Amdo Tibetan speakers the into “farmer (rong skad) and with the many while being more to Chinese while the latter is as more of Tibetan Here, as Ward Tibetan children's language are within social are by language by the of anticolonial of the language histories of within the multiethnic emerging for and educational as well as traditional Tibetan between and The the classic linguistic research and Tibetan who use in their to with one In the children are also as well as about their and urban in to one In this Ward a of how children and with how children their among and how urban between Tibetan and Putonghua Ward provides a linguistic of the Amdo Tibetan a of and a of Ward how children to use and as well as into (Ward 2024, p. In and children's as they where to in their Ward that children to to the by this same by their to the (Ward 2024, p. in the Ward how Amdo Tibetan caretakers use to project and in urban (Ward 2024, pp. Ward also that a whose dominant language to be Amdo Tibetan, to in their The on expressions of as well as of and among Tibetan children and on studies of language and Ward Tibetan of talk about of and children's Ward is not but linguistic and Ward and in between Tibetan children and their children to shift to and to of Here, Ward demonstrates how serves as the for Tibetan children's and which Tibetan later as of Tibetan and The language by the state's language To this “a language that of literary Tibetan over diverse Amdo as a and (Ward 2024, p. Here, Ward as a general among colonial subjects that one's mother tongue and from one's and demonstrates the of this language and in among the “farmer The in an despite Tibetan children's Chinese and a of multilingualism that for Tibetan in (Ward 2024, p. the such might be to a where intergenerational among Tibetans is to However, Ward demonstrates how Tibetan children the Tibetan of from the and to Chinese as well as how they Tibetan in By Ward despite of language ideologies, Tibetan of languages of (Ward 2024, pp. In Amdo Lullaby an of language within a at a time, analysis and that the of linguistic research at its on of language and as well as this book not only to into the of Tibetan but also about Tibetan Buddhist of and in Indigenous Tibetan Amdo Lullaby that and not only but also and the of Tibetan by often by scholars of Tibet Tibetan intellectuals as well as who their concerns on the loss and of Here, the vitality of is In a dialectical process of Putonghua monolingualism continues to with the communities’ equally monolingual anxiety over a Tibetan language as Ward might be to as speakers and language As Amdo Lullaby further in among as well as between children and their are the for the and of Tibetan in a Tibet who for Tibetan rights should the Chinese government on the rights of Tibetan caretakers in their children within their own Here, the that Tibetan are to with their children has compared with their to China's recent policies of boarding in Tibet. linguists and to Tibetans by into of a minority a minority of that To on the of both books, a more What we on to their and these as or Chinese colonialism (as in both Roche 2024, p. Ward 2024, p. What can and of as well as ongoing about general beyond of the biopolitical state and Here, I an of Tibetan sociolinguistics and scholars to think with minority and majority Tibetan who are with state colonialism amid diverse Indigenous languages and This to the that those who under colonialism must be in and by colonial or language we might on the one hand, to and, on the to the but of to the of or to “the p. To be this to while to the by the Chinese which Tibetan voices in scholars outside of from Tibet studies pp. In this I will a and of for such This for such is minority Tibetan scholars many of whom to be in traditional Tibetan How do multilingual Tibetan linguists themselves the multilingual of the Tibetan In a minority Tibetan scholar the of minority Tibetan and from In what calls “a about against and about and p. his of the between his mother tongue, the skad, and the Tibetan from pp. the Cultural and his to an otherwise marginalized language or in Tibetan the skad, as a language for the Tibetan. other minority and majority Tibetan scholars supported this historical linguistic To be is for linguistic anthropologists to this as by “language might also to be a scientific for scholars who are with the linguistic the is also a metalinguistic framework that Tibetan scholars on the one hand, to those who speak around and, on the to the of has by several Chinese state Similarly, in the context of as as the Tibetan scholar has historical and linguistic among minority Tibetans living around Manegacha Here, Manegacha, be recognized as a or skad, and by Tibetan linguists and their A is the of minoritized Tibetans in Tibetan For the prominent Tibetan writer has a of for a young from minority Tibetan whose mother tongue does not fit into the recognized categories of Amdo, Kham, or Lhasa Tibetan. In the of Tibetan multilingualism a in the and between the two The a Tibetan from who speaks the Amdo Tibetan a at a the a that like Tibetan but not Tibetan, which is the p. This with a different Tibetan language has a of in the On the one hand, the own linguistic frustration at not being to the despite being to On the other hand, for the young the Amdo Tibetan and a Tibetan the of the the to in a that written Tibetan. The was by the with a and to the Here, in the and of of a Tibetan literary public is also their humanistic and Tibetan with multilingual on the The trend I introduce is within a of Tibetan and for or around their diverse regional languages or This has through the use of the Tibetan and social By standardized Amdo, Kham, and Lhasa Tibetan and written by broadcasting popular of such as To be colonial as both and language are heavily censored in Chinese social like and the same those concerning minority Tibetan also as of the I the same that both books Language is not only at the of Tibetan but also a for Tibetan and speakers who and spoken at about A Tibetan sociolinguistics would be situated at the of the of and Tibetan Here, the does not the of a but informs through multilingual I for the and from and are I the review to the I about the of one of the book authors, Shannon I that I not the is to
This study explores the impact of annotation inconsistencies in Universal Dependencies (UD) treebanks on typological research in computational linguistics.UD provides a standardized framework for cross-linguistic annotation, facilitating large-scale empirical studies on linguistic diversity and universals.However, despite rigorous guidelines, annotation inconsistencies persist across treebanks.The objective of this paper is to assess how these inconsistencies affect typological universals, linguistic descriptions, and complexity metrics.We analyze systematic annotation errors in multiple UD treebanks, focusing on morphological features.Case studies on Spanish and Dutch demonstrate how differing annotation decisions within the same language create contradictory typological profiles.We classify the errors into two main categories: overgeneration errors (features incorrectly annotated, since do not actually exist in a language) and data omission errors (inconsistent or incomplete annotation of features that do exist).Our results show that these inconsistencies significantly distort typological analyses, leading to false generalizations and miscalculations of linguistic complexity.We propose methodological safeguards for typological research using UD data.Our findings highlight the need for methodological improvements to ensure more reliable cross-linguistic generalizations in computational typology.
The article examines the structure and functionality of the Machine fund of the Bashkir language (MFBL), which was established at the Institute of History, Language and Literature of the UFIC RAS. The MFBL is an integrated system designed to search for linguistic information. It includes several databases. Work on the creation of the MFBL began in 2006. At the moment, the fund consists of ten major sections: general dictionary; lexicographic databases; grammatical bases; experimental phonetic bases; catalogues of handwritten and old printed books; dialectological base; corpus databases containing texts of prose, journalistic, and folklore works in the Bashkir language. In total, the fund contains 75 different linguistic databases. The MFBL system was developed based on the ORACLE database management system. The inclusion of linguistic data in a relational database requires careful analysis and separation of information into its component parts, which makes it possible to efficiently and quickly obtain generalized characteristics. The machine fund of the language has not only scientific, but also practical significance. It is a tool for optimizing and improving the quality of educational materials, including the preparation of language examples for textbooks and teaching aids. The Ministry of Education of the Republic of Bashkortostan actively promotes the use of the Machine Fund among teachers of the Bashkir language. The use of Machine Resources by editors, journalists, and translators certainly contributes to improving the level of Bashkir language proficiency. The machine fund of the language also has significant socio-economic value. Due to the fact that a large number of dictionary and grammar materials are available on the Internet, there is no need for the expensive process of republishing and distributing these materials on paper; Automatic search in the foundation’s databases makes it possible to find philological information faster. This, in turn, accelerates the creation of new linguistic developments and didactic materials. Due to the availability of Bashkir language material on the Internet, residents of the republic can be satisfied with the current language and national policy.
Code-switching presents a complex challenge for syntactic analysis, especially in low-resource language settings where annotated data is scarce. While recent work has explored the use of large language models (LLMs) for sequence-level tagging, few approaches systematically investigate how well these models capture syntactic structure in code-switched contexts. Moreover, existing parsers trained on monolingual treebanks often fail to generalize to multilingual and mixed-language input. To address this gap, we introduce the BiLingua Parser, an LLM-based annotation pipeline designed to produce Universal Dependencies (UD) annotations for code-switched text. First, we develop a prompt-based framework for Spanish-English and Spanish-Guaraní data, combining few-shot LLM prompting with expert review. Second, we release two annotated datasets, including the first Spanish-Guaraní UD-parsed corpus. Third, we conduct a detailed syntactic analysis of switch points across language pairs and communicative contexts. Experimental results show that BiLingua Parser achieves up to 95.29% LAS after expert revision, significantly outperforming prior baselines and multilingual parsers. These results show that LLMs, when carefully guided, can serve as practical tools for bootstrapping syntactic resources in under-resourced, code-switched environments. Data and source code are available at https://github.com/N3mika/ParsingProject
The paper presents the analysis of thematic groups of English youth slang, its functional and pragmatic realization in social networks using the most used Instagram and X. It is noted that one of the most noticeable manifestations of modern linguistic dynamics is the emergence and active functioning of Internet slang is a flexible, creative and multifunctional layer of vocabulary, which vividly reflects linguistic innovations, stylistic experiments and current cultural codes of the digital era. Accordingly, the thematic classification of digital slang units is not just a way of ordering. It enables a deeper understanding of the pragmatic intentions of speakers, the specifics of lexical choice, and the dominant emotional codes that operate in the youth and student environment. Modern English language slang, operating in the digital space of Instagram and X, structurally represents several key thematic domains that demonstrate the pragmatic, emotional, and socio-communicative functions of youth speech. The corpus analysis proved the most common use of slang units of an emotionalevaluative nature, reflecting the need for a quick verbal response, assessment of situations and phenomena; then slangisms of self-presentation and style, which fix the priorities of visual culture and aestheticization of everyday life; then words and phrases related to gender identity, topics of flirting, romantic relationships, and social positioning in the sphere of privacy; then comes the slang of psycho-emotional state and self-awareness, which indicates the normalization of the topic of mental health in public discourse, and, finally, units related to the academic sphere and work processes: the vocabulary of deadlines, academic stress and self-irony about productivity. Thus, digital slang is a multifunctional resource, reflecting both cultural markers and the lexical dynamics of Generation Z, as well as the need for constant stylistic balance between irony, self-presentation, and co-creation. It has been determined that each of these groups, namely: emotional-evaluative vocabulary, selfpresentation and style, gender relations, psycho-emotional states, humor and meme culture, as well as academic and career life, represents not only the content priorities of digital youth communication, but also specific functional and pragmatic strategies: expressive self-expression, irony, social identification, cognitive de-escalation or symbolic opposition to norms. The slang does not simply reflect linguistic innovations, but also captures the transformations in the culture, thinking, and values of Generation Z within the context of the hybrid discourse of the digital age.
Counts of positive and negative valence ratings across 1237 trials Matrix displays counts of each possible combination of positive and negative valence rating across all subjects in Study 2.
Phrygian-KUL is a treebank of the ancient Phrygian language for Universal Dependencies (UD). Having originally only annotated the New Phrygian subcorpus, this dataset is continuously being updated to include the entire epigraphic corpus. For more information, please visit the relevant page at the UD project site or the repository on Github.
= 255) viewed 76 pictures with affective content and rated their experienced affect. Facial muscle activity during picture presentation was assessed via electromyography (EMG) as a direct physiological measure of affective reactions. We used a multilevel model to quantify affective awareness as the strength of the intraindividual relationship between a person's EMG reactions and affect ratings. This relationship was positive on average and differed significantly between participants. These individual differences in affective awareness were reliable and stable over time. Affective awareness was higher for women than for men and went along with generally strong affective EMG reactivity and better socioemotional abilities. (PsycInfo Database Record (c) 2025 APA, all rights reserved).
The article presents the results of a cross-cultural affective images perception study by Americans and Russians and reveals the degree of cultural factor influence on the stimuli assessment by American and Russian men and women. The hypothesis is that assessments of affective images by American and Russian respondents will have statistical differences due to the linguistic and cultural specificity of the ethnic groups; it is also assumed there are cross-cultural gender differences in the assessment. The study used the method of psycholinguistic questioning with seven-point scaling. 84 images from the open American database of affective images (“Open Affective Standardized Image Set”) were used as research material. The respondents were 34 men and 58 women. The results of the analysis did not show significant cross-cultural differences in ratings of affective images with reference to valence type or emotional evaluation/response. In general, Americans and Russians had a similar distribution of image ratings. However, a statistically significant difference has been found in the ratings of images with different valence types (P < 0.001). Negative and positive images were rated higher by Russians in terms of emotional evaluation, in contrast to Americans, most of whose emotional responses had neutral ratings. There was also a statistically significant difference in the ratings of different thematic images (P < 0.05). Nature images were rated by Russians as causing a feeling of comfort, while Americans noted their neutral impact on them. Images of objects, on the contrary, received the opposite ratings from the respondents. Moreover, cross-cultural gender differences have been revealed between Russian and American women in image ratings based on emotional evaluation and valence parameters (P < 0.05). Russian women rated most of the images as having a positive or negative impact, while the majority of American women’s ratings tended to be neutral. This confirms the influence of the emotional stimulus, valence type, image theme, as well as gender factor on the processing of emotionally coloured units by representatives of different cultures.
Detection of out-of-distribution (OOD) samples is cru-cial for safe real-world deployment of machine learning models. Recent advances in vision language foundation models have made them capable of detecting OOD sam-ples without requiring in-distribution (ID) images. How-ever, these zero-shot methods often underperform as they do not adequately consider ID class likelihoods in their detection confidence scoring. Hence, we introduce CLIPScope, a zero-shot OOD detection approach that normalizes the confidence score of a sample by class likelihoods, akin to a Bayesian posterior update. Furthermore, CLIPScope incor-porates a novel strategy to mine OOD classes from a large lexical database. It selects class labels that are farthest and nearest to ID classes in terms of CLIP embedding distance to maximize coverage of OOD samples. We conduct ex-tensive ablation studies and empirical evaluations, demon-strating state of the art performance of CLIPScope across various OOD detection benchmarks. Code is available at https://github.com/ful001hao/CLIPScope.
La langue comme vecteur de responsabilisation Auteur: Camille Guerineau Affiliation: Chercheuse indépendante ORCID: 0009-0009-0364-6396 Description Mise en garde sur la sensibilité du sujet Ce projet aborde des dimensions linguistiques, culturelles et religieuses qui touchent à des identités fortes, notamment la langue arabe et la tradition islamique. L’objectif n’est en aucun cas d’essentialiser une langue ou une croyance, ni de les réduire à des mécanismes de déresponsabilisation. Il s’agit d’analyser, dans un cadre strictement scientifique, comment certaines structures discursives et représentations culturelles peuvent interagir avec les normes françaises de responsabilité individuelle, et parfois générer des malentendus. Cette recherche se veut descriptive et explicative, non normative. Elle vise à mettre en lumière des dynamiques linguistiques et socio culturelles dans un contexte migratoire précis, sans généralisation abusive ni jugement de valeur. En adoptant une approche comparative et en mobilisant la linguistique cognitive et la sociolinguistique, le projet entend contribuer à une meilleure compréhension interculturelle et à un dialogue constructif entre institutions et populations concernées. Plan détaillé de thèse Introduction Présentation du sujet: langue, responsabilité et intégration. Contexte migratoire en France (enjeux sociaux, culturels et institutionnels). Justification scientifique et sociale du projet. Problématique et hypothèses. Méthodologie générale. Chapitre 1: Cadre théorique 1.1 Linguistique cognitive et agentivité Hypothèse Sapir Whorf. Théories de l’agentivité (Talmy, Langacker). 1.2 Sociolinguistique et migration Langue et identité en contexte migratoire. Études sur la perception de la responsabilité dans différentes cultures. 1.3 Religion et destin Al qadar dans l’islam. Interaction entre croyance religieuse et représentations sociales. Chapitre 2: Structures linguistiques de l’arabe et du français 2.1 Tournures impersonnelles en arabe dialectal Exemples contrastifs: « le verre s’est cassé » « j’ai cassé le verre » 2.2 Comparaison avec le français Mise en avant du sujet agentif. 2.3 Effets cognitifs et discursifs Influence de la syntaxe sur la perception de l’action et de la responsabilité. Chapitre 3: Étude socio culturelle en contexte migratoire 3.1 Enquête de terrain Entretiens avec jeunes issus de l’immigration maghrébine. Observation des discours dans différents contextes (école, justice, institutions). 3.2 Analyse des représentations Explications des actes: destin, hasard, agentivité, responsabilité. 3.3 Confrontation avec la norme française Responsabilité individuelle et autonomie comme valeurs centrales. Chapitre 4: Résultats et discussion Impact des structures linguistiques sur la perception de la responsabilité. Rôle de la croyance religieuse dans l’externalisation de la responsabilité. Conflits et malentendus interculturels. Implications pour l’intégration sociale et institutionnelle. Conclusion Synthèse des résultats. Limites de l’étude. Perspectives: pédagogie interculturelle, sensibilisation institutionnelle, ouverture vers d’autres langues et contextes migratoires. Méthodologie Corpus linguistique bilingue (arabe dialectal / français). Analyse qualitative (discours, entretiens, observation). Approche comparative (structures linguistiques et représentations sociales). Outils: analyse du discours, linguistique cognitive, sociolinguistique. Introduction développée La langue n’est pas seulement un outil de communication: elle constitue une matrice cognitive et culturelle qui oriente notre rapport au monde. Selon l’hypothèse Sapir Whorf, les structures linguistiques influencent la manière dont les individus perçoivent, catégorisent et interprètent la réalité. Certaines langues mettent en avant l’agent de l’action (« j’ai cassé le verre »), tandis que d’autres privilégient des tournures impersonnelles (« le verre s’est cassé »), ce qui peut modifier la perception de la responsabilité. De nombreuses recherches ont montré que les différences lexicales et grammaticales façonnent la vision du monde des locuteurs: les Inuits distinguent des dizaines de nuances de neige, le breton possède le terme glaz pour désigner une couleur intermédiaire entre bleu et vert, certaines langues sans futur marqué influencent la manière dont leurs locuteurs envisagent la planification. Dans ce cadre, il est pertinent d’interroger la langue arabe, notamment dans ses formes dialectales, et son interaction avec la croyance religieuse au destin (al qadar), afin de comprendre comment ces représentations linguistiques et culturelles peuvent influencer la perception de la responsabilité individuelle en contexte migratoire en France. Bibliographie sélective Sapir, E. (1921). Language: An Introduction to the Study of Speech. Whorf, B. L. (1956). Language, Thought, and Reality. Duranti, A. (1997). Linguistic Anthropology. Hymes, D. (1972). On Communicative Competence. Bourdieu, P. (1982). Ce que parler veut dire. Bandura, A. (1997). Self efficacy: The Exercise of Control. Sayad, A. (1999). La double absence. Kerbrat Orecchioni, C. (2005). Le discours en interaction Guerineau, C. (2023). La langue comme vecteur de responsabilisation. https://doi.org/10.5281/zenodo.17706182
Online survey using 75 animal images. Japanese adults (N = 102) rated the valence, arousal, approach-avoidance, and kawaii (cuteness) of the images on a 9-point scale. The dataset was used in Nittono, H., Kitamura, A., & Ihara, N. (2025, July 8–11). Can holding a huggable pillow modulate affective facial responses to animal pictures? [Poster session]. The 22nd World Congress of Psychophysiology (IOP 2025), Krakow, Poland.
The advent of ChatGPT has profoundly reshaped scientific research practices, particularly in academic writing, where non-native English-speakers (NNES) historically face linguistic barriers. This study investigates whether ChatGPT mitigates these barriers and fosters equity by analyzing lexical complexity shifts across 2.8 million articles from OpenAlex (2020-2024). Using the Measure of Textual Lexical Diversity (MTLD) to quantify vocabulary sophistication and a difference-in-differences (DID) design to identify causal effects, we demonstrate that ChatGPT significantly enhances lexical complexity in NNES-authored abstracts, even after controlling for article-level controls, authorship patterns, and venue norms. Notably, the impact is most pronounced in preprint papers, technology- and biology-related fields and lower-tier journals. These findings provide causal evidence that ChatGPT reduces linguistic disparities and promotes equity in global academia.
Recent experimental studies have examined GOODNESS IS BRIGHTNESS and a host of other primary metaphors. However, complex mappings such as INTELLIGENCE IS BRIGHTNESS have been largely ignored, nor has there been any attempt to distinguish their effects from those of primary metaphors such as GOODNESS IS BRIGHTNESS. The current study assesses both the nonprimary metaphoric mapping INTELLIGENCE IS BRIGHTNESS and the well-documented primary metaphor GOODNESS IS BRIGHTNESS in a visual priming task. The study finds that a bright background encourages photos of faces to be rated as both more intelligent and wellintentioned, though the background does not significantly affect either attribute alone. This suggests that two metaphors with the same source domain can reinforce each other. The study also underscores the difficulty in assessing a non-primary mapping in isolation from other factors.
Este é um estudo em andamento sobre Processamento de Língua Natural de um corpus de Narrativas Clínicas em português brasileiro com duas versões anotadas: uma pela máquina e outra por humanos. A frequência dos rótulos das classes de palavras e das relações de dependência dos tokens em cada versão é calculada e uma análise guiada pelo corpus é realizada, destacando as correções feitas pelos humanos nas anotações da máquina. A comparação dessas anotações permite a criação de treebanks que podem ser usados para treinar novos modelos usando técnicas de aprendizado de máquina e para aprimorar diversas aplicações de Processamento de Língua Natural com corpus da área biomédica. Além disso, essa comparação permite a análise da consistência teórica de anotação, a fim de identificar o sistema gramatical desse tipo de corpus e criar guias de anotação para Narrativas Clínicas em português brasileiro de acordo com as Dependências Universais.
BACKGROUND AND OBJECTIVES: Narrative discourse is a useful means to organize ideas and create shared understandings. Clinically, performing discourse analysis on disordered spoken language could facilitate researchers and clinicians not only to evaluate one's language abilities but also to foreshadow his/her communication in real-life situations. Given the normative reference data of a specific discourse task, less-biased judgement and evaluation could be made, which could further facilitate assessment and intervention planning. This study aims to first develop norms by analysing the language samples produced by neurotypical Cantonese speakers on two well-familiarized narrative stories, The Boy Who Cried Wolf, and The Tortoise and Hare. Second, we aim to investigate the potential age and education effects on a wide range of micro- and macro-structural linguistics measures. METHOD: Two semi-spontaneous story narratives from the Cantonese AphasiaBank were selected for scoring. A total of 150 neurotypical Cantonese adult speakers produced the spoken discourse samples for each story narrative. All speakers were native Cantonese speakers living in Hong Kong; they were divided into three age groups: young (18-39 years old), middle-aged (40-59 years old), and older (> 60 years old). Audio recordings were transcribed, segmented, and annotated using CHAT conventions. RESULTS: Normative references of various micro- and macro-structural linguistics measures and the standard scoring references for the two narrative stories were established. For the age effect on narrative discourse, the older adults produced less complex, coherent and thematic-related concepts compared to the young group. However, lexical diversity was preserved in the older group, resulting in no significant differences across the three age groups. For education effect, the higher education group outperformed the lower education group in verbal productiveness and content informativeness. Lastly, the two stories were found to be non-comparable to each other, thereby they should not interchange in pre- and post-test arrangements or in monitoring discourse performance. CONCLUSIONS: The Cantonese discourse norms presented here can be applied in both research and clinical settings, facilitating a more objective review of language impairment and treatment planning. Second, this study demonstrated the effect of normal ageing on both the linguistics and conceptual levels specific to discourse production. WHAT THIS PAPER ADDS: What is already known on this subject Discourse analysis is a critical part of evaluating and understanding a person's communication abilities. Studies indicated that narrative discourse is more sensitive to specific linguistic parameters than other genres, and people with stronger narrative skills tended to have more social communication opportunities. An increasing number of studies had been working on setting norms for discourse tasks. Locally, Kong et al. (2025) have recently reported the normative references for descriptive tasks of the Cantonese AphasiaBank. What this study adds to the existing knowledge First, this study completed the norms establishment for all discourse tasks of the Cantonese AphasiaBank. Second, our analysis of the impact of different factors on narrative discourse offers significant value for clinical applications. We found that ageing was not manifested across all microstructural linguistics consistently, while lexical diversity was found to be tolerant to ageing. However, ageing was found to be adversely affecting propositional parameters and discourse informativeness. What are the clinical implications of this study? Narrative discourse, being one of the most popular tasks in clinical assessment but often faced the challenges of a lack of objective references. This study analysed two well-familiarized narrative stories, provides a complete set of norm data readily for front line clinicians and researchers, which could be used for intervention planning, monitoring progress (e.g., used them as control probes for tracking generalization effects) or as an input for investigating the interplay between linguistics and cognitive abilities.
Abstract Previous research on syntactic complexity is primarily focused on the synchronic distribution of clausal and phrasal features and the diachronic shift from clausal elaboration to phrasal compression. However, the interrelationship between clause complexity and phrase complexity remains unexplored. This study investigated syntactic complexity at different linguistic levels across three disciplinary groups (Social Sciences, Humanities and Natural Sciences) using a corpus of research article abstracts. Sentence complexity was measured by the number of clauses per sentence, clause complexity by the number of clausal constituents per clause, and nominal group (NG) complexity by the number of words per NG. The results show that: (1) sentences are the least complex in Natural Science (NS) texts; (2) clauses are also the least complex in NS, despite having the highest average number of clausal constituents; (3) NGs are the most complex in NS texts. Furthermore, the study found that NG complexity could be more accurately measured by the number of premodifiers of the head noun (HN) of the NG. These findings have important implications for instructing English as a Foreign Language (EFL) learners in discipline-specific academic writing.
Abstract Research on the progressive aspect in Germanic and Romance languages has benefited from corpus data. A transparent, objective, reliable and replicable identification of such constructions in corpora is however challenging. The present chapter presents preliminary methodological work in automatically retrieving and counting authentic examples from treebanks, that is, grammatically annotated corpora. It demonstrates how selected constructions that mark the progressive in Italian and Norwegian are collected from treebanks accessible through the INESS platform. Deep syntactic relations such as those between predicates and arguments are factored in and quantified. Corpus queries that exploit syntactic dominance relations are potentially more powerful than queries using only linear precedence, but there is a relative shortage of treebank resources.
This paper is an introduction to using the Treebank Semantics Parsed Corpus (TSPC) and its online interface.The TSPC is a collection of English with hand worked tree annotation for approaching half-a-million words.The annotation gives a resource of general use but is notable as content to feed a calculation of meaning representations for insights beyond surface syntax.The online interface has functionalities of search, summary and visualisation for presenting the source files of the corpus.The paper ends with a case study demonstrating how to access data for a language education task, with pinpoint insights gained because of the detailed analysis offered by the corpus.
This article is dedicated to analyzing the contemporary significance of sociolinguistics and issues related to the development of language in social networks. In today’s era of globalization, social networks have become one of the primary platforms for communication, where new forms and styles of language are emerging. The study examines the dynamics of language in social networks, the emergence of new lexical units and their application in society, the impact on language norms, and changes in language at an international level. The article also analyzes the linguistic features developing in social networks and their sociolinguistic aspects. This research aims to deepen the understanding of the interrelationship between social network language and sociolinguistics and to identify their future developmental directions.
Large language models (LLM) perform outstandingly in various downstream tasks.However, there is limited understanding regarding how these models internalize linguistic knowledge, so various linguistic benchmarks have recently been proposed to facilitate syntactic evaluation of language models (LM) across languages.This paper introduces QFrCoLA (Quebec-French Corpus of Linguistic Acceptability Judgments), a normative binary acceptability judgments dataset comprising 25,153 in-domain and 2,675 out-of-domain sentences.Our study leverages the QFrCoLA dataset and seven other linguistic binary acceptability judgments corpus to benchmark eight LM.The results demonstrate that, on average, finetuned Transformer-based LM are strong baselines for most languages and that zero-shot binary classification LLM perform worse than the naive baseline on the task.However, for the QFrCoLA benchmark, on average, a finetuned Transformer-based LM outperformed other methods tested.It also shows that pretrained cross-lingual LLMs selected for our experimentation do not seem to have acquired linguistic judgment capabilities during their pre-training for Quebec French.Finally, our experiment results on QFrCoLA show that our dataset, built from examples that illustrate linguistic norms rather than speakers' feelings, is similar to linguistic acceptability judgment; it is a challenging dataset that can benchmark LM on their linguistic judgment capabilities.
Horses are depended on as work animals by humans and are used in leisure and sport across the world, but the extent to which humans can recognise pain in horse faces is not known, which could impact their welfare. There are also significant gaps in our understanding of which psychological traits influence recognition of human facial expressions of pain. To address this, one hundred participants, with either some (N = 30) or no prior horse care experience (N = 70), rated thirty human and thirty horse faces for pain, arousal and valence and completed trait measures of empathy and social anxiety. Ten equine behaviour professionals also rated the horse faces as a baseline for assessing accuracy. Overall, accuracy of pain recognition was higher for human faces, but participants with horse experience were more accurate at pain recognition in horse faces, than those without, and years of horse experience predicted horse pain recognition accuracy. Social anxiety traits predicted accuracy of pain recognition in human but not horse faces, while also predicting subjective ratings of pain in horse but not human faces. Empathy and its cognitive and emotional components were not related to pain recognition accuracy or ratings of horse or human faces. Relationships between trait measures and arousal and valence ratings for both species are reported. This study is the first to report the human ability to read pain in horse faces and the factors which influence this and extends current knowledge on face processing in social anxiety.
OBJECTIVE: The term "active larynx" is a nonspecific and subjective term used by otolaryngologists to describe laryngeal inflammation that can influence the timing of airway reconstruction. We sought to measure the reliability of visual assessments of laryngeal inflammation for later scale development. STUDY DESIGN: A cross-sectional study. SETTING: Pediatric tertiary care center. METHODS: We created an image library from a direct laryngoscopy and bronchoscopy database. Blinded judges were asked to rate the characteristics of laryngeal inflammation (edema, erythema, cobblestoned appearance, and ventricular eversion; 5-point Likert scale), the overall "activeness" of the larynx (10-point scale), and whether laryngeal inflammation would influence a delay in reconstructive surgery (yes/no). A tentative scale was also constructed. Intraclass correlations with 2-way random effects, and Fleiss's κ were used to evaluate interrater reliability. The convergent and discriminant validity of the tentative scale were measured. RESULTS: Three pediatric otolaryngologists reviewed 15 larynges for a total of 45 image ratings. Intraclass coefficients indicated substantial agreement for edema (0.76) and erythema (0.83) and moderate agreement for ventricular eversion (0.58). Cobblestoning had low agreement (intraclass correlation coefficient [ICC] < 0.20). The agreement was substantial for overall "activeness" (ICC 0.76) and moderate for whether inflammation would delay surgery (ICC 0.47). By Fleiss's κ, edema and erythema had moderate agreement (0.50 and 0.61, respectively), whereas all others had poor agreement. The convergent and discriminant validity of the tentative scale were reassuring. CONCLUSION: While the reliability of laryngeal inflammation by visual assessment is variable, the creation of an active larynx scale appears feasible.