Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Emotions comprise multiple coordinated responses, including facial expressions and subjective experiences. Although functional magnetic resonance imaging (fMRI) studies have identified brain regions associated with facial or emotional responses, few have simultaneously assessed and statistically dissociated components. Additionally, the functional networks associated with these emotional responses and the dynamic interplay between such networks remain uncertain. To investigate these issues, we measured fMRI while participants viewed emotional films, as well as their facial videos and dynamic valence ratings. Regional activity analysis revealed that facial responses (lip-corner-pulling actions) were associated with activation in the limbic regions and somatosensory motor cortices. Subjective emotional responses (the absolute values of valence ratings) were associated with activity in the medial parietal and lateral temporoparietal cortices. Independent component analysis revealed that the independent components associated with facial and subjective responses included the abovementioned activated regions. Dynamic causal modeling of these independent components supported a model in which the visual/auditory processing component modulated the facial response component, which subsequently influenced the subjective response component. Our findings imply that, during emotional processing, facial responses are initially generated by the limbic and sensorimotor cortical networks; subsequently, these responses give rise to subjective experiences through activity in the medial parietal-lateral temporoparietal networks.
This paper presents a new version of the Spoken Slovenian Treebank (SST), a balanced and representative collection of transcribed spontaneous speech with manually annotated lemmas, part-of-speech tags, morphological features, and syntactic dependencies, recently expanded with over 3,000 newly annotated utterances. After a brief overview of the data sampling, anno-tation, and consolidation processes—presented in detail in previous work—we evaluate the significance of this new language resource for both linguistic research and natural language pro-cessing by first highlighting its distinctive lexical and morphosyntactic features in comparison to writing, and then assessing their impact on the performance of tools for automatic grammatical annotation. Finally, we reflect on the methodological insights gained during treebank creation, discuss the potential of SST for advancing spoken language research, and argue for the necessity of such resources in supporting linguistic diversity in language technology.
Abstract The paper introduces Parallel Trees, a novel multilingual treebank collection that includes 20 treebanks for 10 languages. The distinguishing property of this resource is that the sentences of each language are annotated using two syntactic representation paradigms (SRPs), respectively based on the notions of dependency and constituency. By aligning the annotations of existing resources, Parallel Trees represents an example of exploiting pre-existing treebanks to adapt them to novel applications. To illustrate its potential, we present a case study where the resource is employed as a benchmark to investigate whether and how BERT, one of the first prominent neural language models (NLMs), is sensitive to the dependency- and constituency-based approaches for representing the syntactic structure of a sentence. The case study results indicate that the model’s sensitivity fluctuates across languages and experimental settings. The unique nature of the Parallel Trees resource creates the prerequisites for innovative studies comparing dependency and phrase-structure trees, allowing for more focused investigations without the interference of lexical variation.
This paper presents the creation of a Universal Dependency (UD) treebank for Amahuaca (Peru), marking the first UD treebank within the Headwaters subbranch of the Panoan family, spoken mostly in Peru and Brazil. While the UD guidelines provided a general framework for our annotations, language-specific decisions were necessary due to the rich morphology of the Amahuaca language. The paper also describes specific constructions to initiate a discussion on several general UD annotation guidelines, particularly those concerning clitics and morpheme-level dependencies.
Rapid Eye Movement Sleep (REM) is thought to process emotions via memory reactivation. Such REM reactivation can be triggered by presenting a tone associated with the target memory. This reduces subjective arousal ratings for negative stimuli. Here, we measure arousal objectively in brain and autonomic system. Participants rated negative image-sound pairs, half of which were then re-presented during subsequent REM. All images were re-rated in a Magnetic Resonance Imaging (MRI) scanner with pulse oximetry 48 h after encoding. Reactivation in REM reduced responses in the brain's Salience Network (SN), including Anterior Insula and dorsal Anterior Cingulate Cortex (dACC), and associated emotion-processing regions: orbitofrontal cortex, subgenual cingulate, and left amygdala. Memory reactivation in REM reduced heart rate deceleration (HRD). Subjective arousal ratings were reduced for more upsetting images and increased for less upsetting images. Our findings have implications for the use of memory reactivation to treat depression and anxiety disorders.
Violence has been a significant phenomenon throughout human history and has been studied across various disciplines. With the rise of interdisciplinary approaches, it has become a central issue in many fields. Adapting a psycholinguistic perspective, this study examines how films with violent content influence people’s perception of emotions. It was hypothesized that the violent content in film the participants saw would alter their perception of the emotional content of the words they were shown. To this end, they performed a rating task on a list of positive, negative, and neutral words before and after watching a violent film. A comparison of pre- and post-ratings revealed that valence ratings decreased for all word types after watching the film, while arousal ratings remained unchanged. These findings suggest that exposure to violent content can influence how emotional words are perceived. The results provide valuable insights into the impact of violence on emotional processing.
Timbral blend is a phenomenon that occurs when two or more concurrent acoustic events produced by distinct sources fuse perceptually and give rise to new timbres. Auditory scene analysis proposes that concurrent grouping cues of onset synchrony, harmonicity, and parallel change in pitch and dynamics are involved in the perceptual fusion of events, but research has also shown that several timbral cues can affect concurrent grouping. We investigated potential factors that may cause different degrees of instrumental blend in orchestral excerpts using rating scales ranging from “unity” to “multiplicity” and from “strongly blended” to “not at all blended.” With linear mixed effects modeling, the factors found to affect ratings included the rating scale used, musical training, timbre class (instrument families involved), the degree of parallelism and onset synchrony of melodic lines involved in the blend, the number of different notes present simultaneously, and several acoustic features related to timbre. Musicians differ from nonmusicians in the use of the multiplicity scale, rating excerpts as more multiple, even if they are fairly well blended, whereas nonmusicians ratings are similar for both scales and to musicians’ ratings of blend. Excerpts with bowed strings and/or woodwinds blend the strongest, followed by combinations involving brass instruments, with excerpts involving percussion and plucked strings blending the least. The important finding of this study on real musical excerpts is in demonstrating the relative roles of the score-based and acoustic factors that are associated with the perception of multiplicity and blend in complex orchestral sonorities as well as the influence of musical training.
Zora Neale Hurston’s Their Eyes Were Watching God presents a radical departure from the tragic mulatta trope in African American literature by centering a black feminist protagonist, Janie Crawford, who is neither defined by racial ambiguity nor constrained by the moral expectations imposed on middle-class black women of the 19th and early 20th centuries. Unlike her literary predecessors, Janie speaks in black vernacular, embraces her sexuality, and ultimately finds agency outside of marriage, despite the novel’s exploration of love and relationships. This paper argues that Hurston’s portrayal of Janie’s three marriages illustrates a pessimistic view of black women’s status within love and marriage, revealing that even true love cannot fully liberate them from patriarchal constraints. Through an analysis of Janie’s relationships, this paper demonstrates how Their Eyes Were Watching God challenges intra-community sexism and critiques the internalization of white patriarchal values by black men. Additionally, it explores Hurston’s literary innovations, particularly her use of black dialect and folklore, as an intervention against white literary standards and a foundation for later black feminist narratives. Hurston’s use of black dialect and folklore functions not merely as a literary gesture, but as a deliberate political and aesthetic intervention. The black vernacular, often seen as non-literary or even “primitive” in dominant white and even black literary standards, becomes in Hurston’s hands a medium of authenticity, resistance, and empowerment. By embedding Janie’s voice within this dialect—particularly through her dialogues with other women and her defiance of male authority—Hurston decentralizes white linguistic norms and reclaims black southern oral traditions as legitimate literary forms.By foregrounding the singularity of Janie’s experience, Hurston’s novel marks a turning point in the representation of black women in literature, paving the way for subsequent authors like Alice Walker and Toni Morrison to further explore Black female autonomy and agency.
In the 20th century, the concept of “compliance” gained prominence in academic publications, particularly within the realms of medical and financial-economic discourse. Recently, its application has extended to fields such as pedagogy, psychology, and philology. This concept encompasses two primary substantive dimensions: the first involves the establishment of requirements aligned with existing norms and standards, while the second pertains to the genuine and conscientious adherence to these stipulations. Contemporary sociological studies are beginning to explore the willingness of certain religious adherents to fulfill tax obligations, necessitating a theoretical framework for understanding these practices within the context of academic religious studies. Historically, the term “compliance” — denoting agreement, indulgence, voluntary concession, and self-restraint — originated in English from Latin ecclesiastical terminology. It encapsulated specific norms governing interreligious and interfaith relations that evolved within Christian culture, particularly during the Reformation, a period marked by significant conflict. The imperative to achieve consensus with dissenters for the collective civil good prompted the formulation of new legal and literary standards, addressing challenges such as the transformation of various “religion names” and “confessional names,” wherein derogatory connotations were supplanted by neutral or respectfully affirming alternatives. This analysis draws upon resources from dictionaries, encyclopedias, the linguistic database “National Corpus of the Russian Language,” and additional scholarly sources.
Identifying novel neuromodulatory targets for deep brain stimulation (DBS) in psychiatric disorders is an urgent clinical need. Equally critical is the discovery of simple oscillatory biomarkers that bridge behaviour and clinical symptoms, enabling personalized treatment strategies. The bed nucleus of the stria terminalis (BNST), a pivotal output structure of the amygdala, is a potential candidate for DBS due to its key role in regulating fear, emotional valence, and prosocial behaviour. However, owing to the small size, its neural dynamics and functional contributions are poorly understood, precluding behavioural-clinical relevance. In a cross-sectional design, we acquired BNST neural recordings from 23 patients with depression undergoing DBS during two tasks: pain perception with painful/non-painful scenarios and an affect task with emotionally valenced images. We first localized the electrode contacts in the BNST and using their neural recordings for further analysis. We subjected the preprocessed data to time frequency decompositions to find condition differences. The significant clusters were then used to link to the behavioural ratings and clinical symptom severity. Furthermore, cross-frequency interactions were also undertaken. Pain perception elicited late theta and alpha activity (∼1s), with theta activity linked to subjective pain ratings and alpha correlating with anxiety/depression scores and anxiety symptoms post-DBS. Negative imagery induced early theta (∼250 ms), resembling previously reported amygdalar responses, and which link to valence ratings, depression and anxiety symptom severity. These results reveal distinct BNST dynamics in depression: early theta for rapid threat processing and late theta/alpha for complex socio-cognitive responses. Task-dependent theta and alpha activity linked behavioural profiles and symptom severity, highlighting BNST's role in behaviourally and clinically relevant oscillatory patterns, contributing novel insights for advancing precision neuromodulation strategies.
Mental health issues such as depression, stress, anxiety, and personality disorders are increasingly prevalent, particularly within online communities. This study proposes a lightweight and efficient multi-class classification framework to identify five mental health conditions using Reddit user-generated posts. While previous studies predominantly rely on conventional CNNs or standard machine learning techniques for binary classification, our work introduces a novel Bidirectional Long Short-Term Memory (BiLSTM) model integrated with an attention mechanism. The architecture is further enhanced by synonym-based data augmentation using the WordNet lexical database, which improves semantic diversity and enhances model robustness, particularly for underrepresented classes. Unlike prior works that focus narrowly on binary classification or employ transformer-based models with high computational demands, our model offers a lightweight, high-performance architecture optimized for multi-class detection and real-world deployment. Experimental results demonstrate that the proposed model achieves a peak validation accuracy of 95.02%, along with precision 95.08%, recall 95.02%, and F1-scores of 95.03%. These findings support the advancement of efficient AI-driven diagnostic systems in mental health analytics and lay the groundwork for future integration into mobile or resource-constrained platforms.
The rapid development of Neural Machine Translation (NMT) and Large Language Models (LLMs) marks what can be described as an Algorithmic Turn in translation studies.This shift fundamentally reconfigures translation from a primarily human-centered act of linguistic mediation to a process increasingly shaped by computational systems and algorithmic logic.This paper critically examines the growing influence of Artificial Intelligence on intercultural narratives and on the process of semantic transfer in an interconnected global context.While AI-driven translation technologies offer unprecedented speed, scalability, and accessibility, their widespread adoption also raises significant concerns related to cultural authenticity, linguistic diversity, and the preservation of meaning.Drawing on a critical-theoretical framework, the study argues that AI systems function not merely as neutral tools but as active co-authors in the construction of crosscultural meaning.The analysis focuses on two central issues: first, the ways in which intercultural narratives are shaped by AI models trained on culturally imbalanced datasets, often privileging dominant linguistic norms and contributing to cultural homogenization; and second, the inherent limitations of AI in handling semantic nuance, particularly in relation to irony, implicit power relations, and culturally embedded meanings.The paper ultimately contends that the Algorithmic Turn produces a paradoxical outcome-a translation that is linguistically fluent yet culturally hollow.Addressing this paradox requires moving beyond accuracy metrics toward a critical evaluation of AI's ethical, cultural, and epistemological implications, reaffirming the essential role of the human translator as an intercultural mediator whose expertise complements, rather than competes with, technological innovation.
The article examines the issue of translating idioms in I.S. Turgenev's novel Fathers and Sons in the context of cultural and historical influences on translation decisions. A comparative analysis of three English translations of the novel, produced in the 19th, 20th, and 21st centuries, is conducted to identify changes in the translation of idiomatic expressions. Special attention is given to the concept of «cultural time» as a factor affecting translation strategies. The study demonstrates that each generation of translators adapts the text according to the prevailing cultural and linguistic norms of its time. The analysis highlights how the perception of Russian idioms evolves and how cultural shifts influence the choice of equivalents in the target language. The findings of this research may be valuable for specialists in literary translation and intercultural communication.
This paper outlines the first operation towards a full update of MariTerm, a WordNet-like resource on maritime terminology developed and maintained by CNR-ILC, in preparation for future compliance with FAIR principles (Wilkinson et al., 2016).The project focused on expanding and linking synsets between ItalWordNet (IWN), a general lexical database for Italian, and MariTerm to enrich IWN with maritime concepts.A semi-automatic pipeline was developed to facilitate this process, prioritizing critical semantic relations and automatic evaluation.Key outcomes include an enriched ItalWordNet with links to MariTerm concepts and a revised MariTerm with connections to IWN synsets.While further refinement is needed, this work marks a significant step toward integrating maritime terminology into ItalWordNet.
Based on the English-Chinese Parallel Corpus of Children’s Literature and the Corpus of Chinese Children’s Literature, this study investigates the feature of normalization in the Chinese translation of English children’s literature. Normalization refers to the adaptation of foreign features in the source text to comply with the cultural and linguistic norms of the target culture. The study analyzes both macro and micro levels of language features in translated children’s literature, comparing them with original Chinese and English texts. The findings reveal a clear trend towards normalization, evidenced by shorter sentences, increased repetition of high-frequency words, a lower frequency of hapax legomena, and a higher textual readability in translated Chinese versions. Furthermore, linguistic structures such as reduplication, modal particles, “把” (BA), and “得” (DE) constructions are found to occur at rates comparable to or significantly higher than those in the original Chinese corpus. This paper argues that normalization is a creative outcome, molded by translators aligning with reader expectations, conscientiously considering the psychological characteristics of child readers, and adapting to social, cultural, and market influences. The study contributes to understanding linguistic features of translated children’s literature, sheds light on translation universals, and underscores the dynamic interplay between normalization and translator creativity.
As part of the reintegration of the traditional and computational lexical research and development activities at the University of Gothenburg, the Swedish Academy's lexical databases are being edited for greater consistency and included in the computational lexical infrastructure of Sprkbanken Text, which in turn is undergoing considerable development in order to meet the requirements of tradi tional lexicography.One aim of this work is to formally interlink these databases with Sprkbanken's Lexical Research Infrastructure.In this chapter we focus on the opportunities for diachronic lexical research offered by the inclusion of one such dataset in this infrastructure.SAOLhist Plus brings together digitized versions of successive editions of the SAOL dictionary, 10 editions covering a timespan of almost a century and a half.Together with some other historical and modern dictionaries and combined with large corpus-extracted vocabularies from the same time interval, we can make wide-ranging and detailed investigations of lexical change in relation to the history of lexicography in Late Modern and Contem porary Swedish.
The aim of this study was to determine whether closed-loop vibration stimulation, delivered at +3% of the heart rate frequency at an imperceptible intensity before waking, could reduce sleep inertia. Participants napped on a bed equipped with a woofer that delivered vibration stimulation every 5 min, starting 30 min before their scheduled wake time. The effects of the stimulation were assessed using a Psychomotor Vigilance Task performed immediately upon waking, along with the analysis of salivary cortisol and melatonin levels, as well as subjective arousal ratings based on the Karolinska Sleepiness Scale and the Stanford Sleepiness Scale. The results indicated that vibration stimulation at +3% of the heart rate frequency improved Psychomotor Vigilance Task reaction times and increased self-reported arousal scores, thus reducing sleep inertia compared with the control condition without stimulation. Additionally, salivary melatonin levels were lower immediately after waking. These findings suggest that closed-loop vibration stimulation at +3% of the heart rate frequency before waking could be an effective method to reduce sleep inertia. This non-invasive approach may facilitate cognitive recovery following sleep. Further research is required to investigate the underlying mechanisms, and confirm these findings across different populations and settings.
One of the main challenges in ontology matching is to match ontologies with high accuracy. Therefore, ontology matching systems typically use multiple basic matchers, each targeting a specific ontology component for the matching process. However, optimizing the combination of these matchers remains an open problem. In this paper, we present CroMatcher 2.0, an improved ontology matching system that aims to overcome these challenges. We introduce two new basic matchers. The first matcher determines correspondence between entities by comparing strings obtained from entity IDs and annotations using the English lexical database to find similarities between tokens of the compared strings, considering their mutual relations (synonyms, hypernyms, etc.). The second matcher determines the correspondence between entities using the special mediator ontology, which is very valuable for ontology matching as it can contain additional information about the compared ontologies. In this paper, we tested this matcher on the Ontology Alignment Evaluation Initiative Anatomy track by using the Uberon mediator ontology, which contains a lot of information about anatomical structures. We also introduce a new weighted aggregation method (Autoweight 3.0) that automatically determines the weighting factors of the basic matchers in the parallel composition. CroMatcher 2.0 was evaluated in three test cases of the Ontology Alignment Evaluation Initiative (Benchmark, Anatomy, and Conference) and showed competitive performance compared to other state-of-the-art systems. The results position CroMatcher 2.0 among the best ontology matching systems for these datasets and confirm the effectiveness of the newly introduced methods.
This research explores how Micro, Small, and Medium Enterprises (MSMEs) in West Sumatra employ code-switching in their advertising discourse to construct linguistic identity, express cultural belonging, and project entrepreneurial modernity. Using Fairclough’s Critical Discourse Analysis (CDA) as the analytical framework, this study examines linguistic features, forms of code-switching, and the underlying ideological meanings within promotional banners and billboards that combine English and Indonesian. The findings reveal that code-switching serves as more than a marketing strategy it functions as a socio-symbolic practice through which entrepreneurs negotiate between local authenticity and global aspirations. The frequent use of English, despite notable errors in diction, spelling, and syntax, underscores its symbolic power as a marker of prestige and progress in the post-pandemic economic landscape. However, these linguistic inaccuracies also indicate challenges in language proficiency and access to educational resources, exposing power asymmetries between local entrepreneurs and global linguistic norms. From a sociolinguistic standpoint, code-switching embodies both empowerment and vulnerability: it enables small businesses to gain visibility in global markets while simultaneously revealing structural inequalities in linguistic capital. The study concludes that language operates as a key site of negotiation where identity, economy, and ideology intersect. It recommends enhancing critical language awareness and multilingual marketing literacy in MSME training programs. Future research is encouraged to examine digital advertising discourses to understand how linguistic entrepreneurship evolves in online spaces. This study contributes to the growing body of literature on sociolinguistics, linguistic entrepreneurship, and the politics of language in Indonesia’s evolving marketplace.
The widespread use of the rating system in the life of almost every person raises the question of the value of their use, the objectivity and comprehensiveness of ratings. In this article, the authors summarized the most common and generally accepted definitions of the word “rating” and proposed an integrated definition of the term. The classification of the concept of “rating” on several grounds is presented, approaches (positions) reflecting the place and role of ratings in modern society are systematized. Based on the information received, the pros and cons of using them are described.
Contextual processing enables the brain to integrate environmental cues for adaptive emotional perception. It is crucial to understand how this ability operates in dynamic environments of short video viewing and how the brain supports it. In this study, we utilized behavioral and neuroimaging experiments to examine contextual processing during short video viewing. Compared with single-face clips, face-context-face sequences elicited coherent emotional perceptions of neutral faces from emotional contexts. A distributed brain network was involved in the top-down modulation of contextual processing. Temporoparietal junction (TPJ) and precuneus showed sustained engagement during context integration. Activity in TPJ and insula was associated with valence and arousal ratings, respectively. Under negative contexts, global functional segregation facilitated contextual processing, and weaker TPJ-insula connectivity corresponded to stronger contextual effects. This study advances the understanding of contextual processing in the digital media age and may inform future investigations into neural correlates of contextual processing dysfunction in psychiatric conditions.
Computational linguists have long recognized the value of version control systems such as Git (and related platforms, e.g., GitHub) when it comes to managing and distributing computer code.However, the benefits of version control remain under-explored for a central activity within computational linguistics: the development of annotated natural language resources.We argue that researchers can employ version control practices to make development workflows more transparent, efficient, consistent, and participatory.We report a proof-of-concept, GitHub-based solution which facilitated the creation of a legal English treebank.
This paper introduces foundational resources and models for natural language processing (NLP) of historical Turkish, a domain that has remained underexplored in computational linguistics. We present the first named entity recognition (NER) dataset, HisTR, and the first Universal Dependencies treebank, OTA-BOUN, for a historical form of the Turkish language along with transformer-based models trained using these datasets for named entity recognition, dependency parsing, and part-of-speech tagging tasks. Furthermore, we introduce the Ottoman Text Corpus (OTC), a clean corpus of transliterated historical Turkish texts that spans a wide range of historical periods. Our experimental results demonstrate prominent improvements in the computational analysis of historical Turkish, achieving strong performance on tasks that require understanding of historical linguistic structures -- specifically, 90.29% F1 in named entity recognition, 73.79% LAS for dependency parsing, and 94.98% F1 for part-of-speech tagging. They also highlight existing challenges, such as domain adaptation and language variations between time periods. All the resources and models presented are available at https://hf.co/bucolin to serve as a benchmark for future progress in historical Turkish NLP.
The article explores comic discourse as a multifaceted phenomenon that forms at the intersection of linguistic and cultural aspects. The main purpose of the work is to analyze the specifics of comic utterance in the context of various cultural realities and language systems. Special attention is paid to the influence of cultural peculiarities on the processes of creating and interpreting humor, as well as linguistic tools that ensure the transmission of a comic effect. The analysis of the humorous language is carried out, aimed at studying the mechanisms of the generation and perception of humor within the framework of linguistic and cultural contexts. The key factors determining the originality of a humorous utterance are identified, as well as the interrelationships between linguistic means and cultural features in the process of comic communication are investigated. The results of the study indicate that comic discourse acts as a reflection of cultural traditions and social norms, and also serves as a tool for analyzing the dynamics of intercultural interaction. It is established that the perception of humor is determined not only by linguistic norms, but also by cultural values, stereotypes and contexts in which communication is carried out. Thus, comic discourse is a significant object of study for understanding the interrelationships between language and culture.
This article explores the theoretical foundations and practical implications of adequacy and equivalence in translation studies. Drawing upon the frameworks of prominent scholars such as Y.I. Retsker, L.S. Barkhudarov, the paper examines how different approaches define and apply the concepts of equivalence in relation to linguistic norms, functional correspondence, and communicative effect. The analysis highlights the distinction between formal, dynamic, and functional equivalence, and the roles they play in achieving accurate translation outcomes. Special attention is given to the translator’s linguistic and cultural competence and the inherent asymmetry in bilingualism, which significantly affects translation choices. Furthermore, the article discusses translation criticism as a tool for evaluating the quality and equivalence of translated texts.
Large language models (LLMs) have achieved remarkable success across various natural language processing (NLP) tasks.However, recent studies suggest that they still face challenges in performing fundamental NLP tasks essential for deep language understanding, particularly syntactic parsing.In this paper, we conduct an in-depth analysis of LLM parsing capabilities, delving into the underlying causes of why LLMs struggle with this task and the specific shortcomings they exhibit.We find that LLMs may be limited in their ability to fully leverage grammar rules from existing treebanks, restricting their capability to generate syntactic structures.To help LLMs acquire knowledge without additional training, we propose a selfcorrection method that leverages grammar rules from existing treebanks to guide LLMs in correcting previous errors.Specifically, we automatically detect potential errors and dynamically search for relevant rules, offering hints and examples to guide LLMs in making corrections themselves.Experimental results on three datasets using various LLMs demonstrate that our method significantly improves performance in both in-domain and cross-domain settings.
Over the past two decades, quantitative models and statistical methods have been increasingly applied to syntax research, with treebanks emerging as a vital resource for this field. However, tools specifically designed for calculating syntactic metrics remain relatively scarce compared to traditional word frequency measures. This gap hinders the deeper integration of quantitative methods into syntax research. In response, we have developed a tool called QuanSyn, which is tailored for quantitative syntax analysis. Based on treebanks, QuanSyn can calculate metrics such as dependency distance, hierarchical distance, dependency direction, and valency-related syntactic metrics. Additionally, it facilitates the construction of linguistic networks and rapid fitting of functions. QuanSyn can help lower the barriers to accessing syntactic quantitative analysis, thereby promoting the broader application of quantitative methods in syntax research.
This study investigates the relationship between the affective response evoked by hearing upstairs neighbour footsteps in wooden residential buildings, the acoustic characteristics of those footsteps sounds, and the personal traits of participants. Sound recordings were analysed using parameters extracted from the autocorrelation function (ACF) and interaural cross-correlation function (IACF) to identify temporal and spatial features. A laboratory experiment involving 46 individuals assessed their affective responses in terms of arousal and valence after being exposed to a variety of footsteps sounds. The visual simulation of the living room scenario was generated by inviting participants to wear a head-mounted display (HMD) that showed a 360-degree image of a living room with natural or artificial lighting. Participants self-reported their non-acoustic traits before the test including noise sensitivity, attitude towards neighbours, and circadian rhythm type. Results showed significant correlations between affective responses and all acoustic parameters. Pitch-related parameters (ɸ 1 and τ 1 ) and sound pressure level (SPL) were good predictors of arousal and valence. Without SPL, spatial parameters (IACC and τ IACC ) also contributed to affective ratings. Furthermore, participants with low noise sensitivity and a Morning chronotype reported significantly lower arousal and higher valence compared to those with high noise sensitivity and an Evening chronotype. Finally, participants with a positive attitude towards neighbours exhibited higher valence than those with a negative attitude towards neighbours. This study uniquely explores emotional responses to neighbour sounds in lightweight wooden buildings integrating acoustic and non-acoustic factors and using immersive simulations for a more realistic and ecologically valid assessment. • Arousal and valence ratings were significantly correlated with SPL and parameters extracted from the ACF/IACF functions. • Pitch-related parameters and SPL were good predictors of arousal and valence. • Without SPL, spatial parameters also contributed to affective ratings. • Non-acoustic factors significantly impacted affective responses and their correlations with ACF/IACF parameters.
Thermal imaging technology, known for its noncontact and noninvasive nature, offers distinct advantages in computerized emotion sensing. In the literature, a decrease in nose-tip temperature has been associated with dynamic subjective arousal. However, these studies were limited by their focus on a few regions of interest, neglecting a comprehensive analysis of the entire face, and not accounting for the temporal dynamics of thermal changes. To overcome these limitations, we propose an analytical method for facial thermal images using statistical parametric mapping (SPM), which was developed for functional brain image analysis. We developed semiautomated preprocessing protocols to effectively realign and standardize facial thermal images. To validate these analyses, we recorded the thermal images of participants’ faces and assessed dynamic valence and arousal ratings while they observed emotional films. The proposed SPM analyses revealed significant negative associations with dynamic arousal ratings at the nose tip and forehead. The analyses incorporating temporal disparity revealed more forehead clusters than the analyses assuming no delay. These findings validate the proposed pixel-based facial thermal image analysis method using SPM. The results suggest that computerized pixel-based analysis of facial thermal images can be used to estimate dynamic emotional states, with potential applications in various human behavioral fields, including mental health diagnosis and marketing research. • We developed a pixel-based facial thermal image analytical method using SPM. • The method for emotion sensing was tested with an experiment showing emotional films. • Facial thermal images and dynamic valence/arousal ratings were measured. • Negative associations with arousal ratings were detected at the nosetip and forehead. • The data validate the pixel-based facial thermal image analysis for emotion sensing.
Cross-domain constituency parsing is still an unsolved challenge in computational linguistics since the available multi-domain constituency treebank is limited.We investigate automatic treebank generation by large language models (LLMs) in this paper.The performance of LLMs on constituency parsing is poor, therefore we propose a novel treebank generation method, LLM back generation, which is similar to the reverse process of constituency parsing.LLM back generation takes the incomplete cross-domain constituency tree with only domain keyword leaf nodes as input and fills the missing words to generate the cross-domain constituency treebank.Besides, we also introduce a span-level contrastive learning pretraining strategy to make full use of the LLM back generation treebank for cross-domain constituency parsing.We verify the effectiveness of our LLM back generation treebank coupled with contrastive learning pre-training on five target domains of MCTB.Experimental results show that our approach achieves state-of-theart performance on average results compared with various baselines.
One of the fundamental tasks in natural language processing (NLP) is dependency parsing, which involves analyzing the grammatical structure of sentences by establishing relationships or dependencies between words. In this paper, we examine the difficulties and methods for performing dependency parsing for Bangla text, a language with complex morphology and distinctive syntactic properties. The article addresses the value of dependency parsing in capturing the linguistic subtleties of Bangla sentences and their applications in various NLP tasks. A method has been proposed to develop dependency parsing on Bangla text using a graph-based approach. A parsing tree is generated from a directed graph using Bangla input. The proposed system is achieved overall 68% accuracy which is evaluated using Bangla dependency corpus. This method enriches Bangla language linguistic resources and annotated corpora, facilitating the language’s global use. The resultant tree is also evaluated using evaluation metrics.
International audience
International audience
Arguments, unlike adjuncts, are typically understood as verb-specific dependents, which includes the fact that the morphosyntactic devices used for argument encoding are determined by individual verbs. Building on this observation, we operationalize arguments as dependents whose encoding device occurs with a given verb at a significantly higher-than-average frequency. We apply an argument extraction algorithm to a dataset of 132,221 verb dependents from Russian treebanks available in the Universal Dependencies (UD) platform. To evaluate the algorithm ’ s performance, we compare its results to a manually annotated subset, informed by The Active Dictionary and a detailed semantic understanding of argumenthood. The frequency-based algorithm achieves acceptable precision (approx. 0.83), with particularly few false positives, making it a promising tool for cross-linguistic applications in typologically diverse languages with UD treebanks. Theoretically, we argue that a quantitative distributional approach to valency—originally proposed in Ju. D. Apresjan ’ s early pioneering work—broadly aligns with the in-depth semantic analyses of individual verbs and their meanings found in his later works, including The Active Dictionary.
International audience
The drive for internationalization in higher education has accelerated over the past two decades, reshaping how universities teach, collaborate, and move knowledge across borders. Central to this transformation are not only institutional partnerships and student exchange programs, but the more subtle and powerful roles of language, literature, and culture. These elements shape how students understand the world and how universities construct global learning spaces. English has taken on the role of academic lingua franca, simplifying international communication. Yet, its dominance raises complex questions about equity, inclusion, and cultural diversity. At the same time, we’re witnessing rapid developments in artificial intelligence (AI) that promise to radically alter how we think about access, communication, and language learning in higher education. Against this backdrop, the European Higher Education Area (EHEA) has set a clear goal: by 2030, at least 20% of graduates should have studied or trained abroad (European University Association, 2023). To reach this target, we must reconsider whether English alone can carry the weight of internationalization, or whether a more multilingual, AI-supported approach is needed.
UD labels that are used by generated dependency treebank.