Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
entiment analysis is a field of research that uses NLP/Artificial Intelligence techniques to determine thoughts, emotions, and attitudes expressed about an object or event. This means analyzing the tone or mood of the text, feelings and thoughts. Many sentiment analysis studies and linguistic resources have been developed for the English language. However, there are a limited number of studies on the Uzbek language. The reasons for this are insufficient data collection, direct translation of existing English language resources not being compatible with the Uzbek language, differences in the grammatical structure of the two languages, and regional, religious, and mental differences. The study aims to identify emotional-expressiveness units that express emotional evaluation and to develop a linguistic database for emotional analysis. It deals with the issues of developing the linguistic support for sentiment analysis of Uzbek language texts; the study of context-dependent factors in emotional analysis; and determining the best style for the Uzbek language. Attention is paid to the analysis of expressing emotional evaluation means specific to the Uzbek language levels.
• Pedestrians felt more aroused with larger highly automated vehicles (HAVs) than smaller HAVs. • Higher trust ratings were reported for larger HAVs with static eHMI compared to no eHMI, while the opposite was true for smaller HAVs. • Shorter crossing initiation times when pedestrians interacted with HAVs through dynamic eHMI, which signals automation status and yielding intent. • Pedestrians felt safer and trusted HAVs more with dynamic eHMIs compared to static or no eHMIs. • In non-matching conditions, pedestrians showed more negative evaluations of trust, perceived safety, and affective responses than in matching. Highly automated vehicles (HAVs) will soon be introduced into mixed urban traffic. Pedestrians might have an idea of HAVs. Nevertheless, they probably have never interacted with them before. Moreover, pedestrians will not be able to communicate with HAVs like they are used to with manual vehicles. External human–machine interfaces (eHMIs) are possible design solutions for HAVs to ensure safe interaction with other road users. Light-based eHMIs positively affected pedestrians’ trust ratings, perceived safety, and willingness to cross. However, previous studies often neglected the effect of vehicle size, although larger-sized HAVs could be potentially perceived as the more significant threat. Additionally, the relationship between vehicle kinematics and eHMIs for differently sized HAVs remains an underexplored research topic. This study investigated the effects of vehicle size (small vs. large), eHMI state (dynamic eHMI vs. static eHMI vs. no eHMI), and vehicle kinematics (yielding vs. non-yielding) on pedestrian crossing behavior and their subjective assessment. In virtual reality, we created a shared space traffic scenario, in which the eHMI and vehicle kinematics matched or did not match. For yielding conditions, the results showed that participants felt more aroused with larger HAVs than with smaller HAVs. Moreover, pedestrians initiated their crossing significantly earlier when both vehicle sizes had a dynamic eHMI compared to a static eHMI vs. no eHMI. Additionally, pedestrians evaluated a dynamic eHMI with higher trust ratings, higher perceived safety, and more positive affective reactions. The results manifested that the use of dynamic eHMIs can effectively enhance pedestrian-vehicle communication with a large and a small HAV. For non-matching conditions, the participants tended to rely on the vehicle kinematics for both vehicle sizes. Overall, the study highlighted the potential of eHMIs for pedestrian interactions with HAVs of varying sizes when they are well-coordinated with the vehicle kinematics, aiming to enhance safety and efficiency in mixed-traffic environments.
Abstract To explore if translation-intrinsic features are apparent in other types of bilingualism-influenced constrained language use such as non-native production, this study approaches syntactic and typological properties of constrained English translated from Chinese and written by native Chinese speakers via two cognitively-motivated dependency metrics, viz. mean dependency distance (MDD) and dependency direction (DDir). Results of this study show that translated English (both L1 and L2) and non-native English differ from the non-constrained native English in a similar way yet to a slightly different extent, but not from each other in both indicators. Syntactically, bilingually-constrained varieties exhibit reduced syntactic complexity with shorter MDDs, suggesting a simplification tendency. Typologically, cross-linguistic influences are detected in constrained varieties for being more head-final in word-order primed by the source or native language Chinese. Surprisingly, it seems that language directionality affects, albeit marginally, the affinity between constrained varieties, with non-native English being more syntactically and typologically similar to translated English from L1 than from L2.
Revealing the syntactic structure of sentences in Chinese poses significant challenges for word-level parsers due to the absence of clear word boundaries. To facilitate a transition from word-level to character-level Chinese dependency parsing, this paper proposes modeling latent internal structures within words. In this way, each word-level dependency tree is interpreted as a forest of character-level trees. A constrained Eisner algorithm is implemented to ensure the compatibility of character-level trees, guaranteeing a single root for intra-word structures and establishing inter-word dependencies between these roots. Experiments on Chinese treebanks demonstrate the superiority of our method over both the pipeline framework and previous joint models. A detailed analysis reveals that a coarse-to-fine parsing strategy empowers the model to predict more linguistically plausible intra-word structures.
This study investigates the potential of large language models (LLMs) to provide accurate estimates of concreteness, valence, and arousal for multi-word expressions. Unlike previous artificial intelligence (AI) methods, LLMs can capture the nuanced meanings of multi-word expressions. We systematically evaluated GPT-4o's ability to predict concreteness, valence, and arousal. In Study 1, GPT-4o showed strong correlations with human concreteness ratings (r =.8) for multi-word expressions. In Study 2, these findings were repeated for valence and arousal ratings of individual words, matching or outperforming previous AI models. Studies 3-5 extended the valence and arousal analysis to multi-word expressions and showed good validity of the LLM-generated estimates for these stimuli as well. To help researchers with stimulus selection, we provide datasets with LLM-generated norms of concreteness, valence, and arousal for 126,397 English single words and 63,680 multi-word expressions.
The paper explores the grammatical transformations in the paradigm of nouns from the cycle Za Serafimite (On the Seraphim) by St. John Chrysostom. The transformations detected have been compared to manuscripts from the 12th – 14th centuries: the Yagichev Zlatoust, Tarnovo recension of the Verse Prolog, the Synaxaria for the Triodion by Zacchaeus the Philosopher, the Life of St. Antony, the Chronicle of Manasses, Boril’s Synodikon, the Song of Songs from the Rila Monastery, Patriarch Photios’s Letter to Knyaz Boris, Philotheus Kokkinos’s translation of the Panegyric to All Saints. The paper also shows the chronology and frequency of the changes that took place. It explores the grammatical relations and the continuity between a pre-Euthymian translation and the Tarnovo Literary School, and makes a synchronic inquiry into a word class, which indicates the historical development of the standard Bulgarian language and the tendency towards analyticism.
Abstract The management of severe hoarding is often highly challenging due to lack of collaboration and the need to coordinate a large team of professionals. Although numerous strategies have been developed to manage severe hoarding, the most effective approach has not been established. To evaluate and compare three different approaches to the management of severe hoarding in non-voluntary clients. Naturalistic study of clients treated involuntarily by a Crisis Resolution Home Treatment (CRHT) team for severe hoarding. Three management strategies were compared: (1) case management approach with full and part-time staff (HLH), (2) case management approach based on interprofessional networking collaboration (ICN), and (3) routine social service care with non-specific hoarding management led by a social worker (RSW). The Clutter Image Rating scale (CIR) was used to assess hoarding severity at baseline and at 6-, 12-, and 24-months. The main outcome measure was “case resolution” (CIR score < 4). Of the 271 cases referred to the CRHT, 214 completed all follow-up measures. Resolution was achieved in 84.5%, 36.6%, and 36.4% of cases managed by the HLH, RSW, and ICN strategies, respectively ( p < 0.001). The HLH strategy resulted in the greatest improvement in hoarding behaviour. In this study, the most effective strategy to resolve severe hoarding in non-voluntary clients was the case management approach with a full-time team. These findings suggest that centralizing case management in a team of specialized, highly autonomous professionals using a collaborative approach involving motivational interviewing could be the best strategy to resolve severe hoarding.
The influence of prescriptivism on nineteenth-century English has been a moot point since the last century, where two extreme positions are noted.Traditional history of the English language handbooks, on the one hand, often consider prescriptivism as a pervasive phenomenon in the late eighteenth and the nineteenth centuries, because of a belief that the English language was for some time reluctant to variation and change (Finegan 1998: 545; Anderwald 2014: 13-14).Modern approaches, on the other hand, contend that prescriptivism did not have actual effects on the language of the period, with the assumption that language change was effective as a result of continuous developments from below (Tieken-Boon van Ostade 2009: 79).The issue, however, is now taken to be far from this polarization as there is evidence that prescriptivism certainly helped shape nineteenth-century English (and present-day English), even though the influence 'may have been only partially successful, i.e. on only certain text-types, only on written language but not spoken, or only for a certain period, not for others' (Anderwald 2019: 88-89).The nineteenth century was actually the century of prescriptivism as a result of the common awareness of the need for correctness to have access to a higher social scale.This explains the proliferation of grammar books and usage guides throughout the century which, using a normative approach, 'taught readers that social advancement through language was in principle within reach of everyone and could be obtained with the help of proper guidance' (Tieken-Boon van Ostade 2009: 3).The number of such publications increased notably from the beginning of the century, reaching its peak towards the second half of the century (Tieken-Boon van Ostade 2008: 4).Importantly enough, this intense rate of publication derived from the creation of a discourse community of grammarians who contributed to the development of homogeneous norms eventually accepted on the grounds of the continuous copying and re-writing of previous sources (Anderwald 2012: 29; also Watts 1999;2008).The effects of prescriptivism on language have been hitherto overestimated in some publications.Therefore it is necessary to investigate the development of individual features to evaluate, if any, the likely relationship between linguistic norms and the rise or decline of a particular form in Late Modern English.The present issue thus stems from Anderwald's proposal now two decades ago to combine historical grammaticographythe study of historical grammar books of English -with historical corpus linguistics as the only means to confirm 'with some confidence which features of language were subject to prescriptive influence, and where prescriptivists' attempt at changing (or preserving) the language had little or no effect' (2014: 1).Anderwald's desideratum implies that a detailed
The academic identity is in the focus of the paper. The study is relevant as Academic Identity represented in English determines a young scolar`s entering the global academic community, acknowledgement by the dicourse community and science ambition realisation. The paper is aimed at analising lingvo-didactic ponential of revising academic text in English as an efficient methodological technique in enhancing English proficiency of research students; the way revising facilitates Academic Identity construction. The aim involves the following objectives: (1) considering various viewponts of both foreign and Russian scientists on the notion and major components of Academic Identity, our fomulating the definition of Academic Identity; (2) revealing the lingvotextual peculiarities of a scientific paper which influence Academic Identity construction, analising their potential of representing author and his/her academic status; (3) discussing methodological relevance of using English academic text revising as an efficient methodological technique in enhancing English proficiency of research students and developing Academic Identity; (4) hypothesis testing on the co-working academic case-studies of university teachers and research students at BMSTU; (5) interpreting the obtained results in terms of their verification and relevance in reaching the goal. The authors employ empirical techniques of pedagogical observation, quantitative and qualitative (discourse-interpretation) analyses. The novelty and theoretical value of the study are determined by the fact that it has proved the possibility of constructing Academic Identity at the International level as a holistic structural dynamic process of developing identity by mastering academic communication language characterised by its discourse functional peculiarities. It has been established that Academic Identity of research students may be nurtured during writing academic texts by applying structural-linguistic norms of academic writing, author's self-identification by means of linguistic and text markers, revising texts and manipulating reasonable language resources to vary the degree the author and his vewpoint are visible in a scientific paper. The practical value lies in its relevance for successful Academic Identity constructing of young scholars by English resourse, for research students' awareness at technical higher education institutions how revising academic texts influences Academic Identity degree. Therefore, the co-working of university teachers and research students on revising their scientific papers facilitates mastering academic discourse knowledge while enhancing Academic English proficiency, self-confidence and self-consistency, i.e. constructing Academic Identity.
The article deals with color-term lexemes in the modern Russian and Chinese languages. The focus is on a comparative study of the semantics and functioning of the lexemes ‘zelenyi’ (green), ‘绿’ (green, lü), and ‘青’ (a polysemous color name with the meaning of ‘green’, qing) in the Russian and Chinese mass media; their language connotations are investigated. The study employed the descriptive and comparative methods and the method of contextual analysis. The article analyzes the main and implied meanings of green-color terms through the examples of expressions and phrases with a green lexeme found in Russian and Chinese linguistic databases such as the Standardized terminology base for foreign translation of Chinese characteristic discourse, Russian National Corpus, General Internet Corpus of Russian, etc. The author seeks to investigate the multiple meanings of green-color lexemes in the Russian and Chinese languages, notes the expansion of their meaning, especially in the Russian and Chinese mass media. The conducted research shows that the emergence of the new meaning of the green-color lexeme is closely connected with the cultural and social changes in Russia and China. The importance of using linguocultural color lexemes in the language of modern media is emphasized by the fact that the lexeme ‘green’ as a bright language feature performs the function of expressive informing, conveys the national value. The initial analysis of the polysemy and metaphoricity of the Chinese lexeme ‘青’(qing) carried out in the article fills a gap in the study of the use of the polysemous color name ‘青’ (qing) in translation of texts written by the Russian classics into Chinese and contributes to further in-depth lexicosemantic, lexicographical, and translation studies on the translation of this lexeme.
Abstract The stimulant ± 3,4-methylenedioxymethamphetamine (MDMA) has been shown to enhance the perceived pleasantness of touch. However, the underlying neural processes contributing to touch-related effects of MDMA are not well understood. Using a double-blind, randomized, within-subject design, this study used fMRI to examine hemodynamic changes following MDMA (1.5 mg/kg) vs. lactose placebo administration during gentle touch stimulation in a healthy sample (N = 18). Participants were stroked on the forearm at a slower, more pleasant (3 cm/s), and a faster (30 cm/s), less pleasant speed. For the MDMA session, participants’ affective ratings of touch stimulation were higher than their placebo ratings. Increase in plasma oxytocin (OT) levels was also greater during the MDMA session. On the neural level, primary sensorimotor areas showed greater hemodynamic changes during the MDMA than during the placebo session for both touch speeds, indicating a relatively early influence within somatosensory pathways. Changes in OT levels showed an interaction with drug in an occipitotemporal region, area MT+, associated with motion perception. However, posterior insula did not show preferential activation for the slower stroking speed. These initial findings provide a basis for extending our knowledge of the neural processes underlying the effect of MDMA on affective touch.
This study is a survey of the English-language words that are used when speaking about meaning with specific focus on the categories of function, ritual and myth. Such words can be used in interviews, questionnaires, measurement metrics and other forms of ethnography and testing. Understanding why consumers perceive designed artefacts to be personally relevant is a commercial imperative. Previous research has suggested that three categories of meaning are commonly encountered, i.e. function, ritual and myth. They cover a spectrum from the purely instrumental to the purely symbolic. However, despite the logical and philosophical groundwork, there has been little analysis of the actual words and phrases that are in everyday use by people when describing the meanings of designed artefacts. The objectives of the study described here were (1) to identify the words and phrases that are most frequently encountered in everyday language when discussing meaning, (2) to determine for each word or phrase its degree of belonging to the formal categories of function, ritual and myth and (3) to thematically group the words and phrases into macro-components of meaning. Three different analysis were performed. The first was based on the contents of major online dictionaries and thesauri, the second was based on the results from queries of the online lexical database WordNet and the third was based on a corpus analysis approach involving neural network word embedding algorithms. Thematic grouping of the database of extracted words and phrases suggested that in all three cases the macro-components of the concept of ‘function’, ‘ritual’ and ‘myth’ cover a spectrum that can be considered to be from an essential property (‘intention’, ‘ceremonial’ and ‘belief’) to an emergent property (‘action’, ‘spiritual’ and ‘symbolism’). The list of words, phrases and macro-components provides a first empirically established vocabulary of meaning for use in design activity.
Linguistic studies of mass media discourse prove that it presents a complex blend of verbal and visual characteristics. To this end, the methodological procedures should also imbibe structural methods combined with cognitive modelling. These forms of text and its analysis tend to be unjustly given less attention in educational process although teaching media literacy has already become a popular trend. The present paper is aimed at investigating the level of students’ awareness as far as certain semantic and semiotic elements are concerned, and to show the methodological steps, students can further rely on. As one of the sources, frequent in modern British mass media discourse, we have chosen criminal conceptual metaphor applied to Jeremy Corbyn. Thus, we show both the linguistic analysis procedure and the results of the pedagogical experiment. In our linguistic investigation we applied a combination of cognitive and structural linguistic methods. For teaching purposes, an experimental linguistic database of verbal and polycode texts was formed. The procedure described herein can be applied in various pedagogical cases: for teaching students of linguistics, journalism and mass media communication studies, pedagogics, politics, cross-cultural and linguocultural studies. The results of the analysis contribute to the investigations of political metaphors and broaden our understanding of polycode mass media texts.
In this paper, we present a Universal Dependencies (UD) treebank for the Standard Albanian Language (SAL), annotated by expert linguistics supported by information technology professionals. The annotated treebank consists of 24,537 tokens (1,400 sentences) and includes annotation for syntactic dependencies, part-ofspeech tags, morphological features, and lemmas. This treebank represents the largest UD treebank available for SAL. In order to overcome annotation challenges in SAL within the UD framework, we delicately balanced the preservation of the richness of SAL grammar while adapting the UD tagset and addressing unique language-specific features for a unified annotation. We discuss the criteria followed to select the sentences included in the treebank and address the most significant linguistic considerations when adapting the UD framework conform to the grammar of the SAL. Our efforts contribute to the advancement of linguistic analyses and Natural Language Processing (NLP) in the SAL. The treebank will be made available online under an open license so that to provide the possibility for further developments of NLP tools based on the Artificial Intelligence (AI) models for the Albanian language.
Phrygian-KUL is a treebank of the ancient Phrygian language for Universal Dependencies (UD). Having originally only annotated the New Phrygian subcorpus, this dataset is continuously being updated to include the entire epigraphic corpus. For more information, please visit the relevant page at the UD project site or the repository on Github.
We introduce a detailed annotation scheme for argument structure constructions (ASCs) along with a manually annotated ASC treebank.This treebank encompasses 10,204 sentences from both first (5,936) and second language English datasets (1,948 for written; 2,320 for spoken).We detail the annotation process and evaluate inter-annotation agreement for overall and each ASC category.
In this article, the beta version 0.1.0 of Opera Graeca Adnotata (OGA), the largest open-access multilayer corpus for Ancient Greek (AG) is presented. OGA consists of 1,687 literary works and 34M+ tokens coming from the PerseusDL and OpenGreekAndLatin GitHub repositories, which host AG texts ranging from about 800 BCE to about 250 CE. The texts have been enriched with seven annotation layers: (i) tokenization layer; (ii) sentence segmentation layer; (iii) lemmatization layer; (iv) morphological layer; (v) dependency layer; (vi) dependency function layer; (vii) Canonical Text Services (CTS) citation layer. The creation of each layer is described by highlighting the main technical and annotation-related issues encountered. Tokenization, sentence segmentation, and CTS citation are performed by rule-based algorithms, while morphosyntactic annotation is the output of the COMBO parser trained on the data of the Ancient Greek Dependency Treebank. For the sake of scalability and reusability, the corpus is released in the standoff formats PAULA XML and its offspring LAULA XML.
Tolosa Treebank for Occitan Tolosa Treebank is the first dependency treebank for Occitan, developed as part of the EFA 227/16 LINGUATEC Project, financed by the POCTEFA Interreg European funds. The current version of the treebank contains 25K tokens annotated for PoS tags, lemmas and syntactic dependencies. Linguistic annotation follows Universal Dependencies guidelines (https://universaldependencies.org/#language-u). The corpus files are stored in the ConLL-U format. Each sentence is preceded by a sentence ID and the original, non-tokenized text of the sentence. The annotation is provided in a column-based format defined as follows: 1. ID: Word index, integer starting at 1 for each new sentence; may be a range for multiword tokens.2. FORM: Word form or punctuation symbol.3. LEMMA: Lemma or stem of word form.4. UPOS: Universal part-of-speech tag.5. XPOS: Language-specific part-of-speech tag; underscore if not available.6. FEATS: List of morphological features from the universal feature inventory or from a defined language-specific extension; underscore if not available.7. HEAD: Head of the current word, which is either a value of ID or zero (0).8. DEPREL: Universal dependency relation to the HEAD9. DEPS: Enhanced dependency graph in the form of a list of head-deprel pairs.10. MISC: Any other annotation. The corpus is built out of data from four major Occitan dialects: Gascon, Lengadocian, Lemosin and Provençau. The corpus content is stratified by dialect. File naming conventions are indicated in the readme file. AUTHORS: Aleksandra Miletic, Myriam Bras, Marianne Vergez Couret, Jean Sibille, Louise Esher, Clamença Poujade LICENSE: The corpus is distributed under the Creative Commons BY-NC-SA 4.0 license (https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en).REFERENCING: If you use this corpus in your work, please cite the following paper:Aleksandra Miletic, Myriam Bras, Marianne Vergez-Couret, Louise Esher, Clamença Poujade, and Jean Sibille. 2020. A Four-Dialect Treebank for Occitan: Building Process and Parsing Experiments. In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects, pages 140–149, Barcelona, Spain (Online). International Committee on Computational Linguistics (ICCL).
This paper explored the cultural and linguistic aspects of health exchanges between healthcare practitioners and patients among the Jukuns of Wukari in Nigeria, within the health centres in the town. It focused on patient-healthcare-provider dynamics and found out how language and culture influenced healthcare communication within formal settings. Integrating ethnographic, sociolinguistic, and anthropological approaches, the study unveiled how language and culture impacted interactions and health-seeking behaviours in these centres. It revealed the roles of language and culture in understanding health information, healthcare provider-patient exchanges, and treatment adherence within the distinct sociolinguistic context of the Jukun. Using such qualitative techniques as interviews and observations in the health centres, the study captured the intricate verbal and nonverbal communication, specific cultural discourse patterns, and communication strategies used by patients and healthcare practitioners. Findings highlighted diverse cultural and linguistic methods employed by Jukuns, such as using proverbs, ironies, metaphors, and nonverbal cues, to express themselves in healthcare settings. The research showed that these methods could facilitate communication with familiar practitioners but might complicate interactions with those from different ethnic backgrounds. Ultimately, it offered crucial perspectives for refining healthcare provision, aligning with the precise linguistic and cultural contexts of the Jukun community within formal healthcare settings in Wukari and other parts of Jukunland. Based on the foregoing, the researchers recommended that health practitioners should make use of interpreters and familiarise themselves with the cultural and linguistic norms of their immediate communities for effective health discourse that would enhance quality healthcare delivery.
This article presents a set of standardised corpora of poetry comprising over 330,000 poems in ten languages (Czech, English, French, German, Hungarian, Italian, Portuguese, Russian, Slovenian, and Spanish). Each corpus has been deduplicated, enriched with Universal Dependencies, provided with additional metadata, and converted into a unified json structure.
The present project endeavors to enrich the linguistic resources available for Italian by constructing a Universal Dependencies treebank for the KIParla corpus (Mauri et al., 2019, Ballarè et al., 2020), an existing and well known resource for spoken Italian.
Public discourse often excludes the erroneous speech of migratory subjects, thus foreclosing social and political rapprochement. This paper argues that the position outside of pregiven, linguistic norms provides migratory speech with an improvisatory quality that can serve as a catalyst for community formation. Simulating freestyle forms in writing, Feridun Zaimoğlu’s Kanak Sprak seeks to find a new language for the critique of xenophobia and to establish belonging based on precarious conditions. In a close reading of Fikret’s monologue “Pity is that true vitamin,” I show how improvisation disrupts established discourses and transforms the meaning of conventional hate speech tropes to forge transethnic alliances. The paper then turns to the subsequent volume Koppstoff and problematizes the commodification of Kanak speech in neoliberal pop culture. Çağıl’s monologue “If you’re smart, you take our side” hints at a different understanding of improvisation that reframes the relation between mainstream society and its others. Drawing on critical improvisation studies, the paper contributes to the understanding of linguistic interventions into social orders that determine who can say what, in which speech form, and according to which norms of belonging.
In this paper we present PARSEME-AR, the first openly available Arabic corpus manually annotated for Verbal Multiword Expressions (VMWEs). The annotation process is carried out based on guidelines put forward by PARSEME, a multilingual project for more than 26 languages. The corpus contains 4749 VMWEs in about 7500 sentences taken from the Prague Arabic Dependency Treebank. The results notably show a high degree of discontinuity in Arabic VMWEs in comparison to other languages in the PARSEME suite. We also propose analyses of interesting and challenging phenomena encountered during the annotation process. Moreover, we offer the first benchmark for the VMWE identification task in Arabic, by training two state-of-the-art systems, on our Arabic data.
Phrygian-KUL is a treebank of the ancient Phrygian language for Universal Dependencies (UD). Having originally only annotated the New Phrygian subcorpus, this dataset is continuously being updated to include the entire epigraphic corpus. For more information, please visit the relevant page at the UD project site or the repository on Github.
The development of formal models of decision making under risk has been shaped largelyby decisions between options with monetary outcomes. The most prominentmodel—cumulative prospect theory (CPT)—is good at describing choices betweenmonetary lotteries, but performs less well with nonmonetary and nonnumerical outcomes(e.g., medications with possible side effects). We suggest that affective processes, which arenot considered in CPT, play a larger role in nonmonetary than in monetary choices, andpropose two psychologically motivated modifications to CPT’s modeling framework tocapture these differences: (a) using affect ratings rather than monetary equivalents torepresent the subjective value of nonmonetary outcomes (affective valuation); and (b)allowing the probability weighting of an outcome to depend on the amount of affecttriggered in a choice problem (affective probability weighting). We compared model variantsof CPT implementing the proposed modifications in four empirical datasets (totalN = 240). For choices between options with negative nonmonetary outcomes (medicationswith possible side effects), these modifications substantially improved model performancerelative to the standard implementation of CPT. The same did not hold for monetarychoices. Further, an eye-tracking study on nonmonetary choice (N = 68) provided evidencefor two key behavioral and cognitive predictions of affective probability weighting—namely,that risk aversion increases and attention to probability information decreases as theaffective value of the worst outcome in a choice problem increases. Our work integratesprevious ideas on how affect guides and modulates preference construction within acomputational model and delineates an important context in which these mechanismsapply.
Discourse relations play a pivotal role in establishing coherence within textual content, uniting sentences and clauses into a cohesive narrative.The Penn Discourse Treebank (PDTB) stands as one of the most extensively utilized datasets in this domain.In PDTB-3 (Webber et al., 2019), the annotators can assign multiple labels to an example, when they believe that multiple relations are present.Prior research in discourse relation recognition has treated these instances as separate examples during training, and only one example needs to have its label predicted correctly for the instance to be judged as correct.However, this approach is inadequate, as it fails to account for the interdependence of labels in real-world contexts and to distinguish between cases where only one sense relation holds and cases where multiple relations hold simultaneously.In our work, we address this challenge by exploring various multi-label classification frameworks to handle implicit discourse relation recognition.We show that multi-label classification methods don't depress performance for single-label prediction.Additionally, we give comprehensive analysis of results and data.Our work contributes to advancing the understanding and application of discourse relations and provide a foundation for the future study.
Phrygian-KUL is a treebank of the ancient Phrygian language for Universal Dependencies (UD). Having originally only annotated the New Phrygian subcorpus, this dataset is continuously being updated to include the entire epigraphic corpus. For more information, please visit the relevant page at the UD project site or the repository on Github.
Skin conductance response (SCR) serves as a dependable marker of sympathetic activation used to measure emotional arousal. This study investigates the impact of presentation modality (face or word) on the degree of emotional discrimination elicited by SCR. Facial expressions or words associated with six basic emotions-anger, happiness, disgust, fear, sadness, and surprise-were studied among 102 participants. The amplitude of SCR was accurately predicted by subjective arousal ratings of these stimuli, but not by valence ratings. The habituation process to emotional and neutral stimuli across six successive presentations was characterized by an exponential decay function, capturing the rate at which SCR response diminishes in relation to the preceding trial of the same stimulus. Through the subtraction of the response to neutral stimuli from the emotion-evoked SCR, it was demonstrated that the initial presentation of each emotion elicits a substantial response, particularly attributable to the emotional content. Notably, the initial emotional response to faces expressing happiness, disgust, and sadness surpassed that of words conveying the same emotions. The results indicate that different emotional responses can be quantified using a simple electrical instrument.
Muitas das línguas indígenas brasileiras estão ameaçadas de extinção. Na maioria dos casos, estratégias de revitalização e de conservação dessas línguas são imprescindíveis (Crystal, 2002; Harrison, 2007), necessitando de processos contínuos de promoção de políticas linguísticas e de ações voltadas à educação escolar indígena. Este artigo apresenta o uso de ferramentas linguísticas associadas à construção de treebanks (corpus de textos com anotações sintáticas e morfológicas) e à descrição de duas línguas minoritárias do tronco linguístico Tupí faladas no sudoeste Amazônico. Os treebanks, parte das Dependências Universais (De Marneffe et al., 2021; Duran et al. 2022), são a base de algumas das atividades do projeto “Educação, Linguística, História e Comunidades Indígenas” vinculado ao Programa Institucional de Bolsas de Iniciação Científica (2021-2022) da Universidade Federal da Paraíba (UFPB). Neste artigo, discutimos a aplicação dessas ferramentas na descrição linguística e exploramos a interseção da linguística computacional com a linguística descritiva.
The present research focuses on analyzing the mechanisms of politeness and the role of adjectives in interpersonal relationships, considering them as important tools for expressing attitudes, emotions, and cultural values. Based on the research results, it has been established that politeness strategies and the use of adjectives vary significantly depending on cultural contexts, underscoring the need for a deeper understanding of linguistic norms and variations. Special attention has been paid to the role of adjectives, which serve not only as descriptive elements but also as means of expressing evaluations and emotions, thus, playing a significant role in intercultural interaction. The research conclusions underscore the importance of integrating intercultural understanding into the process of linguistic interaction. It has been revealed that successful intercultural communication requires not only language proficiency but also a deep understanding of cultural differences in politeness strategies. The research also points to the need for further exploration in this field, in particular, in developing practical recommendations for enhancing intercultural communication. In this context, knowledge and application of relevant linguistic strategies can contribute to better understanding and respect among representatives of different cultures.
Computer lexicography is one of the important directions of modern domestic linguistics and translation studies.Nowadays, scientists face important questions related to the theoretical and practical aspects of compiling computer dictionaries, which, undoubtedly, have significant scientific significance.A necessary stage in solving these questions is to understand the peculiarities of the formation of this section of linguistic science -its preconditions, methodological base; directions of the scientific research.The article is devoted to highlighting some aspects of the historical development of domestic and foreign lexicography.The task of the article is to consider the main stages of the development of computer technologies for compilation of dictionaries and to determine the prerequisites that led to the emergence of such a direction in linguistics as computer lexicography.The advent of computers actively influenced the development of lexicography.Initially, they were used to prepare paper dictionaries, in other words, they served as a typewriter.But later it turned out that computers can perform such functions as editing, storing any lexicographic information, and therefore computer corpora of texts appeared, and then machine-readable dictionaries.emergence of linguistic databases, electronic libraries and card libraries.Automated lexicographic databases in the form of electronic dictionaries are now an integral part of systems of machine translation, information search, editing and correction of texts, as well as processing of large text arrays and their storage as a separate task of creating electronic libraries.Computer dictionaries on optical media enabled translators and scientists to quickly find any information about a word (translation, interpretation, etc.).
Natural Language Toolkit (NLTK) is a comprehensive Python library designed to facilitate the exploration, analysis, and processing of human language data. With its extensive collection of tools, NLTK provides researchers, developers, and educators with a powerful platform for tasks ranging from basic text processing to advanced natural language understanding and machine learning. The toolkit includes modules for tokenization, stemming, lemmatization, part-of-speech tagging, named entity recognition, syntactic parsing, semantic analysis, and more. Furthermore, NLTK offers access to numerous linguistic resources such as corpora, lexicons, and treebanks, making it an invaluable resource for both learning and research in the field of natural language processing (NLP). NLTK serves as an indispensable tool for unlocking the complexities of human language
In everyday life, music is increasingly being listened to through headphones and mobile devices in public situations. While a large body of research has demonstrated that music may influence the emotional states of listeners and affect multimodal perceptions i.e. in films, less is known about the music’s impact on environments and the interpretation of social situations. We conducted an online experiment to investigate the influence of music on evaluations, considering individuals’ emotional states (emotion congruence) and group perception. Participants were randomly assigned to one of three experimental conditions (music with positive valence and high arousal, music with negative valence and low arousal, and no music) while viewing images of two different social group types that varied in perceived group characteristics (group members being familiar or unfamiliar with each other). Images were rated on four bipolar scales measuring affective quality and cognitive evaluation of social situations. Results show that individuals who listened to negative music provided lower valence ratings and also judgded social environments lower in terms of pleasantness and cheerfulness (affective) than individuals in the other experimental conditions. In contrast, ratings of crowdedness and familiarity (cognitive) did not differ between experimental conditions. The effect of music on affective evaluations was shaped by social group types, such that participants were more influenced by music when viewing intimacy groups (e.g., friends) than when viewing transitory groups (e.g., strangers). Overall, our results support the assumption of mood congruency for affective evaluations and emphasize the need to consider social information when studying the influence of music on the perception of environments.
Abstract Classical models of tool knowledge and use are centred on dorsal and ventral parietal pathways. Theories of semantic cognition implicate a “hub-and-spoke” network, centred on the anterior temporal lobe (ATL), that underpins all concepts including tools. Despite their prominence, the two theoretical frameworks have never been brought together and the large discrepancy in the functional neuroanatomy addressed. We undertook a multiple-regression Representational Similarity Analysis (RSA) of task fMRI data with four (motor action, broad function, mechanical function, object structure) feature-based models. The motor action model correlated with the activation patterns in bilateral superior parietal lobules (SPL), while the models of broad function and mechanical effect aligned with the activation patterns in bilateral ATLs. The object-structure model correlated with activation patterns in bilateral middle occipital gyri. The results also showed that the ventral ATL activation patterns corresponded simultaneously with all RDM models except object structure. Furthermore, a standard univariate analysis using tool-familiarity ratings for parametric modulation revealed that classical tool-network regions (frontal, inferior parietal, and posterior middle temporal cortices) were increasingly active as the tool familiarity reduced. These results demonstrate that parietal and ATL regions are both crucial and motivate a major extension and revision of the neuroanatomical framework for tool use. Significance Statement This study provides definitive evidence for convergent tool representation in human anterior temporal lobe (ATL), outside the traditionally focused parietal lobe as the critical centre for human tool-use ability. For many years the parietal lobe was considered crucial in recognizing and planning use of familiar objects, while regions in temporal lobe received little attention. Our advanced multi-voxel analysis with artefact-resistive fMRI scanning revealed that both non-motor (tool-function) and motor (kinematics for tool use) information convergently represented in the left ventral ATL, while showing other distributed regions encoding distinct types of tool information in anterior temporal and parietal regions. These findings highlight the ATL’s crucial role in tool representation and necessitate a significant expansion of neuroanatomical framework for human tool-use ability.
Direct dependency parsing of the speech signal -- as opposed to parsing speech transcriptions -- has recently been proposed as a task (Pupier et al. 2022), as a way of incorporating prosodic information in the parsing system and bypassing the limitations of a pipeline approach that would consist of using first an Automatic Speech Recognition (ASR) system and then a syntactic parser. In this article, we report on a set of experiments aiming at assessing the performance of two parsing paradigms (graph-based parsing and sequence labeling based parsing) on speech parsing. We perform this evaluation on a large treebank of spoken French, featuring realistic spontaneous conversations. Our findings show that (i) the graph based approach obtain better results across the board (ii) parsing directly from speech outperforms a pipeline approach, despite having 30% fewer parameters.
Abstract Human creativity originates from brain cortical networks that are specialized in idea generation, processing, and evaluation. The concurrent verbalization of our inner thoughts during the execution of a design task enables the use of dynamic semantic networks as a tool for investigating, evaluating, and monitoring creative thought. The primary advantage of using lexical databases such as WordNet for reproducible information-theoretic quantification of convergence or divergence of design ideas in creative problem solving is the simultaneous handling of both words and meanings, which enables interpretation of the constructed dynamic semantic networks in terms of underlying functionally active brain cortical regions involved in concept comprehension and production. In this study, the quantitative dynamics of semantic measures computed with a moving time window is investigated empirically in the DTRS10 dataset with design review conversations and detected divergent thinking is shown to predict success of design ideas. Thus, dynamic semantic networks present an opportunity for real-time computer-assisted detection of critical events during creative problem solving, with the goal of employing this knowledge to artificially augment human creativity.
Social decision-making is known to be influenced by predictive emotions or the perceived reciprocity of partners. However, the connection between emotion, decision-making, and contextual reciprocity remains less understood. Moreover, arguments suggest that emotional experiences within a social context can be better conceptualised as prosocial rather than basic emotions, necessitating the inclusion of two social dimensions: focus, the degree of an emotion's relevance to oneself or others, and dominance, the degree to which one feels in control of an emotion. For better representation, these dimensions should be considered alongside the interoceptive dimensions of valence and arousal. In an ultimatum game involving fair, moderate, and unfair offers, this online study measured the emotions of 476 participants using a multidimensional affective rating scale. Using unsupervised classification algorithms, we identified individual differences in decisions and emotional experiences. Certain individuals exhibited consistent levels of acceptance behaviours and emotions, while reciprocal individuals' acceptance behaviours and emotions followed external reward value structures. Furthermore, individuals with distinct emotional responses to partners exhibited unique economic responses to their emotions, with only the reciprocal group exhibiting sensitivity to dominance prediction errors. The study illustrates a context-specific model capable of subtyping populations engaged in social interaction and exhibiting heterogeneous mental states.
несмотря на лингвистическое разнообразие, наблюдаемое в Тунисе, французский язык сохраняет значительную роль в жизни страны. В рамках данной статьи исследуются локальные преобразования французской лексики в тунисском варианте языка. Целью исследования является выявление и анализ этих преобразований, а также определение влияющих на них лингвистических, социальных и историко-культурных факторов. Исследование использует интегрированный подход, сочетающий лингвистический, социолингвистический и историко-культурный анализ. На основе корпуса текстов, написанных на тунисском варианте французского языка, были выделены два основных типа лексических преобразований. Результаты исследования показывают, что эти преобразования являются результатом взаимодействия французского языка с местной культурой, социальными нормами и историческими событиями. В частности, отмечено влияние местных языков. Исконно французские слова приобретают новые этнические коннотации и адаптируются к местным реалиям, отражая уникальный характер тунисского варианта французского языка. Данное исследование вносит вклад в понимание многоязычных ситуаций и территориальной вариативности языка. Оно демонстрирует важность местных социокультурных факторов в формировании уникальных вариантов языка, что имеет значение для межкультурной коммуникации и лингвистического образования. despite the linguistic diversity observed in Tunisia, French retains a prominent role in the nation's life. This article examines the local transformations of French vocabulary within Tunisian French. The study aims to identify and analyze these transformations, as well as to ascertain the linguistic, social, and historical-cultural factors that influence them. The study employs an integrated approach, combining linguistic, sociolinguistic, and historical-cultural analysis. Based on a corpus of texts written in Tunisian French, two main types of lexical transformations were identified. The results of the study show that these transformations are the result of the interaction between French and the local culture, social norms, and historical events. In particular, the influence of local languages is noted. Originally French words acquire new ethnic connotations and adapt to local realities, reflecting the unique character of Tunisian French. This research contributes to the understanding of multilingual situations and territorial variation of language. It demonstrates the importance of local sociocultural factors in the formation of unique language variants, which has implications for intercultural communication and language education.
This paper identifies a micro-cue correlating to verb second word order (V2) in two closely related Medieval Romance languages. As V2 is asymmetrically distributed in main rather than subordinate clauses, an asymmetry would be expected in phenomena assumed to relate to V2, such as subject inversion, null subject and enclisis. The loss of that asymmetry should therefore indicate the loss of the V2 word order rule. These assumptions are tested here by a quantitative analysis of a treebank of calibrated data covering the crucial period of change (from the 14th to the 16th century) for Medieval French and Venetian. The hard quantitative evidence provided demonstrates that the main versus embedded asymmetry is indeed a micro-cue of V2 structure, and of its loss in one of the two investigated languages.
Quantization is one of the efficient model compression methods, which represents the network with fixed-point or low-bit numbers. Existing quantization methods address the network quantization by treating it as a single-objective optimization that pursues high accuracy (performance optimization) while keeping the quantization constraint. However, owing to the non-differentiability of the quantization operation, it is challenging to integrate the quantization operation into the network training and achieve optimal parameters. In this paper, a novel multi-objective convex quantization for efficient model compression is proposed. Specifically, the network training is modeled as a multi-objective optimization to find the network with both high precision and low quantization error (actually, these two goals are somewhat contradictory and affect each other). To achieve effective multi-objective optimization, this paper designs a quantization error function that is differentiable and ensures the computation convexity in each period, so as to avoid the non-differentiable back-propagation of the quantization operation. Then, we perform a time-series self-distillation training scheme on the multi-objective optimization framework, which distills its past softened labels and combines the hard targets to guarantee controllable and stable performance convergence during training. At last and more importantly, a new dynamic Lagrangian coefficient adaption is designed to adjust the gradient magnitude of quantization loss and performance loss and balance the two losses during training processing. The proposed method is evaluated on well-known benchmarks: MNIST, CIFAR-10/100, ImageNet, Penn Treebank and Microsoft COCO, and experimental results show that the proposed method achieves outstanding performance compared to existing methods.
Purpose We explored whether (1) an informational intervention improves ratings of individuals on the autism spectrum (IotAS) in a job interview by curbing salience bias and whether expert-based influence amplifies this effect (Study 1); (2) the effect of disclosure of autism on ratings depends on a candidate’s presentation as IotAS or neurotypical (Studies 1 and 2) and (3) social desirability bias affects ratings of and emotional responses to disclosers (Study 2). Design/methodology/approach In two studies, participants, randomly assigned to experimental conditions, watched a mock job interview of a candidate presenting as an IotAS or neurotypical and reported their perception of his job suitability and selection decision. Study 2 additionally measured participants’ traits associated with social desirability bias, self-reported emotions and involuntary emotions gauged via face-reading software. Findings In Study 1, the informational intervention improved ratings of the IotAS-presenting candidate; delivery by an expert made no difference. Disclosure increased ratings of both the IotAS-presenting and neurotypical-presenting candidates, especially the former, and information mattered more in the absence of disclosure. In Study 2, disclosure improved ratings of the IotAS-presenting candidate only; no evidence of social desirability bias emerged. Originality/value We explain that an informational intervention works by attenuating salience bias, focusing raters on IotAS' qualifications rather than on their unexpected behavior. We also show that disclosure is less helpful for IotAS who behave more neuronormatively and social desirability bias affects neither ratings of nor emotional responses to IotAS-presenting job candidates.
Natural language processing for Greek and Latin, inflectional languages with small corpora, requires special techniques.For morphological tagging, transformer models show promising potential, but the best approach to use these models is unclear.For both languages, this paper examines the impact of using morphological lexica, training different model types (a single model with a combined feature tag, multiple models for separate features, and a multi-task model for all features), and adding linguistic constraints.We find that, although simply fine-tuning transformers to predict a monolithic tag may already yield decent results, each of these adaptations can further improve tagging accuracy.1 For example, for each type (unique word form) in the GUM English Universal Dependencies Treebank (see https://universaldependencies.org/) there are 10.7 tokens.For the Latin PROIEL treebank there are only 6.5, and for the Greek Perseus treebank even less, viz.4.8 (note that they are all roughly similar in size: 212K, 205K and 202K tokens respectively).
<p>Listening to music often leads to physiological responses. Do these physiological responses contain sufficient information to infer emotion induced in the listener? The current study explores this question by attempting to predict judgments of “felt” emotion from physiological responses alone using linear and neural network models. We measured five channels of peripheral physiology from 20 participants—heart rate (HR), respiration, galvanic skin response, and activity in corrugator supercilii and zygomaticus major facial muscles. Using valence and arousal (VA) dimensions, participants rated their felt emotion after listening to each of 12 classical music excerpts. After extracting features from the five channels, we examined their correlation with VA ratings, and then performed multiple linear regression to see if a linear relationship between the physiological responses could account for the ratings. Although linear models predicted a significant amount of variance in arousal ratings, they were unable to do so with valence ratings. We then used a neural network to provide a non-linear account of the ratings. The network was trained on the mean ratings of eight of the 12 excerpts and tested on the remainder. Performance of the neural network confirms that physiological responses alone can be used to predict musically induced emotion. The non-linear model derived from the neural network was more accurate than linear models derived from multiple linear regression, particularly along the valence dimension. A secondary analysis allowed us to quantify the relative contributions of inputs to the non-linear model. The study represents a novel approach to understanding the complex relationship between physiological responses and musically induced emotion.</p>
Background: Despite the frequent comorbidity of affective and addictive disorders, the significance of affective dysregulation in problematic pornography use (PPU) is commonly disregarded. The objective of this study is to investigate whether individuals with PPU demonstrate increased sensitivity to negative emotional stimuli in comparison to healthy controls (HCs). Methods: Electrophysiological responses were captured via event-related potentials (ERPs) from 27 individuals with PPU and 29 HCs. They completed an oddball task involving the presentation of deviant stimuli in the form of highly negative (HN), moderately negative (MN), and neutral images, with a standard stimulus being a neutral kettle image. To evaluate participants' subjective feelings of valence and arousal, the Self-Assessment Manikin (SAM) was employed. Results: Regarding subjective evaluations, individuals with PPU indicated diminished valence ratings for HN images as opposed to HCs. Concerning electrophysiological assessments, those with PPU manifested elevated N2 amplitudes in response to both HN and MN images when contrasted against neutral images. Additionally, PPU participants displayed an intensified P3 response to HN images in contrast to MN images, a distinction not evident within the HCs. Discussion: These outcomes suggest that individuals with PPU exhibited heightened reactivity toward negative stimuli. This increased sensitivity to negative cues could potentially play a role in the propensity of PPU individuals to resort to pornography as a coping mechanism for managing stress regulation.
Conventional continuous emotion prediction systems are typically trained to predict the ‘average’ of affect ratings obtained from multiple human annotators. These systems, however, ignore the ambiguity inherent in the perceived emotions, which is not captured by the ‘average rating’. This paper presents a novel ambiguity-aware continuous emotion prediction system that predicts the time-varying emotion state as a series of beta distributions. Our recent work has shown beta distributions to be an effective parametric model of a collection of affect ratings. This work develops an appropriate cost function that enables neural networks to be trained to predict beta distributions. It also investigates the choice of parameterization of the beta distribution, the choice of activation functions of the output layer, and the tractability of gradient definitions in combination with the loss function. The proposed framework is implemented using a Bag-of-Audio-Words front-end and an LSTM-based back-end and evaluated on the RECOLA dataset. In addition to comparison with baseline systems that only predict the ‘average rating’, the effectiveness with which the predictions represent ambiguity in perceived emotions is also evaluated. Experimental results reveal that the proposed approach outperforms other ambiguity-aware systems, especially when predicting valence.
Semantic change is a universal phenomenon in human language. This article analyzes the trends of semantic change and the possible reasons behind them from a lexical perspective. By illustrating the five different types of changes that happen on the lexicon, the study finds that factors like language contacts, cultural trends and social norms significantly impact semantic changes. Besides, people’s cognitive mode and their concept of collocations also contribute to the changes in the lexicon.