Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Nowadays, tree-structured deep learning classifier models have been widely used in different applications to ensure effective feature representation and learning. Amongst, dimensional sentiment analysis is the most interactive research field, which intends to identify continuous numerical values in the valence-arousal (VA) space. To achieve this, a tree-structured regional convolutional neural network with long short-term memory (T-CNN-LSTM) model was developed, which predicts the VA ratings of the texts for sentiment analysis. In contrast, the effect of a low prediction rate and difficulty of feature learning in a small number of class samples was not analyzed. Hence, this manuscript proposes an adversarial T-CNN-LSTM (A-T-CNN-LSTM) model for predicting the VA to achieve more fine-grained sentiment analysis. This model develops a semantic-enabled frequency-aware generative adversarial network (SFGAN) to produce more adversarial samples using the generator network and decrease the spectral data loss of the discriminator. It embeds the frequency-aware categorizer (FAC) into the discriminator to determine the input veracity in the spatial and spectral domains. Besides, semantic restricted sampling is employed in SFGAN for synthesizing the image subject to a semantic mask. Further, the created samples are classified by the T-CNN-LSTM for predicting the VA scores of sentences. Finally, the experimental results exhibit that the A-T-CNN-LSTM on stanford sentiment Treebank (SST) and CIFAR-10 databases achieves 90.12% and 91% accuracy than the other tree-structured CNNs.
The principle of DEPENDENCY LENGTH MINIMIZATION, which seeks to keep syntactically related words close in a sentence, is thought to universally shape the structure of human languages for effective communication. However, the extent to which dependency length minimization is applied in human language systems is not yet fully understood. Preverbally, the placement of long-before-short constituents and postverbally, short-before-long constituents are known to minimize overall dependency length of a sentence. In this study, we test the hypothesis that placing only the shortest preverbal constituent next to the main-verb explains word order preferences in Hindi (a SOV language) as opposed to the global minimization of dependency length. We characterize this approach as a least-effort strategy because it is a cost-effective way to shorten all dependencies between the verb and its preverbal dependencies. As such, this approach is consistent with the bounded-rationality perspective according to which decision making is governed by "fast but frugal" heuristics rather than by a search for optimal solutions. Consistent with this idea, our results indicate that actual corpus sentences in the Hindi-Urdu Treebank corpus are better explained by the least effort strategy than by global minimization of dependency lengths. Additionally, for the task of distinguishing corpus sentences from counterfactual variants, we find that the dependency length and constituent length of the constituent closest to the main verb are much better predictors of whether a sentence appeared in the corpus than total dependency length. Overall, our findings suggest that cognitive resource constraints play a crucial role in shaping natural languages.
BACKGROUND: Cognitive behavioral therapy (CBT) is a moderately efficacious treatment for hoarding disorder (HD), with most individuals remaining symptomatic after treatment. The Joining Forces Trial will evaluate whether 10 weeks of in-home decluttering can significantly augment the outcomes of group CBT. METHODS: A randomized controlled trial of in-home decluttering augmentation of group CBT for HD. Adult participants with HD (N = 90) will receive 12 weeks of protocol-based group CBT for HD. After group CBT, participants will be randomized to either 10 weeks of in-home decluttering led by a social services team or a waitlist. The primary endpoint is 10 weeks post-randomization. The primary outcome measures are the self-reported Saving Inventory-Revised and the blind assessor-rated Clutter Image Rating. Participants on the waitlist will cross over to receive the in-home decluttering intervention after the primary endpoint. Data will be analyzed according to intention-to-treat principles. We will also evaluate the cost-effectiveness of this intervention from both healthcare and societal perspectives. DISCUSSION: HD is challenging to treat with conventional psychological treatments. We hypothesize that in-home decluttering sessions carried out by personnel in social services will be an efficacious and cost-effective augmentation strategy of group CBT for HD. Recruitment started in January 2021, and the final participant is expected to reach the primary endpoint in December 2024. TRAIL REGISTRATION: ClinicalTrials.gov NCT04712474. Registered on 15 January 2021.
OBJECTIVE: Hoarding behaviour is a common but poorly characterised problem in real-world clinical practice. Although hoarding behaviour is the key component of Hoarding Disorder (HD), there are people who exhibit hoarding behaviour but do not suffer from HD. The aim of the present study was to characterise a clinical sample of patients with clinically relevant hoarding behaviour and evaluate the differential characteristics between patients with and without HD. METHODS: This study included patients who received treatment at the home visitation program in Barcelona (Spain) from January 2013 through December 2020, and scored ≥ 4 on the Clutter Image Rating scale. Sociodemographic, DSM-5 diagnosis, clinical data and differences between patients with and without an HD diagnosis were assessed. RESULTS: A total of 243 subjects were included. Hoarding behaviour had been unnoticed in its early stages and the median length in the sample was 10 years (IQR 15). 100% of the cases had hoarding-related complications. HD was the most common diagnosis in 117 patients (48.1%). CONCLUSIONS: The study found several differential characteristics between patients with and without HD diagnosis. Alcohol use disorder could play an important role among those without HD diagnosis. Home visitation programs could improve earlier detection, preventing hoarding-related complications.
The strongest formulations of grounded cognition assume that perceptual intuitions about concepts involve the re-activation of sensorimotor experience we have made with their referents in the world. Within this framework, concreteness and imageability ratings are indeed of crucial importance by operationalising the amount of perceptual interaction we have made with objects. Here we tested such an assumption by asking whether visual intuitions about concepts are provided accurately even when direct visual experience is absent. To this aim, we considered concreteness and imageability intuitions in blind people and tested whether these judgments are predicted by Image-based Frequency (IF, i.e. a data-driven estimate approximating the availability of the word referent in the visual environment). Results indicated that IF predicts perceptual intuitions with a larger extent in sighted compared to blind individuals, thus suggesting a role of direct experience in shaping our judgements. However, the effect of IF was significant not only in sighted but also in blind individuals. This indicates that having direct visual experience with objects does not play a critical role in making them concrete and imageable in a person’s intuitions: people do not need visual experience to develop intuition about the availability of things in the external visual environment and use this intuition to inform concreteness/imageability judgments. Our findings fit closely the idea that perceptual judgments are the outcome of introspection/abstraction tasks invoking high-level conceptual knowledge that is not necessarily acquired via direct perceptual experience.
This article in addition to introducing and defining taboos, examines the existing and well-known taboos in the collection of short stories” To whom should I greet” by the famous Iranian novelist SiminDaneshvar, during which, it refers to the use of taboo words by the characters of gender(female/male) in social situations. The authors have tried to include social behaviors such as: good/bad, holy/ unholy, polite/ impolite, etc. in social relations and dialogues of the work, based on the accepted norms of the Iranian society in front of the readers. In adition to be able to explain the relationship between culture and language and the intractions between the two in terms of prohibition in the sociology of language based on gender differences and how they are used. Besides, to prove that in this work, Simin while paying attention to the values of the Persian society, has consciously tried to break the linguistic norms in the daily individual and social life of the characters of some of stories in many cases. This study also tried to show the results obtained from the frequency of using taboos and inconveniences in the actions and speech of speakers by gender in the form of a graph at the end of the article and between gender and taboos in terms of their use in speakers there was a significant relationship.
In this paper, we propose a method for removing linguistic information from speech for the purpose of isolating paralinguistic indicators of affect. The immediate utility of this method lies in clinical tests of sensitivity to vocal affect that are not confounded by language, which is impaired in a variety of clinical populations. The method is based on simultaneous recordings of speech audio and electroglotto-graphic (EGG) signals. The speech audio signal is used to estimate the average vocal tract filter response and amplitude envelop. The EGG signal supplies a direct correlate of voice source activity that is mostly independent of phonetic articulation. These signals are used to create a third signal designed to capture as much paralinguistic information from the vocal production system as possible-maximizing the retention of bioacoustic cues to affect-while eliminating phonetic cues to verbal meaning. To evaluate the success of this method, we studied the perception of corresponding speech audio and transformed EGG signals in an affect rating experiment with online listeners. The results show a high degree of similarity in the perceived affect of matched signals, indicating that our method is effective.
An analysis of robots (simulators) in education is provided. Promising directions for their development are highlighted, such as realism, interactivity, adaptation and personalization. The features of using simulators in dentistry are considered. The main disadvantages of existing simulators in dentistry have been identified, namely the lack of a communicative component and imitation of patient behavior. The anthropomorphic dental simulator is based on the Robo-C robot, which is a unique combination of advanced technologies and human facial expressions, which allows it to communicate with people, reproduce movements of different parts of the body and express emotions. As dental components, the following components were created and implemented into the Robo-C control system: a Smart jaw, including cameras and a temperature sensor, and a Smart tooth, including a pressure sensor. The Robo-C control system has been upgraded taking into account the Smart jaw and Smart tooth, which made it possible to connect the dental treatment process with the robot’s servos through its linguistic base. The process of analyzing data obtained from Smart jaw cameras using a neural network is described. A two-stage classification scheme for dental defects has been proposed and its effectiveness has been proven. The linguistic base contains a set of rules with the help of which devices (microphone, speakers, servos, Smart jaw, Smart tooth) interact with each other. An example of compiling a linguistic database rule is given. The linguistic base, Smart-jaw and Smart-tooth are configured for one of four cases: caries treatment, tooth preparation for a crown, tooth extraction, endodontic treatment. Treatment quality control is carried out using a comprehensive assessment of communication interaction with the robot and analysis of Smart-jaw and Smart-tooth data. An example of work in one of the cases is given. The anthropomorphic dental simulator presented in the article allows the use of new technologies in the training of dentists, as well as the simulation of various dental procedures, which will significantly improve the practical preparation of students for working with patients.
In the investigation of musical features that influence musical affect, timbre has received relatively little attention. We studied the acoustic properties describing the timbral qualities of sound and analyzed how they predict perceived and induced affect. First, we considered the timbre of single tones played by different instruments by re-analyzing two previously published studies on perceived affect by Eerola et al. (2012, Mus. Percept.) and McAdams et al. (2017, Front. Psychol.) and comparing them to Experiment 1 from Korsmit et al. (2023, submitted), which investigated both perceived and induced affect on valence, tension arousal, and energy arousal ratings. For all datasets, positive valence and decreased tension were predicted by an increasingly prominent fundamental frequency. In experiments with pitch variation, energy arousal was predicted by increased pitch and decreased inharmonicity. In experiments with variations in playing technique, energy arousal was predicted by a faster attack or less sustain. Second, we compared Experiment 1 results on dimensional affect (valence, tension, energy) to results on discrete affect (anger, fear, sadness, happiness, tenderness). Like valence and tension, angrier and more fearful sounds and less happy and tender sounds showed a less prominent fundamental frequency. Happiness and tenderness had a shorter perceived duration, and sad sounds were more sustained. Third, Experiment 2 from Korsmit et al. (2023) tested the affective response to chromatic scales. As with single tones, energy was predicted by an increase in pitch and decrease in inharmonicity. However, dimensional and discrete affect were most frequently predicted by median inharmonicity and spectral spread range. The synthesis of multiple datasets revealed consistent findings, but also discrepancies that may be explained by differences in stimulus selection. Furthermore, although findings on perceived and induced affect were largely similar, some findings on discrete affect and chromatic scales were not revealed with dimensional affect and single tones.
This paper provides a comparative analysis of the patterns of formation and qualities of a modern linear text and an Internet text. The article is the result of the study of Internet text stylistics, based mainly on Russian-language texts from Russia and Ukraine. The paper considers the process of formation of a new type of text - the Internet text. Being essentially different from the classical linear text, the Internet text does not lend itself well to the description based on the classical text theory. Thus, the Internet text is not complete, vectorial, not united by the completeness of thought expression. An important characteristic of the Internet text is its interactivity, which means that the roles of addressee and addressee are constantly changing. In addition, it is difficult, or even impossible, to define the boundaries of the Internet text due to its hypertextuality, which has become habitual intertextuality. All the above-mentioned aspects make up the pragmatics of the Internet text as a subspecies of the media text. Another crucial problem addressed in the paper is the study of the regularities of Internet communication in general and the stylistics of the Internet text. In the course of the research it became obvious that speech aggression and violation of norms of speech culture are stylistic dominants of online communication. This influenced the formation of other stylistic dominants such as, hate speech, fake, hype, clickbait, etc. Internet style is clearly characterized by being provocative, aggressive, hostile. The problems of bullying, humiliation of human dignity, invective and obscenity are actively studied from the standpoint of linguoecology, because, according to most researchers, the constant neglect of communicative and linguistic norms leads to the degradation of the national language style and literary norms. Despite the fact that there is still a division between public and interpersonal online communication, i.e. formal and informal, the problems of speech behavior of Internet users are becoming increasingly relevant.
This chapter synthesizes the most prominent Natural Language Processing (NLP) studies conducted on Persian, focusing on text processing. The first section contains selected tasks from the NLP pipeline, such as text preprocessing, tokenization, POS tagging, syntactic parsing, treebank annotation or semantic analysis along with examples of how researchers approached the problem for Persian and, where applicable, examples of tools developed to perform given tasks. The following section discusses the application of Persian NLP like spell-checking, information retrieval, machine translation or sentiment analysis. Finally, the last section summarizes the Persian NLP corpora and other resources.
The onset of the COVID-19 pandemic accentuated the need for access to biomedical literature to answer timely and disease-specific questions. During the early days of the pandemic, one of the biggest challenges we faced was the lack of peer-reviewed biomedical articles on COVID-19 that could be used to train machine learning models for question answering (QA). In this paper, we explore the roles weak supervision and data augmentation play in training deep neural network QA models. First, we investigate whether labels generated automatically from the structured abstracts of scholarly papers using an information retrieval algorithm, BM25, provide a weak supervision signal to train an extractive QA model. We also curate new QA pairs using information retrieval techniques, guided by the clinicaltrials.gov schema and the structured abstracts of articles, in the absence of annotated data from biomedical domain experts. Furthermore, we explore augmenting the training data of a deep neural network model with linguistic features from external sources such as lexical databases to account for variations in word morphology and meaning. To better utilize our training data, we apply curriculum learning to domain adaptation, fine-tuning our QA model in stages based on characteristics of the QA pairs. We evaluate our methods in the context of QA models at the core of a system to answer questions about COVID-19.
The purpose of the article is to consider the morphological peculiarities of the system of the nouns in New Bulgarian translation of the “Catechismos” written by Theodore the Studite, which is a part of the manuscripts no. 1/154 kept in Odessa National Scientific Library. The subject of the research is the morphological specifics of nouns in Odessa copy of the “Catechismos” dating from the 18th. The morphological peculiarities of nouns is considered in the context of the formation of a linguistic norm, which allowed the combination with different intensity of linguistic means of several language systems functioning at the time (traditional Middle Bulgarian written language, Church Slavonic Eastern recensions, and vernacular language form). The analysis proposed in this paper presents the extensive system of cases, which does not reflect the real vernacular Bulgarian speech in the 18th; the specifics of the functioning of the gramemes of the case paradigm of masculine, feminine and neuter nouns in the singular and plural forms is analyzed. Usage case endings mistakes, which indicates their artificial nature, are considered. The lack of article of nominal parts of speech is noted; the predominance of compound declension form of the adjectives and participles over short forms is revealed; the relatively high frequency of use of active present participles is registered. The results of the study make it possible to outline some probable factors that determine the writer’s preference for using the linguistic tools of the so-called “bookish”, “traditional”, “archaic” writing systems. An another reason which to some extent explains he usage of case inflections in the text of this relatively late stage of the historical development of the Bulgarian language might be the use of East Slavic copies of the Studite’s sermons by the scriber. Еhe comparison of “Catechismos” copies of South and East Slavic origin is necessary for verification of this assumption, in which we see prospects for further research.
We present an approach for assessing how multilingual large language models (LLMs) learn syntax in terms of multi-formalism syntactic structures. We aim to recover constituent and dependency structures by casting parsing as sequence labeling. To do so, we select a few LLMs and study them on 13 diverse UD treebanks for dependency parsing and 10 treebanks for constituent parsing. Our results show that: (i) the framework is consistent across encodings, (ii) pre-trained word vectors do not favor constituency representations of syntax over dependencies, (iii) sub-word tokenization is needed to represent syntax, in contrast to character-based models, and (iv) occurrence of a language in the pretraining data is more important than the amount of task data when recovering syntax from the word vectors.
The modern linguistic studies confirm that keeping pace with modern linguistic norms requires careful usage of terminology in order to be properly employed through which the translation process is promoted and the problem of translating the specialized terminology in general and particularly economic terms are among the problematic encountered in transferring knowledge from foreign languages to Uzbek language. The purpose of this study is to analyze these problems in order to propose solutions for the unification of the term in English, refuting the reasons for the unification of the English terminology, because the terms are the keys to science, so if there are multiple terms equivalent to a single one, this leads to a disturbance in understanding and reflects negatively on the assimilation of knowledge, and contributes to the confusion of the entire translation process.
Decoding Part of Speech(POS) tagging directly from electroencephalography (EEG) signals whilst user overtly spoke (voiced speech) sentences could improve direct speech brain-computer interfaces (BCIs) using imagined or inner speech. To the best of our knowledge, earlier work uses machine learning approach using 74,953 sentences/tokens recorded in 75 EEG sessions. The tokens can be found in 4,479 phrases consisting of terms from the English Online treebank which contains the record of weblogs, newsgroups, reviews, and Yahoo Answers. The results demonstrated the feasibility of POS decoding from EEG based on word class, word frequency, and word length with accuracy of 71%, 86%, 89%, respectively. We believe that there is significant room for improvement with more advanced artificial intelligence. In this paper, we further extend the existing work with end-to-end transformers. Our results presents transformer model outperforms benchmark traditional ML results with +20% in length, +13% for the open vs closed class and +12% in frequency. In our empirical analysis, we find the decoding performance was better when using multi-electrode recordings as compared to single-electrode recordings.
Previous studies have made great advances in RST discourse parsing through specific neural frameworks or features, but they usually split the parsing process into two subtasks and heavily depended on gold discourse segmentation. In this article, we introduce an end-to-end method for sentence-level RST discourse parsing via transforming it into a text-to-text generation task, which can also be simply applied to document-level parsing. Our method unifies the traditional two-stage parsing and generates the parsing tree directly from the input text through our constrained decoding and postprocessing algorithms, without requiring a complicated model. Moreover, the discourse segmentation can be simultaneously generated and extracted from the parsing tree. Experimental results on the RST Discourse Treebank demonstrate that our proposed method outperforms existing methods in both the tasks of discourse parsing and segmentation. We further carry out ablation studies and more targeted comparisons with traditional patterns to analyze our method in more detail. Considering the lack of annotated data in RST parsing, we also create high-quality augmented data and implement self-training, which further improves the performance of our method.
Russian constructicon is an open-access linguistic database containing detailed descriptions of over 3,800 Russian grammatical constructions. In this paper we present a new, enlarged and updated version of Russian Constructicon (RusCxn) as well as new trajectories of development which were opened for the resource after the update. Since its first release, RusCxn, has undergone many significant changes. Our team has expanded the number of constructions present in the database 1,5 times, introduced new meta-information features such as glosses, significantly reworked the architecture and the design of Russian Constructicon’s website, and improved the search facilities. The above-mentioned changes not only make RusCxn more attractive and convenient-to-use, but they can also greatly facilitate typological research in the field of Construction Grammar and improve the mapping between constructicography-orinented resources for different languages.
This corpus-based study investigates the distributions of Korean multiple anaphors, with respect to their morphological types and discourse-pragmatic properties in Long-Distance (LD)-binding. The study is based on the theory of Long-Distance Anaphors (LDAs), such as ‘form-function correlation’ argument (Cole, Hermon, & Sung, 1990; Reinhart & Reuland, 1993; Reuland, 2011, 2017), as well as exempt binding and logophoricity (Sells, 1987; Pollard & Sag, 1992). Based on Sejong Treebank (Parsed Corpus), six hundred sentences containing various Korean anaphors went through manual coding of 5 linguistic factors related to LD-binding: locality, discourse, exempt, logophoric, and logophoric roles. The encoded sentences in distinct conditions were analyzed in terms of frequency by using Chi-square tests. The overall results demonstrated the following: 1) Korean anaphors did not show form-function correlations in terms of binding type and morphological form. 2) Korean anaphors can occur even in syntactically non-exempt positions. 3) Logophoricity conditions were found with the LD-antecedents of the anaphors. The results seem to support discourse-pragmatic analysis with Korean anaphors.
The circumplex model posits a circular representation of affect and some personality traits. There is an increasing need to examine the viability of the circumplex model with multivariate time series data collected on the same individuals due to the development of new data collection methods such as smartphone applications and wearable sensors. Estimating the circumplex model with time series data is more complex than with cross-sectional data because scores at nearby time points tend to be correlated. We adapt Browne’s circumplex model to accommodate time series data. We illustrate the proposed method with an empirical data set of daily affect ratings of an individual over 70 days. We conducted a simulation study to explore the statistical properties of the proposed method. The results show that the method provides more satisfactory confidence intervals and test statistics than a method that treats time series data as if they were cross-sectional data.
In this study, we collected affective ratings of emotional valence and arousal for 882 Serbian words and compared their values at three points in time: before the onset of the COVID-19 pandemic (2018), during the COVID-19 lockdown (2020) and after the government measures were abandoned (2022). Although valence ratings were more stable than arousal ratings, we did not observe a significant change in either valence or arousal ratings across the time points. A more detailed look into the data revealed the change in arousal that was different across the valence values. Our analyses demonstrated that, upon the onset of the COVID-19 pandemic, emotionally negative words elicited higher arousal ratings, whereas emotionally positive words elicited lower arousal ratings. It revealed that our participants became more sensitive to the negative content and less sensitive to the positive content. We hypothesized that this pattern could be linked to reduced resilience and consequently could represent a mental health risk. (This article is published in Applied Psycholinguistics. Popović Stijačić, M., Mišić, K., & Filipović Đurđević, D. (2023). Flattening the curve: COVID-19 induced a decrease in arousal for positive and an increase in arousal for negative words. Applied Psycholinguistics, 44(6), 1069–1089. doi:10.1017/S0142716423000425)
Embodied cognition research identifies mechanisms by which our cognitive activity is connected to body experiences. This approach encompasses not only experimental manipulations but also the quantification of variables related to group and individual differences, i.e., participant-related variables. Moreover, stimuli-related characteristics, such as sensorimotor word ratings, can either be used for the selection of experimental materials or can be the main output of a study themselves. This quantitative information about individuals or stimuli can be collected through non-experimental methods, such as questionnaires and cognitive tests. This chapter gives an overview of questionnaires and cognitive tests often used in embodied cognition research. A questionnaire is a list of questions asking participants to provide information on certain aspects, such as their sociodemographic or medical status. A test is a series of tasks which participants perform for further evaluation by researchers, such as tests of mathematical ability, reading speed, or counting direction. Rating studies collect subjective evaluations of various parameters, typically for large sets of items. The present chapter is divided into two main sections: Participant-related variables and stimuli-related characteristics. We present examples from cognitive linguistics, psycholinguistics, psychophysics, as well as from research on numerical cognition, peripersonal space, and attitudes towards social robots.
PURPOSE: Slow speech rate and abnormal temporal prosody are primary diagnostic criteria for differentiating between people with aphasia who do and do not have apraxia of speech. We sought to identify appropriate cutoff values for abnormal word syllable duration (WSD) in a word repetition task, interpret them relative to a data set of people with chronic aphasia, and evaluate the extent to which manually derived measures could be approximated through an automated process that relied on commercial speech recognition technology. METHOD: Fifty neurotypical participants produced 49 multisyllabic words during a repetition task. Audio recordings were submitted to an automated speech recognition (ASR) service (IBM Watson) to measure word duration and generate an orthographic transcription. The transcribed words were compared to a lexical database, and the number of syllables was identified. Automatic and manual measures were compared for 50% of the sample. Results were interpreted relative to WSD scores from an existing data set of 195 people with mostly chronic aphasia. RESULTS: ASR correctly identified 83% of target words and 98% of target syllable counts. Automated word duration calculations were longer than manual measures due to imprecise cursor placement. Upon applying regression coefficients to the automated measures and examining the frequency distributions for both manual and estimated measures, a WSD of 303-316 ms was found to indicate longer-than-normal performance (corresponding to the 95th percentile). With this cutoff, 40%-45% of participants with aphasia in our comparison sample had an abnormally long WSD. CONCLUSIONS: We recommend using a rounded WSD cutoff score between 303 and 316 ms for manual measures. Future research will focus on customizing automated WSD methods to speech samples from people with aphasia, identifying target words that maximize production and measurement reliability, and developing WSD standard scores based on a large participant sample with and without aphasia.
This paper attempts to examine 'World Englishes' (WE) with connectivity to English as an International Language (EIL), Applied Linguistics and socio-linguistics. In the light of Kachru's model of English Language in the late 20th century. This model has three circles, inner circle, where English is used as native language, Outer Circle, mostly former colonies of British Empire, such as Singapore, India, Kenya, Ghana, Malaysia, Pakistan and others, and 3rd is Expanding Circle, include countries in which English is known as Foreign Language in schools and universities, mostly for communication and business or economic purposes as well with Inner and Outer circles. The term "English language" refers to various interesting and notable features, patterns, or aspects of the English language. These phenomena can encompass a wide range of linguistic phenomena, including grammar, vocabulary, pronunciation, syntax, idioms, and more. English holds significant importance around the world because English is the most widely spoken language globally. It serves as a common language of communication among people from different linguistic backgrounds. Proficiency in English enables individuals to connect with a broader range of people, both in personal and professional contexts. English is the language of international business and economics as well. It facilitates global trade, negotiations, and collaboration between companies and individuals from different countries. Proficiency in English enhances employability and career opportunities, particularly in multinational corporations and industries with international reach. It recognizes the importance of both native and non-native varieties of English and acknowledges that each circle has its own linguistic norms, purposes, and language development. The study informs us that Kachru was an original thinker not in the field of English Language including applied linguistics, multilingualism, bilingualism, language policy, language creativity, code mixing, code switching, cross-cultural communication, sociolinguistics but also in the domain of politics of language and so many other issues including cross-cultural awareness.
The goal of this contribution is to present The Digital Rosetta Stone, which is a project developed at Leipzig University by the Chair of Digital Humanities and the Egyptological Institute/Egyptian Museum Georg Steindorff in collaboration with the British Museum and the Digital Epigraphy and Archaeology Project at the University of Florida. The aims of the project are to produce a collaborative digital edition of the Rosetta Stone, address standardization and customization issues for the scholarly community, create data that can be used by students to understand the language and content of the document, and produce a high-resolution 3D model of the stone. First, the three versions of the text were transcribed and encoded in XML according to the EpiDoc guidelines. Next, the versions were aligned with the Ugarit iAligner tool that supports the alignment of ancient texts with modern languages, such as English and German. All three texts were then parsed syntactically and morphologically through Treebank annotation. Finally, the project explored new 3D-digitization techniques of the Rosetta Stone in the British Museum in order to enhance traditional archaeological methods and facilitate the study of the artifact. The results of this work were used in different courses in Digital Humanities, Digital Philology, and Egyptology.
User ratings are widely used in web systems and applications to provide personalized interaction and to help other users make better choices. Previous research has shown that rating scale features and user personality can both influence users' rating behaviour, but relatively little work has been devoted to understanding if the effects of rating scale features may vary depending on users' personality. In this paper, we study the impact of scale granularity and colour on the ratings of individuals with different personalities, represented according to the Big Five model. To this aim, we carried out a user study with 203 participants, in the context of a web-based survey where users were assigned an image rating task. Our results confirm that both colour and granularity can affect user ratings, but their specific effects also depend on user scores for certain personality traits, in particular agreeableness, openness to experience and conscientiousness.
Fear overgeneralization and perceived uncertainty about future outcomes have been suggested as risk factors for clinical anxiety. However, little is known regarding how they influence each other. In this study, we investigated whether different levels of threat uncertainty influence fear generalization. Three groups of healthy participants underwent a differential fear conditioning protocol followed by a generalization test. All groups learned to associate one female face (conditioned stimulus, CS+) with a female scream (unconditioned stimulus, US) while the other face (CS-) was not associated with the scream. In order to manipulate threat uncertainty, one group (low uncertainty, n = 26) received 80%, the second group (moderate uncertainty, n = 32) received 60%, and the third group (high uncertainty, n = 30) 40% CS-US contingency. In the generalization test, all groups saw CS+ and CS- again as well as four morphs that varied in similarity with the CS+ in steps of 20%. Subjective (expectancy, valence, and arousal ratings), psychophysiological (skin conductance response, SCR), and visuocortical (steady-state visual evoked potentials, ssVEPs) indices of fear were registered. Participants expected the US in accordance with their reinforcement schedules but displayed stronger skin conductance with more uncertainty. However, acquisition of conditioned fear was not evident in ssVEPs. During the generalization test, we found no effect of threat uncertainty in any of the measured variables, but the strength of generalization for threat expectancy ratings was positively correlated with dispositional intolerance of uncertainty. This study suggests that mere threat uncertainty does not modulate fear generalization.
- output-{ciep,treebanks}-full.csv: frequency and entropy for all the categories, using four types of combinations of layers;<br> - plots.R: R script to draw plots from the output files;<br> - readReport-{CIEP+,treebanks}.R: R script to extract frequency and compute entropy from the report files (not included);<br> - ud-wordorder.py: Python script to extract word order pairs from conllu files and write them in report files. Unfortunately, I cannot include the report files, as CIEP+ is protected by copyright; the analysis can be however replicated with respect to the UD Treebanks.
Introduction A study was conducted to investigate if an individual’s trust in law enforcement affects their perception of the emotional facial expressions displayed by police officers. Methods The study invited 77 participants to rate the valence of 360 face images. Images featured individuals without headgear (condition 1), or with a baseball cap (condition 2) or police hat (condition 3) digitally added to the original photograph. The images were balanced across sex, race/ethnicity (Asian, African American, Latine, and Caucasian), and facial expression (Happy, Neutral, and Angry). After rating the facial expressions, respondents completed a survey about their attitudes toward the police. Results The results showed that, on average, valence ratings for “Angry” faces were similar across all experimental conditions. However, a closer examination revealed that faces with police hats were perceived as angrier compared to the control conditions (those with no hat and those with a baseball cap) by individuals who held negative views of the police. Conversely, participants with positive attitudes toward the police perceived faces with police hats as less angry compared to the control condition. This correlation was highly significant for angry faces ( p &lt; 0.01), and stronger in response to male faces compared to female faces but was not significant for neutral or happy faces. Discussion The study emphasizes the substantial role of attitudes in shaping social perception, particularly within the context of law enforcement.
В статье анализируется историческая динамика политической корректности, ее положительные и отрицательные стороны, а также прослеживается ее связь с вежливостью, которая заключается в неиспользовании конфронтационных, ликоповреждающих коммуникативных стратегий при указании на гендер, расу, этнос, возраст, физическое состояние и социальное положение адресата. Исследуются факторы, затрудняющие формирование политкорректного русского языкового сознания: 1) не вполне сформировавшееся понятие ПК применительно к отечественному социальному контексту, осложненное его концептуализацией через призму западного восприятия; 2) отсутствие правовых механизмов ПК, несмотря на недопустимость дискриминации, закрепленную в Конституции РФ; 3) неразработанность языковых основ применения ПК в российском публичном дискурсе; 4) англоцентричность правил ПК для международного общения. Сделан вывод о необходимости выработки российских норм политической корректности с участием широкого лингвистического сообщества. The paper explores the historical dynamics of political correctness (PC), its positive and negative aspects, as well as its connection with politeness as an avoidance of confrontational, face-threatening strategies in reference to the interlocutor’s gender, race, ethnic background, age, physical condition and social status. The study also deals with the factors hindering the formation of the Russian PC awareness, which include: 1) the incompletely formed notion of political correctness in the Russian social context complicated by its conceptualization through the prism of Western comprehension; 2) absence of legal PC mechanisms, in spite of non-discrimination enshrined in the RF Constitution; 3) underdeveloped linguistic norms of political correctness in Russian public discourse; 4) anglocentrism of PC rules in intercultural communication. The article concludes by proposing a wide discussion of Russian political correctness norms involving a wider linguistic community.
Abstract We propose a Slovak language model for the spaCy library in Python. These models are easy-to-use for basic natural language processing tasks in a single package. The package contains several components for basic preprocessing tasks, such as tokenization, sentence boundary detection, syntactic parsing, lemmatization, named entity recognition, morphology analysis, and word vectors. It is based on the state-of-the-art monolingual SlovakBERT model. Named entity recognition is trained on a separate, publicly available WikiAnn database. The other statistical classifiers use a Slovak Dependency Treebank corpus. Morphological tags are compatible with the conventions of the Slovak National Corpus. The part of speech tags use conventions of the Universal Dependencies framework. We trained a separate word vector model on a web-based corpus. The training uses fastText with Floret modification. We present a series of experiments that confirm that the model performs similarly to other languages for all tasks. Training scripts and data are publicly available.
Numerous studies have been conducted on the interpretation and translation of English terms into other languages. The purpose of this study was to identify the adequate Indonesian equivalent terminology for hotel amenities, services, and facilities applied in English and the strategies utilized by both domestic and international hotel guests in understanding the equivalent terms in their native language. Qualitative research methodology was used. The subjects included 10 domestic guests from a 5-star hotel, 10 domestic guests from a 4-star hotel, 5 international guests from a 3-star hotel, and 2 hotel staff from a 5-star hotel, 3 staff from a 4-star hotel, and 1 staff from a 3-star hotel. The findings demonstrated that some of the English terms commonly used in hotels had Indonesian equivalents, and some did not. The international guests strategies were: 1) searching in an online dictionary or a Google search; 2) asking people they met nearby immediately; and 3) guessing the meaning. Domestic guests’ strategies included: (a) asking other guests or hotel staff for clarification; and (b) guessing the meaning. Future research should overcome the limitations of this study, considering translations and linguistic norms training strategies.
The Wall Street Journal section of the Penn Treebank has been the de-facto standard for evaluating POS taggers for a long time, and accuracies over 97\% have been reported. However, less is known about out-of-domain tagger performance, especially with fine-grained label sets. Using data from Elder Scrolls Fandom, a wiki about the \textit{Elder Scrolls} video game universe, we create a modest dataset for qualitatively evaluating the cross-domain performance of two POS taggers: the Stanford tagger (Toutanova et al. 2003) and Bilty (Plank et al. 2016), both trained on WSJ. Our analyses show that performance on tokens seen during training is almost as good as in-domain performance, but accuracy on unknown tokens decreases from 90.37% to 78.37% (Stanford) and 87.84\% to 80.41\% (Bilty) across domains. Both taggers struggle with proper nouns and inconsistent capitalization.
The paper traces the dynamics of the interpretation of the grammatical nature of the vocative in Ukrainian grammars from the 16th century until the present. The subject of the analysis is the content and presentation of this category in two sections of Ukrainian grammar books: (1) morphological, which clarifies the status of the vocative in the inflectional paradigm of the noun, and (2) syntactic, in which the means of expressing address are characterized. Based on the findings of the research, various trends in the description of the vocative in different historical periods have been identified, in particular: (1) until the beginning of the 20th, it was unequivocally qualified as an equal member of the inflectional paradigm of the noun, equal to other cases; (2) from the beginning of the 20th century to 1933 was a period of competition between two theories (the vocative is a case the same as others or the vocative is not a true case, but a “special” form in the inflectional paradigm of the noun); (3) the canonization of the “fake case” status theory; (4) from 1991 to the present there has been an unanimity of authors in qualifying the vocative as a case. Comparing the stages of fundamental changes in the scientific definition of the vocative in grammars with defining events in the history of Ukraine provides the basis for discussions about the influence of socio–political factors on the representation of linguistic theories and the codification of the linguistic norms.
Abstract Stress, anxiety, and depressive symptoms can be reduced by listening to music, but the underlying mechanisms remain unclear. To address this gap, we measured brain connectivity while participants listened to songs of different genres: ambient, pop, and metal. Additionally, affective ratings were obtained while participants ( n = 30) listened to the six different songs, and subjective ratings of state anxiety were solicited at the terminus of each song. Electroencephalography (EEG) connectivity combining weighted Phase Lag Index and graph theory was utilised to document brain activity during listening. Repeated-measures ANOVA indicated that listening to more pleasant and less arousing songs was associated with lower self-reported state anxiety levels than songs rated unpleasant and highly arousing. Of interest, EEG alpha connectivity differed across two ambient songs, particularly in the frontal lobes, despite being from the same genre and rated as highly pleasant and low in arousal. We also observed a sex effect on EEG results, where female participants ( n = 18) displayed stronger connectivity than male participants ( n = 12). Combined, these results suggest that ambient songs reduce state anxiety but have divergent brain responses, possibly reflecting the complex nature of music listening, including sensory processing, emotion and cognition.
Abstract The present study analyzes the transformation of the vowel system and especially the process of vowel mergers based on the Latin inscriptions of the Gallic and Germanic provinces. With the help of the Computerized Historical Linguistic Database of the Latin Inscriptions of the Imperial Age ( http://lldb.elte.hu/ ), it tries to draw and then compare the phonological profiles of the selected provinces and to describe the dialectal position of Gaul and the Germanic provinces regarding vocalism in three periods (AD 1–300, 301–500 and 501–700). The analysis, which also covers comparisons with certain provinces of Italy, Spain and Dalmatia, is carried out considering four aspects: the ratio of vocalic versus consonantal changes, the ratio of vowel mergers compared to vocalic changes, the ratio of e-i and o-u mergers compared to each other, and the ratio of vowel mergers by stressed and unstressed syllable. As a result of the present study, it was revealed that Gallic provinces cannot be treated as a unit or as clearly separate from the other areas studied according to either aspect of the study, especially not in the early, pre-Christian period. Gallic provinces appear to behave in the same or a levelled manner at most in the later and/or latest periods. The Germanic provinces, especially Germania Superior, have, albeit with some delay, adapted to the Gallic provinces in their late development. The present study, which continued József Herman's research, managed to explore the hitherto little-known linguistic and dialectological features of Latin in the Gallic and Germanic provinces.
In this paper, we present a grammar-based natural language framework for robot programming, specifically for pick-and-place tasks. Our approach uses a custom dictionary of action words, designed to store together words that share meaning, allowing for easy expansion of the vocabulary by adding more action words from a lexical database. We validate our Natural Language Robot Programming (NLRP) framework through simulation and real-world experimentation, using a Franka Panda robotic arm equipped with a calibrated camera-in-hand and a microphone. Participants were asked to complete a pick-and-place task using verbal commands, which were converted into text using Google's Speech-to-Text API and processed through the NLRP framework to obtain joint space trajectories for the robot. Our results indicate that our approach has a high system usability score. The framework's dictionary can be easily extended without relying on transfer learning or large data sets. In the future, we plan to compare the presented framework with different approaches of human-assisted pick-and-place tasks via a comprehensive user study.
Hoarding disorder is characterised by the acquisition of, and failure to discard large numbers of items regardless of their actual value, a perceived need to save the items and distress associated with discarding them, significant clutter in living spaces that render the activities associated with those spaces very difficult causing significant distress or impairment in functioning. To aid development of an intervention for hoarding disorder we aimed to identify current practice by investigating key stakeholders existing practice regarding identification, assessment and intervention associated with people with hoarding disorder. Two focus groups with a purposive sample of 17 (eight male, nine female) stakeholders representing a range of services from housing, health, and social care were audio recorded, transcribed verbatim and analysed thematically. There was a lack of consensus regarding how hoarding disorder was understood and of the number of cases of hoarding disorder however all stakeholders agreed hoarding disorder appeared to be increasing. The clutter image rating scale was most used to identify people who needed help for hoarding disorder, in addition to other assessments relevant to the stakeholder. People with hoarding disorder were commonly identified in social housing where regular access to property was required. Stakeholders reported that symptoms of hoarding disorder were often tackled by enforced cleaning, eviction, or other legal action however these approaches were extremely traumatic for the person with hoarding disorder and failed to address the root cause of the disorder. While stakeholders reported there was no established services or treatment pathways specifically for people with hoarding disorder, stakeholders were unanimous in their support for a multi-agency approach. The absence of an established multiagency service that would offer an appropriate and effective pathway when working with a hoarding disorder presentation led stakeholders to work together to suggest a psychology led multiagency model for people who present with hoarding disorder. There is currently a need to examine the acceptability of such a model.
Vision-Language Pre-training (VLP) has advanced the performance of many visionlanguage tasks, such as image-text retrieval, visual entailment, and visual reasoning. The pre-training mostly utilizes lexical databases and image queries in English. Previous work has demonstrated that the pre-training in English does not transfer well to other languages in a zero-shot setting. However, multilingual pre-trained language models (MPLM) have excelled at a variety of single-modal language tasks. In this paper, we propose a simple yet efficient approach to adapt VLP to unseen languages using MPLM. We utilize a cross-lingual contextualized token embeddings alignment approach to train text encoders for non-English languages. Our approach does not require image input and primarily uses machine translation, eliminating the need for target language data. Our evaluation across three distinct tasks (image-text retrieval, visual entailment, and natural language visual reasoning) demonstrates that this approach outperforms the state-of-the-art multilingual vision-language models without requiring large parallel corpora. Our code is available at https://github.com/Yasminekaroui/CliCoTea.
The Japanese CCGBank serves as training and evaluation data for developing Japanese CCG parsers. However, since it is automatically generated from the Kyoto Corpus, a dependency treebank, its linguistic validity still needs to be sufficiently verified. In this paper, we focus on the analysis of passive/causative constructions in the Japanese CCGBank and show that, together with the compositional semantics of ccg2lambda, a semantic parsing system, it yields empirically wrong predictions for the nested construction of passives and causatives.
Tastes affect the body and our emotions. We used tasteless, sweet, and bitter stimuli to induce participants' moods, and we examined the effect of mood on an emotional evaluation of pleasant, neutral, and unpleasant images using event-related potentials, N2, N400, and late positive potential (LPP), which reflect emotional evaluation in the brain. The results indicated that mood valence was most positive for sweetness and most negative for bitterness. Moreover, there was no significant mood effect on subjective valence ratings of emotional images. Furthermore, the N2 amplitude, which is related to the early semantic processing of preceding stimuli, was unaffected by the taste induced mood. In contrast, we found that the N400 amplitude, which is related to the mismatch of emotional valence between stimuli, increased significantly for unpleasant images when participants were in a positive rather than negative mood state. Also, the LPP amplitude, which is related to the emotional valence of images, showed only the main effect of the images' emotional valence. The N2's results suggest that the early semantic processing of taste stimuli might have had a negligible impact on emotional evaluation because taste stimuli minimize semantic processing that accompanies mood induction. In contrast, the N400 reflected the effects of the induced mood, and the LPP reflected the impact of the valence of emotional images. The use of taste stimuli to induce mood revealed different brain processing of taste-induced mood effects on emotional evaluation, including N2's involvement in semantic processing, N400's involvement in matching emotions between mood and stimuli, and LPP's involvement in subjective evaluations of stimuli.
Aesthetic evaluations, including beauty and attractiveness, have an important role in our lives. Despite its importance in our every-day life, enough attention has not been devoted to the assessment of place attractiveness in previous studies. We assume that changes in elements of square attractiveness are associated with changes in brain functional connectivity patterns. In this study, we have tried to explore the relationship between elements of square attractiveness and individuals' emotional perception as well as the brain mechanism involved in the process of cognitive development. There has been a focus on using objective measures of physiological rather than using self-reported data of an individual's emotions because people cannot understand their emotions properly and it is needed to compare self-report emotions with physiological processes. Classification of the five main elements of attractiveness was performed using the Delphi technique. Subsequently, twenty-four healthy young adults were exposed to the visual stimuli consists of five elements. A 32-channel EEG system was used to record the brain activities of participants while watching the stimuli. The subjects' feelings about valence and arousal levels of the elements were evaluated using the Self-Assessment Manikin (SAM) technique. The findings showed that “visual openness” is the most important element to increase the square attractiveness of everyday landscape in residential areas. The analysis revealed a significant difference (p = 0.048) in arousal ratings between more attractive (more openness) (M = 4.77) and less attractive (less openness) (M = 4.52). Attractiveness elements of the stimuli have a region-specific association with brain functional connectivity networks. This pattern is mainly found in the functional connections between central parts of the brain.