Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In our society men are considered more impulsive than women, especially in the violent and sexual domain. This correlation of sex and impulsivity might trace back to enhanced male impulsivity in general or a domain specific effect of emotions on impulsivity. The evidence for sex differences in the interaction of emotional or sexual stimuli and impulsivity has been relatively inconclusive so far. In this study, we investigated the effects of various emotional stimuli on responsivity in a Go/No-Go task. Participants had to respond quickly to a visual cue and withhold their response to another visual cue, while different emotional pictures were presented in the background, including sexual stimuli, non-sexual positive stimuli and negative stimuli. Both men (N = 37) and women (N = 38) made most commission errors in the sexual condition, indicating a disinhibiting effect in both genders. On top of this, men made even more commission errors than women, specifically in the sexual condition and not in other conditions. Men rated sexual stimuli as more positive, but did not differ from women in arousal ratings and pupil dilation. These findings may partly indicate increased impulsive behavior under sexual arousal in men, most likely driven by enhanced approach motivation due to more positive value but not higher arousal of sexual stimuli. The results are consistent with the theory of evolutionarily based concealment of sexual interest in women.
Abstract Background and Objective Dyspnoea is a debilitating symptom in individuals with chronic obstructive pulmonary disease (COPD) and a range of other chronic cardiopulmonary diseases and is often associated with anxiety and depression. The present study examined the effect of visually‐induced mood shifts on exertional dyspnoea in individuals with COPD. Methods Following familiarization, 20 participants with mild to severe COPD (age 57–79 years) attended three experimental sessions on separate days, performing two 5‐min treadmill exercise tests separated by a 30‐min interval on each day. During each exercise test, participants viewed either a positive, negative or neutral set of images sourced from the International Affective Picture System (IAPS) and rated dyspnoea or leg fatigue (0–10). Heart rate (HR) and peripheral oxygen saturation (SpO 2 ) were measured at 1‐min intervals during each test. Mood valence ratings were obtained using Self‐Assessment Manikin (SAM) scale (1–9). Results Mood valence ratings were significantly higher when viewing positive (end‐exercise mean ± SEM = 7.6 ± 0.3) compared to negative IAPS images (2.4 ± 0.3, p < 0.001). Dyspnoea intensity (mean ± SEM = 5.8 ± 0.4) and dyspnoea unpleasantness (5.6 ± 0.3) when viewing negative images were significantly higher compared to positive images (4.2 ± 0.4, p = 0.004 and 3.4 ± 0.5, p = 0.003). Eighty‐five percent of participants ( n = 17) met the minimal clinically important difference (MCID) criteria for both dyspnoea intensity and unpleasantness. HR, SpO 2 and leg fatigue did not differ significantly between conditions. Conclusion These findings indicate that the negative affective state worsens dyspnoea in COPD, thereby suggesting strategies aimed at reducing the likelihood of negative mood or improving the mood may be effective in managing morbidity associated with dyspnoea in COPD.
Primary visual cortex (V1) is generally thought of as a low-level sensory area that primarily processes basic visual features. Although there is evidence for multisensory effects on its activity, these are typically found for the processing of simple sounds and their properties, for example spatially or temporally-congruent simple sounds. However, in congenitally blind individuals, V1 is involved in language processing, with no evidence of major changes in anatomical connectivity that could explain this seemingly drastic functional change. This is at odds with current accounts of neural plasticity, which emphasize the role of connectivity and conserved function in determining a neural tissue's role even after atypical early experiences. To reconcile what appears to be unprecedented functional reorganization with known accounts of plasticity limitations, we tested whether V1's multisensory roles include responses to spoken language in sighted individuals. Using fMRI, we found that V1 in normally sighted individuals was indeed activated by comprehensible spoken sentences as compared to an incomprehensible reversed speech control condition, and more strongly so in the left compared to the right hemisphere. Activation in V1 for language was also significant and comparable for abstract and concrete words, suggesting it was not driven by visual imagery. Last, this activation did not stem from increased attention to the auditory onset of words, nor was it correlated with attentional arousal ratings, making general attention accounts an unlikely explanation. Together these findings suggest that V1 responds to spoken language even in sighted individuals, reflecting the binding of multisensory high-level signals, potentially to predict visual input. This capability might be the basis for the strong V1 language activation observed in people born blind, re-affirming the notion that plasticity is guided by pre-existing connectivity and abilities in the typically developed brain.
This paper deals with an electronic resource under construction. The objective is to construct a lexical database of Malagasy adjectives. Malagasy is an agglutinative language of Austronesian origin, spoken in the African island of Madagascar. The method used to construct the resource is adapted from the approach of Gross (1989) to electronic dictionaries. The content of the resource is based on linguistic analysis and encoded so as to be used by language-processing software. However, care has been taken that all linguistic information is easily readable and updatable. The resource allows for morphological analysis and generation of adjectives, removing obstacles to the construction of computer applications to process the Malagasy language. The originality of this paper also comes from our proposal of a distinction between adjectives in the usual sense and adjectival forms of other parts of speech.
Syntactic parsing is a highly linguistic processing task whose parser requires training on treebanks from the expensive human annotation. As it is unlikely to obtain a treebank for every human language, in this work, we propose an effective cross-lingual UD parsing framework for transferring parser from only one source monolingual treebank to any other target languages without treebank available. To reach satisfactory parsing accuracy among quite different languages, we introduce two language modeling tasks into the training process of dependency parsing as multi-tasking. Assuming only unlabeled data from target languages plus the source treebank can be exploited together, we adopt a self-training strategy for further performance improvement in terms of our multi-task framework. Our proposed cross-lingual parsers are implemented for English, Chinese, and 29 UD treebanks. The empirical study shows that our cross-lingual parsers yield promising results for all target languages, approaching the parser performance which is trained in its own target treebank.
Abstract In today’s multilingual lexical databases, the majority of the world’s languages are under-represented. Beyond a mere issue of resource incompleteness, we show that existing lexical databases have structural limitations that result in a reduced expressivity on culturally-specific words and in mapping them across languages. In particular, the lexical meaning space of dominant languages, such as English, is represented more accurately while linguistically or culturally diverse languages are mapped in an approximate manner. Our paper assesses state-of-the-art multilingual lexical databases and evaluates their strengths and limitations with respect to their expressivity on lexical phenomena of linguistic diversity.
Affect intensity refers to the intensity with which people experience their emotional response. Individual differences in affect intensity are supposed to be related to the strength of the response to emotional stimuli. Previous studies showed that participants with high affect intensity responded to emotional stimuli with stronger or more intense affective reactions than participants scoring low in affect intensity However, previous studies are mainly limited to the impact of affect intensity on consumer responses to advertising appeals or are limited to the use of life events descriptions as emotional stimuli. No previous studies used behavioural measures of the emotional response to standardized stimuli, varying in terms of arousal. In the present study the predictive value of affect intensity, measured by a self-report questionnaire, the Affect Intensity Measure (AIM), on the emotional response to standardized pictures and sounds has been investigated. In particular, the predictive value of affective intensity measured by the AIM, using both the total AIM total score and the four subscales scores, on subjective arousal ratings of different categories of standardized emotional pictures and sounds was assessed on a nonclinical sample. The total AIM score has been found to be predictive for subjective arousal scores for low unpleasant pictures while, using the AIM subscales scores, results showed that the Negative Reactivity subscale was predictive for arousal scores to high negative pictures and sounds. These findings seem to show that the use of the total AIM score can obscure the relationships between specific features of affect intensity and other variables. Moreover, the present results didn’t show a general effect of affect intensity on behavioural responses to emotional standardized stimuli but an emotion specific effect for high negative stimuli.
Considerando a importância que dependências sintáticas vêm assumindo em tarefas de Processamento de Linguagem Natural (PLN) e, consequentemente, nos estudos linguísticos voltados para o processamento automático das línguas, apresentamos aqui uma avaliação qualitativa de um treebank padrão ouro recém lançado para a língua portuguesa, com o objetivo de identificar (a) os padrões linguísticos que apresentam maior dificuldade para anotadores automáticos, (b) os motivos que podem levá-los a errar essas análises, e (c) ampliar as possibilidades de diálogo entre os estudos linguísticos e a linguística computacional. A anotação sintática foi realizada conforme as diretrizes do projeto Universal Dependencies (UD), e a avaliação da anotação foi realizada utilizando ferramentas de código aberto, em três etapas: em primeiro lugar, fizemos uma avaliação intrínseca de um modelo de dependências sintáticas, assumindo que este tipo de avaliação reflete indiretamente a consistência da anotação do corpus com o qual o modelo foi treinado; em seguida, detalhamos os resultados desta avaliação, apresentando o índice de acertos de cada classe linguística individualmente, o que nos deu um panorama das dificuldades linguísticas para o aprendizado automático e, também, informação quanto à confiança na análise automática de cada classificação linguística. Por fim, selecionamos as classes com maior número de erros e analisamos todos os casos errados. Os resultados sugerem que, do lado linguístico, já podemos contar com análises consistentes e em quantidade, ao menos aparentemente, suficiente. No que se refere à qualidade dos parsers automáticos, o espaço para melhorias linguísticas é cada vez menor.
These data include COVID history scores, image ratings, and pandemic disruption scores. Covariate measures including demographic factors are also included.
Automated Multiple-Choice Question (MCQ) generation is a rapidly growing field in Natural Language Processing (NLP) that aims to assist educators and trainers in creating high-quality, efficient, and effective assessment materials. This is achieved by analyzing large amounts of textual data, such as educational content, and identifying key concepts and relationships between them. The process of generating MCQs automatically can be broken down into several steps, including text summarization, keyword extraction, distractor generation, and sentence mapping. A new approach is proposed for Text summarization is based on NLP models like XLNet and YAKE is used for keyword extraction. Distractor generation is done using lexical databases such as ConceptNet and WordNet. Sentence mapping is used to identify the main concepts and relationships within a text, which can then be used to formulate questions and options for the multiple-choice questions. The output is a set of MCQs that are semantically related to the input text and can be used for educational and training purposes.
The folder contains subjective arousal ratings, eye-tracking data (fixation times) and EEG data relative to the processing of emotional body expressions presented at the end of a virtual promenade within different architectural forms. Scripts are also provided for the reproducibility of statistical data analysis. Please refer to the "Readme.txt" file for more detailed information
* Introduction This is the Khmer ALT of the Asian Language Treebank (ALT) Corpus. Please refer to<br> http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/index.html<br> for an introduction of the ALT project. The process of building the Khmer ALT began with sampling about 20,000 sentences from English Wikinews, and then these sentences were translated into Khmer language.<br> <br> The English Wikinews<br> https://en.wikinews.org/wiki/Main_Page<br> is available under the terms of the Creative Commons Attribution 2.5 License.<br> https://creativecommons.org/licenses/by/2.5/ * License Khmer ALT has been developed by NICT and CADT (a.k.a. NIPTICT). The license of Khmer ALT is Academic Research Non-Commercial Limited CC-BY-NC-SA Reference-Type License Agreement Terms and Conditions The use of this material (copyrighted material, data, etc.) is permitted under the conditions similar to the Creative Commons "Attribution-NonCommercial-ShareAlike" License. However, the purpose of use shall not only be "NonCommercial" but must be “for Academic Research and Non-Commercial”.<br> The detailed legal provisions are listed at the end of the document, but the outline is as follows. You are free to:<br> Share — copy and redistribute the material in any medium or format<br> Adapt — remix, transform, and build upon the material<br> The licensor cannot revoke these freedoms as long as you follow the license terms. Under the following terms:<br> Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.<br> Academic Research&Non-Commercial — You may use the material only for academic research purposes and may not use it for commercial purposes.<br> ShareAlike — If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original. In this Terms and Conditions, academic research purposes are not associated with the interpretation of the Creative Commons License, and shall refer to "the purpose of intellectual creative activities conducted by persons belonging to institutions of higher education and research institutes, etc. in order to develop a better and affluent society through exploration of the truth about nature, humans, society, etc., discovery of new principles and laws, and application of these in a wide range of fields from natural sciences to human and social sciences." Non-commercial means: not for the purpose of any contribution to a for-profit or commercial business. The outcome of research and development activities initially undertaken without the purpose of contributing in any way to a for-profit or commercial business shall not be permitted to be used later to contribute to one. The legal provisions of this license are: the terms of Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode) plus the following rider clauses. On that note, the Terms and Conditions of this license are not equal to that of the Creative Commons license—they were inspired by them and share most of the terms and conditions with them, but shall be considered as a different license. Rider Clauses<br> 1. "Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License" in the title and preamble shall read "Academic Research Non-Commercial Limited CC-BY-NC-SA Compliant License Agreeement Terms and Conditions".<br> 2. Article 1 c and g shall be deleted.<br> 3. Article 1 k shall be amended as follows:<br> k. "Academic Research Purposes" shall refer to the purpose of intellectual creative activities conducted by persons belonging to institutions of higher education and research institutes, etc. in order to develop a better and affluent society through exploration of the truth about nature, humans, society, etc., discovery of new principles and laws, and application of these in a wide range of fields from natural sciences to human and social sciences, and does not consider contribution to for-profit or commercial businesses as its main or secondary purpose. Activities initially undertaken without the purpose of contributing to a for-profit or commercial business, shall no longer be considered as "academic research purposes" as soon as they decide to contribute to one.<br> 4. Article 3 b.1. shall be replaced by the following:<br> 1. The Adapter's License that You apply must be this License.<br> 5. In all provisions, "NonCommercial" shall read "Academic Research Purposes" and "this Public License" shall read "this License". Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License<br> <https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode> * Usage 1. Compile the source codes.<br> 2. Run with no option to see the instruction.<br> 3. Follow the instruction to generate the treebank.
Datasets\nIMDb Large Movie Review Dataset - This dataset contains movie reviews along with their associated binary sentiment polarity labels. It is intended to serve as a benchmark for sentiment classification.\nLink: http://ai.stanford.edu/~amaas/data/sentiment/\n\nSentiment140 - This dataset contains 1.6 million tweets extracted using the Twitter API. The tweets have been annotated (0 = negative, 2 = neutral, 4 = positive) and they can be used to detect sentiment.\nLink: http://help.sentiment140.com/for-students/\n\nYelp Reviews - An open dataset released by Yelp for learning purposes. It consists of millions of reviews with star ratings that can be used for sentiment analysis.\nLink: https://www.yelp.com/dataset\n\nAmazon Reviews for Sentiment Analysis - This dataset contains product reviews and metadata from Amazon, including 142.8 million reviews spanning May 1996 - July 2014. The dataset includes reviews (ratings, text, helpfulness votes), product metadata (descriptions, category information, price, brand, and image features), and links (also viewed/also bought graphs).\nLink: http://jmcauley.ucsd.edu/data/amazon/\n\nTwitter US Airline Sentiment - A sentiment analysis job about the problems of each major U.S. airline. Twitter data was scraped from February of 2015 and contributors were asked to first classify positive, negative, and neutral tweets, followed by categorizing negative reasons.\nLink: https://www.kaggle.com/crowdflower/twitter-airline-sentiment\n\nStanford Sentiment Treebank - This dataset includes fine-grained sentiment labels for 215,154 phrases in the parse trees of 11,855 sentences.\nLink: ctrader vs mt4\n
Abstract: Parent–child attachment is robustly associated with typical patterns of emotion regulation but rarely examined in relation to changes in emotion in response to events. We studied how attachment is related to emotion reactivity to positive and negative events and to immediate and delayed emotion recovery from social exclusion. The sample (78% White, 46% girls) included 110 children (9–12 years). Children completed a story stem measure that was scored for security, avoidance, ambivalence, and disorganization. Emotion reactivity and recovery were assessed with child-reported and observer positive affect and negative affect ratings. Parents rated child temperament. More avoidant children showed dampened emotion responding (low reactivity and recovery), whereas more ambivalent children showed heightened emotion responding (high reactivity and recovery). Attachment security, disorganization, and temperament were not consistently related to emotion reactivity or recovery. The findings highlight that emotion regulation occurs in response to contextual changes and is related to attachment.
Abstract Introduction “Sleep to forget and sleep to remember” model postulates that the content of emotional memory is strengthened during sleep while the arousal component is attenuated. Experimental sleep deprivation studies in healthy individuals have supported the model, but it is not known if sleep in individuals with insomnia, a condition of chronic poor sleep, provides the same effect on emotional memory. Methods Individuals reporting insomnia and good sleepers (n = 76 and n = 83 respectively; 83% female, mean age of 43.2 years) completed a picture-rating task to assess a) emotional valence ratings pre and post sleep and b) recognition accuracy post-sleep. In addition, an autobiographical memory task was completed during the day to assess a) the number of recalled memories in a 2-minute period for good and bad past events respectively, and b) emotional intensity ratings of the retrieved memories at the time the event occurred and how they feel now. Results Wilcoxon rank-sum test was used to assess group differences. Sleep diary variables on the experimental night showed that individuals with insomnia experienced a significantly more disrupted sleep than good sleepers (SOL: r = 0.537; WASO: r = 0.620; TST: r = 0.669). In the picture-rating task, there was no significant group difference for recognition accuracy (neutral stimuli: r = 0.016; negative stimuli: r = 0.018) or change in emotional valence pre-to-post sleep (neutral stimuli: r = 0.003; negative stimuli: r = 0.008). In the autobiographical memory task, individuals with insomnia remembered significantly fewer good days (r = 0.245) and rated their bad days at the time of its occurrence as significantly worse than good sleepers (r = 0.235). There was no significant group difference in the extent to which emotional intensity faded more for bad days than good days (r = 0.113). Conclusion We have shown that autobiographical memory is altered in insomnia which may have important influence on mental health. Our findings also suggest that overnight memory consolidation for negative stimuli is not affected in insomnia versus good sleepers. This may be a reflection of our one-night protocol; future work is needed across consecutive days. Support (if any) Dr-Mortimer-&-Theresa-Sackler-Foundation
The article focuses on the analysis of anglicisms in the professional speech of language teachers on the basis of educational and methodological materials posted on the popular educational platforms “Vseosvita” and “Na urok”. The main focus is on the compounds with the element “бук”, as the number of such titles has increased significantly in recent years. The speech of teachers of language and literature abounds in the following anglicisms: буктрейлер, буккросинг, артбук, воркбук, лепбук, скрапбук. It is noteworthy that most of these new words are recorded only in the Dictionary of Modern Anglicisms, which was published in 2022. Other lexicographical sources do not list them, and also do not contain the component “бук” either as part of other compounds or as an independent word. Anglosurzhyk words appear in teachers’ speech due to various factors: learning from the professional experience of fellow teachers from English-speaking countries, using working methods from such fields as economics, business and management in the educational process, translation difficulties, the desire to make it concise, the desire to embellish their speech with unusual words and evoke interest of the students. Functioning in speech, these new words play the role of professionalisms. The latter are known for their stylistically colouring, they exist beyond the linguistic norm and may disappear over time. However, the widespread usage of these names in educational and methodological works indicates the desire to make these words normative terms. Compound words of English origin with the component “бук” used by teachers of language and literature do not denote new concepts, and therefore need to be replaced by those with more transparent semantics. The teachers frequently use foreign words to denote concepts that already have or may have Ukrainian names. Transliteration of English words does not enrich our language, but contaminates it and leads to the formation of another type of language mix – anglosurzhyk.
I delivered (24.03.2026) a seminar to over 100 teachers from Greece’s Model–Experimental Schools, introducing Helex Kids 2 alongside the Word Tool a structured lexical database I developed to support learners with reading and spelling difficulties. The session focused on translating research into classroom practice, equipping educators with evidence-informed strategies and practical digital resources to enhance literacy instruction. The strong engagement and discussion that followed highlighted both the demand for targeted literacy support tools and the value of bridging innovation with everyday teaching practice.<br/><br/>Feedback received: Participant feedback highlighted the clarity and accessibility of the seminar, with teachers noting the use of simple, comprehensible language, concise and well-structured delivery, and clear explanations of complex concepts such as neuronal processes in spelling and the involvement of multiple brain centres in reading and writing; they particularly valued how the session illuminated the difficulties faced by children while addressing a wide range of issues in a focused and highly effective way.<br/>Professor Anthi Revithiadou (Linguistics, Aristotle University of Thessalonica) and event coordinator described the seminar as “outstanding,” adding, “I truly have no words to thank you, excellent in every respect. I say this with complete sincerity. Thank you once again.”<br/>
Abstract The labor market is a key part of an economy. Several existing online platforms allow the upload of resumes and the search for a job. One of their limitations, however, is that obtaining the best opportunity can be hard because certain jobs need some experiences, abilities, and features that an applicant might not know. The recent diffusion and employment of conversational agents definitely have proven to benefit this kind of issue. For example, ChatGPT has shown impressive outcomes in different domains and for a variety of tasks. It has weaknesses, although, related to the veracity of the responses it generates, which might deceive the user interacting with it. The usage of external domain knowledge is the direction we suggest in this chapter. Several lexical databases and taxonomies have already been collected and designed by different organizations. We illustrate a list of such resources and provide a solution that integrates conversational agents with relevant information extracted from one of such resources showing the benefits and the impact that our proposal can generate.
The internal representations of large language models (LLMs) remain largely opaque, hindering interpretability, alignment, and cross-lingual transfer. A core obstacle is polysemanticity—the phenomenon whereby individual neurons activate for semantically unrelated concepts—which arises from the superposition of more features than there are physical dimensions. Existing approaches, such as Sparse Autoencoders (SAEs), address this by decomposing embeddings into millions of monosemantic features, yet they recover no macroscopic coordinate system that organizes these features. In this work, we introduce the Atlas Autoencoder (Atlas AE), a topology-preserving autoencoder trained on approximately 100,000 English dictionary definitions, which compresses 768-dimensional transformer embeddings onto a 128-dimensional latent manifold. Linear probing within this latent space recovers eleven interpretable cognitive axes—Concreteness, Agency, Sentiment, Intentionality, Sociality, Power, Temporality, Dynamicity, Negation, Quantity, and Perceptuality—each grounded in established psycholinguistic norms (AUROC 0.75–0.96; all permutation ). Residual subspace analysis confirms that the semantic space saturates at exactly eleven stable dimensions. To assess universality, we project embeddings from five independent transformer models spanning five typologically distinct language families—Indo-European (English, French), Uralic (Finnish), Altaic (Turkish), and Japonic (Japanese)—through the same frozen, English-trained manifold. The resulting coordinate systems exhibit striking geometric isomorphism: core structural axes such as Power and Agency show cross-model variance, and all five languages agree on directional polarity for nine of eleven axes. Inter-axis correlation analysis reveals that the recovered basis is structurally oblique, with stable covariances (e.g., Power–Sociality ) that challenge the widespread orthogonality assumption in representation learning. These findings establish the Universal Semantic Manifold as a compact, interpretable coordinate system that bridges connectionist representations and symbolic cognition, offering a rigorous geometric framework for cross-lingual alignment, controlled semantic steering, and the quantitative study of conceptual structure.
This article in addition to introducing and defining taboos, examines the existing and well-known taboos in the collection of short stories” To whom should I greet” by the famous Iranian novelist SiminDaneshvar, during which, it refers to the use of taboo words by the characters of gender(female/male) in social situations. The authors have tried to include social behaviors such as: good/bad, holy/ unholy, polite/ impolite, etc. in social relations and dialogues of the work, based on the accepted norms of the Iranian society in front of the readers. In adition to be able to explain the relationship between culture and language and the intractions between the two in terms of prohibition in the sociology of language based on gender differences and how they are used. Besides, to prove that in this work, Simin while paying attention to the values of the Persian society, has consciously tried to break the linguistic norms in the daily individual and social life of the characters of some of stories in many cases. This study also tried to show the results obtained from the frequency of using taboos and inconveniences in the actions and speech of speakers by gender in the form of a graph at the end of the article and between gender and taboos in terms of their use in speakers there was a significant relationship.
Sociolects, as social varieties of language, should be classified and described on the basis of both linguistic and sociological criteria. In practice, however, most often only either the first one (e.g. according to Antoni Furdal and Danuta Buttler) or the second one is used (e.g. according to Aleksander Wilkoń and Tomasz Piekot). Both these criteria are less frequently combined (Stanisław Grabias). The article proposes those aspects of the sociological and linguistic functioning of language varieties that should be used by sociolinguistics especially in the characterisation of sociolects, but also in their typology. The sociological criteria include: the type of contacts in the group (including contacts made via the Internet), the degree of group formalisation, the durability of the group and the type of bond that binds the group together. Among the linguistic criteria the following were distinguished: nomination and expressiveness, the method of acquiring a sociolect and the degree of its codification and the rank of the linguistic norm. The article also clarifies the concept of the distinguishing features of the sociolectic vocabulary.
Focusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet.In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g.elements of wordnet taxonomy, quantifier phrases, certain collocations).In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches.We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one).Language model also proves to be better than a feature-based logistic regression model.
The paper examines team building in multicultural and multiethnic work environments – particularly in NGOs – by linking classic accounts of effective teams with sociolinguistic and socio-emotional perspectives. It conceptualises workplace teams as speech communities in which members share, negotiate, and sometimes contest linguistic norms, expectations, and emotional display rules. Drawing on management and organisational studies, the text outlines key structural conditions for effective teams, including clear goals, appropriate leadership, resource allocation, and mechanisms for accountability and cooperation. These insights are then integrated with research on emotions, culture, and communication, which highlights the role of emotional attachment, trust, and creative problem-solving in sustaining team cohesion. The paper argues that effective team building in such contexts requires not only formal structures but also deliberate cultivation of shared communicative practices and affective climates that support participation, innovation, and mutual understanding across cultural and linguistic differences.
Nouns are one of the most frequently used word classes in English. As one variety of “World Englishes”, China English has received much concern in recent years. However, most research on nouns is still at the cognitive aspect or in the traditional qualitative way. As a new quantitative method, dependency grammar reveals the internal relations among the words in the sentence. Therefore, the syntactic features of nouns in China News English can provide a new perspective. The raw material of China English in this study is randomly crawled from China Daily, and the News part in the Corpus of Contemporary American English (COCA) from 2013 to 2016is selected as the reference corpus. By using the Jupyter Notebook based on Python, the dependency treebanks are generated. After comparing the proportion of the word classes, dependency direction, and dependency distance of different dependents of nouns in China Daily and COCA, this paper finds that 1) On the whole, the categories of dependents in China English and American English are similar. Determiners, adjectives, nouns, prepositions, verbs, numbers, proper nouns, conjunctions, pronouns and particles are the top ten frequently-used modifiers of nouns. Compared with American English, China English tends to use more adjectives and nouns, while the determiners especially the articles and pronouns are relatively less used. 2) Both in China English and American English, the percentage of HI constructions and HF constructions is close to 50%. China English tends to be more right-embedded than American English. Compared with AmE, in ChiE, prepositions, conjunctions and participles are more frequently used as post-modifiers. 3) The mean dependency distance value of China English (2.68) is larger than American English (2.61). Both in China English and American English, the mean dependency distance of pre-modifiers is shorter than in post-modifiers. The longer the mean dependency distance value is, the more difficult is in processing the information. It indicates that in the News domain, China English is more difficult to be processed in expression. This research provides a new perspective to studying English variants and enriches the current research on nouns.
This paper provides a comparative analysis of the patterns of formation and qualities of a modern linear text and an Internet text. The article is the result of the study of Internet text stylistics, based mainly on Russian-language texts from Russia and Ukraine. The paper considers the process of formation of a new type of text - the Internet text. Being essentially different from the classical linear text, the Internet text does not lend itself well to the description based on the classical text theory. Thus, the Internet text is not complete, vectorial, not united by the completeness of thought expression. An important characteristic of the Internet text is its interactivity, which means that the roles of addressee and addressee are constantly changing. In addition, it is difficult, or even impossible, to define the boundaries of the Internet text due to its hypertextuality, which has become habitual intertextuality. All the above-mentioned aspects make up the pragmatics of the Internet text as a subspecies of the media text. Another crucial problem addressed in the paper is the study of the regularities of Internet communication in general and the stylistics of the Internet text. In the course of the research it became obvious that speech aggression and violation of norms of speech culture are stylistic dominants of online communication. This influenced the formation of other stylistic dominants such as, hate speech, fake, hype, clickbait, etc. Internet style is clearly characterized by being provocative, aggressive, hostile. The problems of bullying, humiliation of human dignity, invective and obscenity are actively studied from the standpoint of linguoecology, because, according to most researchers, the constant neglect of communicative and linguistic norms leads to the degradation of the national language style and literary norms. Despite the fact that there is still a division between public and interpersonal online communication, i.e. formal and informal, the problems of speech behavior of Internet users are becoming increasingly relevant.
This chapter synthesizes the most prominent Natural Language Processing (NLP) studies conducted on Persian, focusing on text processing. The first section contains selected tasks from the NLP pipeline, such as text preprocessing, tokenization, POS tagging, syntactic parsing, treebank annotation or semantic analysis along with examples of how researchers approached the problem for Persian and, where applicable, examples of tools developed to perform given tasks. The following section discusses the application of Persian NLP like spell-checking, information retrieval, machine translation or sentiment analysis. Finally, the last section summarizes the Persian NLP corpora and other resources.
Decoding Part of Speech(POS) tagging directly from electroencephalography (EEG) signals whilst user overtly spoke (voiced speech) sentences could improve direct speech brain-computer interfaces (BCIs) using imagined or inner speech. To the best of our knowledge, earlier work uses machine learning approach using 74,953 sentences/tokens recorded in 75 EEG sessions. The tokens can be found in 4,479 phrases consisting of terms from the English Online treebank which contains the record of weblogs, newsgroups, reviews, and Yahoo Answers. The results demonstrated the feasibility of POS decoding from EEG based on word class, word frequency, and word length with accuracy of 71%, 86%, 89%, respectively. We believe that there is significant room for improvement with more advanced artificial intelligence. In this paper, we further extend the existing work with end-to-end transformers. Our results presents transformer model outperforms benchmark traditional ML results with +20% in length, +13% for the open vs closed class and +12% in frequency. In our empirical analysis, we find the decoding performance was better when using multi-electrode recordings as compared to single-electrode recordings.
В статье анализируется историческая динамика политической корректности, ее положительные и отрицательные стороны, а также прослеживается ее связь с вежливостью, которая заключается в неиспользовании конфронтационных, ликоповреждающих коммуникативных стратегий при указании на гендер, расу, этнос, возраст, физическое состояние и социальное положение адресата. Исследуются факторы, затрудняющие формирование политкорректного русского языкового сознания: 1) не вполне сформировавшееся понятие ПК применительно к отечественному социальному контексту, осложненное его концептуализацией через призму западного восприятия; 2) отсутствие правовых механизмов ПК, несмотря на недопустимость дискриминации, закрепленную в Конституции РФ; 3) неразработанность языковых основ применения ПК в российском публичном дискурсе; 4) англоцентричность правил ПК для международного общения. Сделан вывод о необходимости выработки российских норм политической корректности с участием широкого лингвистического сообщества. The paper explores the historical dynamics of political correctness (PC), its positive and negative aspects, as well as its connection with politeness as an avoidance of confrontational, face-threatening strategies in reference to the interlocutor’s gender, race, ethnic background, age, physical condition and social status. The study also deals with the factors hindering the formation of the Russian PC awareness, which include: 1) the incompletely formed notion of political correctness in the Russian social context complicated by its conceptualization through the prism of Western comprehension; 2) absence of legal PC mechanisms, in spite of non-discrimination enshrined in the RF Constitution; 3) underdeveloped linguistic norms of political correctness in Russian public discourse; 4) anglocentrism of PC rules in intercultural communication. The article concludes by proposing a wide discussion of Russian political correctness norms involving a wider linguistic community.
Abstract We propose a Slovak language model for the spaCy library in Python. These models are easy-to-use for basic natural language processing tasks in a single package. The package contains several components for basic preprocessing tasks, such as tokenization, sentence boundary detection, syntactic parsing, lemmatization, named entity recognition, morphology analysis, and word vectors. It is based on the state-of-the-art monolingual SlovakBERT model. Named entity recognition is trained on a separate, publicly available WikiAnn database. The other statistical classifiers use a Slovak Dependency Treebank corpus. Morphological tags are compatible with the conventions of the Slovak National Corpus. The part of speech tags use conventions of the Universal Dependencies framework. We trained a separate word vector model on a web-based corpus. The training uses fastText with Floret modification. We present a series of experiments that confirm that the model performs similarly to other languages for all tasks. Training scripts and data are publicly available.
Two distinct literatures have evolved to study within-person changes in affect over time. One literature has examined affect dynamics with millisecond-level resolution under controlled laboratory conditions, and the second literature has captured affective dynamics across much longer timescales (e.g., hours or days) within the relatively uncontrolled but more ecologically valid conditions of daily life. Despite the importance of linking these literatures, very little research has been done so far. In the laboratory, peak affect intensities and reaction durations were quantified using a paradigm that captures second-to-second changes in subjective affect elicited by provocative images. In two studies, analyses attempted to link these micro-dynamic indexes to fluctuations in daily affect ratings collected via daily protocols up to 4 weeks later. Although peak intensity and reaction duration scores from the laboratory did not consistently relate to daily scores pertaining to affect variability or instability, the total magnitude of changes in affect following images did display relationships of this type. In addition, higher peaks in the laboratory predicted larger intensity reactions to salient daily events. Together, the studies provide insights into the mechanisms through which correspondences and noncorrespondences between laboratory reactivity indices and daily affect dynamic measures can be expected.
Extinction training has proved effective to diminish the expectancy of the aversive unconditioned stimulus (US). However, the negative valence of the conditioned stimulus (CS) may still stay intact. In fact, several studies have suggested that the CS negative valence may be a factor that promotes the return of fear. Our study focuses on the role of changes in the CS valence as a potential mechanism to reduce the spontaneous recovery of threat expectancies. To do that, we evaluated counterconditioning (CC), a technique aimed to reduce the CS negative valence by paring it with a positive stimulus and compared its efficacy to that of a novelty-facilitated extinction (NFE) and a standard extinction interventions. Using a 2-day protocol, participants first learned the relationship between a figure and an aversive sound, using a differential conditioning paradigm, and were then randomly assigned to one of three different groups. For the CC group, CS+ or cue A was paired with a positive US. The standard extinction group was exposed to cue A alone. For a third NFE group, cue A was followed by a neutral US. Finally, on the second day, spontaneous recovery was tested. Our findings did not provide evidence to suggest that CC could be more effective to prevent or reduce the return of threat expectancies or influence valence ratings when compared with NFE and standard extinction.
Abstract We propose a novel graph-based approach for semantic parsing that resolves two problems observed in the literature: (1) seq2seq models fail on compositional generalization tasks; (2) previous work using phrase structure parsers cannot cover all the semantic parses observed in treebanks. We prove that both MAP inference and latent tag anchoring (required for weakly-supervised learning) are NP-hard problems. We propose two optimization algorithms based on constraint smoothing and conditional gradient to approximately solve these inference problems. Experimentally, our approach delivers state-of-the-art results on GeoQuery, Scan, and Clevr, both for i.i.d. splits and for splits that test for compositional generalization.
Individuals differ in the tendency to derive pleasure out of motive-specific incentives, such as being socially included or attaining power. Multiple theoretical approaches have proposed that such motive-specific positive affective contingencies (PACs) are central building blocks of motive dispositions and personality more broadly. In the current research, we put this claim to test and investigated individual differences with regard to motive-specific PACs in the affiliation and power domains. We measured PACs via spontaneous emotional reactions to motive-specific cues, as assessed by affect ratings and electromyographic (EMG) recordings of smile responses. Both of these PAC operationalizations were highly internally consistent and moderately to highly stable across time. Furthermore, motive-specific PACs were linked in a manner consistent with theory to measures of motive dispositions and to personality traits with motivational underpinnings (i.e., extraversion, agreeableness, and narcissism). Finally, in the affiliation domain, motive-specific PACs were linked to objectively assessed, key motivational outcomes (i.e., attentional orientation, behavior in daily life, and in the laboratory). Taken together, the findings underscore the relevance of affective contingencies for the understanding of personality and motivated behavior.
Abstract This study represents the first stage of evaluating whether cognitive training interventions may be facilitated by the presence of a socially assistive robot (SAR) and gamification. Our experimental setup involves using a SAR providing feedback to a gamified visuospatial working memory task, administered according to a differential outcomes training (DOT) protocol. The study’s main objective was to investigate whether performance and attitude towards the task would be affected by different robotic setups (none, simulated or physical) and in relation to different challenge levels. We measured performance accuracy on the gamified visuospatial memory task and self-reported affective ratings, which are relevant for assessing attitude towards the task and providing indicators to the potential for using a SAR for a longer-term cognitive intervention. Additionally, we conducted exploratory analyses of eye movement strategies for memory encoding during the task. The results demonstrated a significant differential outcomes effect (DOE) on memory performance accuracy, regardless of Robot type and Challenge level, providing evidence that a DOE can still be obtained when a SAR interacts with participants. Moreover, the results from the affective ratings revealed that participants accompanied by the physical robot reported lower levels of stress and increased levels of control. Our results demonstrate, for the first time, a DOE using a SAR in a gamified context. This result, coupled with positive subjective reporting of the human–robot interactive experience of participants, demonstrates the potential for using a SAR to: (i) promote positive attitudes for a DOT-based cognitive intervention, without (ii) negatively affecting task performance.
Abstract We introduce an extensive dataset for multilingual probing of morphological information in language models (247 tasks across 42 languages from 10 families), each consisting of a sentence with a target word and a morphological tag as the desired label, derived from the Universal Dependencies treebanks. We find that pre-trained Transformer models (mBERT and XLM-RoBERTa) learn features that attain strong performance across these tasks. We then apply two methods to locate, for each probing task, where the disambiguating information resides in the input. The first is a new perturbation method that “masks” various parts of context; the second is the classical method of Shapley values. The most intriguing finding that emerges is a strong tendency for the preceding context to hold more information relevant to the prediction than the following context.
CONTEXT: Congenital adrenal hyperplasia (CAH) is a genetic disorder that results in hormonal imbalances and decreased brain volumes in regions important for emotional processing. OBJECTIVE: To examine whether emotion perception differs between youth with CAH and control youth, and if these differences relate to brain volumes. METHODS: In this cross-sectional study of 27 youths with CAH (mean age = 12.63 years, 16 female) and 35 age- and sex-matched controls (mean age = 13.03 years, 20 female), each participant rated picture stimuli and completed a 3T structural brain scan. Valence and arousal ratings and reaction times of 61 affective images were assessed. Gray matter volumes were measured by MRI. RESULTS: Youth with CAH had lower valence ratings for negative (P =.007) and neutral (P =.019) images. Controls showed differences in reaction times and arousal ratings across stimuli conditions, but youth with CAH did not. Brain volumes of the right amygdala (P =.025) and left hippocampus (P =.002) were associated with valence ratings. Left rostral middle frontal (P <.001) and right medial orbitofrontal cortex (P =.002) volumes were negatively related to valence scores only in youth with CAH, whereas left medial orbitofrontal cortex (P <.001) volumes were associated with valence scores positively in youth with CAH and negatively in controls. CONCLUSION: Findings suggest that youth with CAH perceive emotive stimuli as more unpleasant. Decreased brain volumes in the amygdala, hippocampus, and prefrontal cortex are associated with these measures of altered emotion perception in youth with CAH.
This dataset contains EEG, ECG and audio recordings of 11 individual musicians playing emotional music on their instrument. The dataset consists of two parts: Experiment: Musicians’ self-reported ratings, audio recordings, and physiological recordings where the 11 expert musicians were asked to play at least 4~2-minute unfamiliar (non-popularly known) musical pieces. Participants were asked to play at least once one of the following emotions: happiness, sadness, relaxation, and anger. For the physiological recordings EEG, ECG, and GSR signals were recorded. Each musican’s data is denoted by MS_ followed by the order in which they were recorded. Self-report questionnaire: A self-assessment questionnaire and their answers where 11 expert musicians were asked to rate musical pieces recorded based on: Objective valence & arousal they felt the piece had; Felt valence & arousal during playing. For a more detailed explanation of the dataset, its recording procedure, and its contents, see L. Turchet, B. O'Sullivan, R. Ortner & C. Gugher (2024). Emotion Recognition of Playing Musicians from EEG, ECG, and Acoustic Signals. IEEE Transactions on Human-Machine Systems. File Listing The following files are available (each explained in more detail below): Name Format Contents EEG_ECG_data_for_each_musician mat This folder contains the raw EEG, ECG, & GSR data for all 11 musicians for each piece they played as well as a resting state recording, which was recorded while a neutral audio stimulus was played. audio_data_for_each_musician wav, JSON, csv This folder contains 3 subfolders: 1) wav_audio_files_original: This is the raw audio data recorded for each musician and includes all 56 pieces included in the paper reported above, please see below regarding rejected trials. 2) wav_audio_files (original split_into_3_parts): Here, the 56 pieces are appropriately split into 3 separate parts.3) analysis_audio_files: This folder contains the acoustic features extracted, 1714 acoustic features were extracted from each split trial. Each result is stored in.JSON, however, the collated results can be seen in all_results.csv. self_report_questionnaire pdf, xls Two files exist in this folder:1. The self-reported questionnaire given to the musicians of the questionnaire during the experiment.2. File Details EEG_ECG_data_for_each_musician These are the original raw data recordings. EEG data were recorded using a g.GAMMAcap2 by g.tec Medical Engineering, a 64-channel cap with g.SCARABEO active electrodes, with two g.GAMMAsys reference active ear clip electrodes. Two g.GAMMAbox electrode connector boxes were used to connect the active electrodes to two g.USBamp biosignal amplifiers with a sampling frequency of 256 Hz.The following 31 EEG channels were used: Fp1, Fp2, AFz, AF3, AF4, AF7, AF8, Fz, F3, F4, F7, F8, Cz, C3, C4, CP3, CP4, CP5, CP6, P1, P2, P3, P4, P5, P6, P7, P8, PO7, PO8, O1, and O2. AFz was used as a ground electrode and Cz was used as a re-reference electrode. The right-side ear clip electrode was used as a reference electrode. ECG data were recorded using a single g.GAMMAclip active electrode clip connected directly to the g.GAMMAbox, sharing the same ground electrode with the EEG cap and placed on position V4 of the subjects.GSR data was recorded using the g.GSRsensor² box which contains two small dry electrodes placed underneath the participant’s toes (due to the amount of hand movement required for playing an instrument). The g.GSRsensor² was connected directly to a g.GAMMAbox, using jumper cables connected to different ports in the g.USBAMPs to share the same reference and ground electrodes as the EEG cap. The locations of the channels and their corresponding number in the raw data is as follows: Channel Number Channel Name 1 Time series 2 AF3 3 AF4 4 AF7 5 AF8 6 CP3 7 CP4 8 CP5 9 CP6 10 P1 11 P2 12 P5 13 P6 14 P7 15 P8 16 O1 17 O2 18 Cz 19 Fp1 20 Fp2 21 F3 22 F4 23 F7 24 F8 25 C3 26 C4 27 P3 28 P4 29 PO7 30 PO8 31 AFz 32 GSR 33 ECG audio_data_for_each_musician/wav_audio_files_original This folder contains all of the raw audio files recorded during the experiment. All pieces were recorded using the software Audacity and exported as WAV files encoded with a bit depth of 32-bits and a sampling rate of 44.1 kHz. 56 of the included trials are present in this folder. audio_data_for_each_musician/wav_audio_files (original split_into_3_parts) This folder contains the above mentioned 56 raw audio pieces separated into 3 appropriately sized recordings. audio_data_for_each_musician/analysis_audio_filesThis folder contains the 1714 acoustic features from each of the 3 separated trials from 56 accepted pieces, these are denoted by MS_ followed by the order in which the musicians were recorded and the order in which the trails were split, the individual features are in JSON format. Within this folder, we have collated all results in a csv file (all_results.csv) where the columns show the intended emotion, trial name and number, followed by the names of the acoustic features. self_report_questionnaire This pdf is the questionnaire which each musician was given following each trial relating to the emotions communicated and felt during each trial. Most* questions in the questionnaire were multiple-choice and speak pretty much for themselves. The answers for which were collated and are described below. self_report_questionnaire_anwsers This.xls file contains the results of all the musicians self-reported ratings from the above described questionnaire. Column name Description Subject The subject code of the musician, denoted by MS_ recording_filename The name of the trial denoted by MS_01_ followed by the trial number. intended_emotion The intended emotion the musician was instructed to communicate. The emotions are as follows: Angry Sad Relaxed Happy They are intended to be reported by using the valence-arousal space. valence_communicated The valence rating (integer between 1 and 5), that participants were asked to objectively rate how they thought the music played would be perceived. arousal_communicated As above, but relating to arousal rather than valence. valence_felt The valence rating (integer between 1 and 5), that participants were asked to rate how they felt while playing the piece. arousal_felt As above, but relating to arousal rather than valence. instrument_played The instrument played for each trial.
Treebanks annotated with Universal Dependencies (UD) are currently available for over 100 languages and are widely utilized by the community. However, their inherent characteristics are hard to measure and are only partially reflected in parser evaluations via accuracy metrics like LAS. In this study, we analyze a large subset of the UD treebanks using three recently proposed accuracy-free dataset analysis methods: dataset cartography, 𝒱-information, and minimum description length. Each method provides insights about UD treebanks that would remain undetected if only LAS was considered. Specifically, we identify a number of treebanks that, despite yielding high LAS, contain very little information that is usable by a parser to surpass what can be achieved by simple heuristics. Furthermore, we make note of several treebanks that score consistently low across numerous metrics, indicating a high degree of noise or annotation inconsistency present therein.
Music is capable of conveying many emotions. The level and type of emotion of the music perceived by a listener, however, is highly subjective. In this study, we present the Music Emotion Recognition with Profile information dataset (MERP). This database was collected through Amazon Mechanical Turk (MTurk) and features dynamical valence and arousal ratings of 54 selected full-length songs. The dataset contains music features, as well as user profile information of the annotators. The songs were selected from the Free Music Archive using an innovative method (a Triple Neural Network with the OpenSmile toolkit) to identify 50 songs with the most distinctive emotions. Specifically, the songs were chosen to fully cover the four quadrants of the valence arousal space. Four additional songs were selected from DEAM to act as a benchmark in this study and filter out low quality ratings. A total of 277 participants participated in annotating the dataset, and their demographic information, listening preferences, and musical background were recorded. We offer an extensive analysis of the resulting dataset, together with a baseline emotion prediction model based on a fully connected model and an LSTM model, for our newly proposed MERP dataset.
Collection of Ancient Greek annotated trees of Artemidorus' Oneirocritica Book 5. Part of the Open Projects in Digital Classics at the College of Letters and Sciences of the State University of São Paulo in Araraquara, São Paulo, Brazil. The trees were annotated manually on Perseids Platform using the Arethusa tool. The treebank tagset and guidelines used were those from The Ancient Greek Dependency Treebank with the morphological and syntactic layer, which was based on Bamman's and Crane's 2008 Guidelines for the Syntactic Annotation of the Ancient Greek Dependency Treebank (1.1). We translated it into Portuguese with a few additions from some specifications provided by a forum maintained in 2013 by Alpheios.net, which is no longer online. The trees are visible in Perseids Collection as UNESP-trees at https://perseids-publications.github.io/unesp-trees.