Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This article delves into the literary canon, a concept shaped by social biases and influenced by successive receptions. The canonization process is a multifaceted phenomenon, emerging from the intricate interplay of sociological, economic, and political factors. Our objective is to detect the underlying textual dynamics that grant certain works exceptional longevity while jeopardizing the transmission of the majority. Drawing on various criteria, we present an operational framework for defining the French literary canon, centered on its contemporary reception and emphasizing the role of institutions, particularly schools, in its formation. Leveraging natural language processing and machine learning techniques, we unveil an intrinsic norm inherent to the literary canon. Through statistical modeling, we achieve predictive outcomes with accuracy ranging from 70% to 74%, contingent on the chosen scale of canonicity. We believe that these findings detect what Charles Altieri calls a “cultural grammar”, referring to the idea that canonical works in literature serve as foundational texts that shape the norms, values, and conventions of a particular cultural tradition. We posit that this linguistic norm arises from biased latent selection mechanisms linked to the role of the educational system in the canon-formation process.
Humanlike androids can function as social agents in social situations and in experimental research. While some androids can imitate facial emotion expressions, it is unclear whether their expressions tap the same processing mechanisms utilized in human expression processing, for example configural processing. In this study, the effects of global inversion and asynchrony between facial features as configuration manipulations were compared in android and human dynamic emotion expressions. Seventy-five participants rated (1) angry and happy emotion recognition and (2) arousal and valence ratings of upright or inverted, synchronous or asynchronous, android or human agent dynamic emotion expressions. Asynchrony in dynamic expressions significantly decreased all ratings (except valence in angry expressions) in all human expressions, but did not affect android expressions. Inversion did not affect any measures regardless of agent type. These results suggest that dynamic facial expressions are processed in a synchrony-based configural manner for humans, but not for androids.
Abstract Proposals such as continuity and causality-by-default relate the level of expectedness of a relation to its linguistic marking as an explicit or implicit relation. We investigate these two proposals with regard to the English transcripts of six TED Talks and their Lithuanian, Portuguese and Turkish translations in the TED-Multilingual Discourse Bank (TED-MDB), annotated for discourse relations, following the Penn Discourse Treebank style of annotation. Our data shows that the discontinuous relations contrast and concession are indeed frequently explicit in all languages. But continuous relations show differences per relation and language. For instance, cause is frequently conveyed implicitly in English and Portuguese, but not in Lithuanian and Turkish. We explore temporal continuity by analysing whether the forward-order sense result is more frequently implicit than the backward-order reason. The hypothesis is confirmed by English and Portuguese, but not Lithuanian and Turkish. However, in Turkish, the arguments of the backward-order relation reason are frequently presented by the reversed order of arguments, retaining the linear order of events even in the presence of the connective. The causality-by-default hypothesis is not confirmed, as cause is not the most frequent implicit relation in the four languages.
This study explores the nuanced effects of social media on society, emphasising how it has both positive and negative aspects. Positively, social media has revolutionised global connectivity by democratising journalism, encouraging participation across great distances, and giving companies access to low-cost advertising channels. The report does admit many drawbacks, too, such as privacy issues, cyberbullying, and the quick dissemination of false information. The study promotes digital literacy, user education, and proactive actions from social media companies to address these problems. It highlights how critical thinking abilities are necessary to successfully traverse the internet environment. Furthermore, the study draws attention to the linguistic influence of social media by presenting acronyms, abbreviations, and colloquial language, prompting concerns about possible negative effects on written language proficiency and the significance of maintaining linguistic norms. Overall, the study highlights how social media has a significant impact on a variety of fields, including activism, politics, marketing, and education. It also highlights the need for a balanced strategy to maximise social media’s advantages while minimising its drawbacks
The Covid-19 pandemic in the last 3 years has strengthened E-commerce growth, making online shopping the new norm due to restricted offline activities. To aid buyers and sellers in conducting transactions in E-commerce, E-commerce platforms have introduced features like product descriptions, product photos, ratings, and reviews. These features have created a competitive landscape, benefiting sellers who can use them effectively. Nevertheless, many sellers still struggle to optimize these features and market their products effectively to buyers. Failure to optimize these features correctly restricts the marketing strategy's effectiveness and puts sellers at risk for unanticipated difficulties that may be prevented by determining the various effects of these features on customers’ purchase intentions. Therefore, this research aims to analyze the impact of product descriptions, product photos, and ratings & reviews on customers' purchase intention in E-commerce. A quantitative approach is used in this study, where the data is analyzed through descriptive statistics and PLS-SEM. The result of this study suggested that all three features of product description, product photo, and rating & review significantly and positively influence purchase intention in E-commerce. In addition, the author also found that moderation of perceived trust significantly affects product description and rating & review on purchase intention, while the moderation of perceived risk only significantly affects rating & review on purchase intention. The finding of this research is expected to give insights to E-commerce sellers on optimizing the features in E-commerce to increase the customers’ purchase intention.
In this article, I engage in a discussion of the approaches of the normativists postulating the absence of a linguistic norm and assuming the identification of the norm with usus. I have made an attempt to prove the existence of the norm as a level of internal organisation of language, based on linguistic and mathematical-computational considerations. I accept the need for separate models of language: linguistic and mathematical, which serve different purposes and have different properties. I ponder the dual nature of language manifested by its finite infinity, assuming that only the usus is infinite. I postulate the adoption of a theory (formulated by K. Kłosińska) assuming the existence of an invariant norm together with ‘allonorms’. I also propose the introduction of the probabilistic Gaussian model into the codification procedures of the linguistic norm as a method objectifying the procedure and removing the (qualitative and quantitative) arbitrariness of codifiers.
Abstract This paper deals with several aspects of context in lexicography. Section 1 briefly mentions some different approaches to the concept context in various fields. Section 2 puts the focus on different uses and perceptions of the concept context in lexicography, contrasting it with related concepts, such as cotext, contextualization and contextual information. A more comprehensive discussion also covers different aspects of the occurrence of the concept context in dictionary research, with specific reference to central aspects of the so-called inner and outer context. Various portals, dictionaries and dictionary entries will illustrate the above-mentioned approaches. Section 3 approaches the subject from a user perspective. Section 4 addresses the question How can contextual data be extracted or generated? To answer this question, some methods and tools for (automatic) acquisition and analysis of contextual data, – in particular of the local contextual data in terms of Faber and León-Araúz (2016) – are introduced. Examples of these are lexical databases or semantic networks, like WordNet, and corpora, like Sketch Engine, or predictive methods, like Word2vec and similar ones. Some advantages and disadvantages of specific data acquisition tools used for the analysis of local contextual data are indicated. This section also contributes to a more detailed discussion of the automatic generation of the so-called local syntactic-semantic context or word environment, specifically of the building of syntactic-semantic argument patterns and their examples.
The current literature suggests that some women are uniquely vulnerable to negative effects of hormonal contraception (HC) on affective processes. However, little data exists as to which factors contribute to such vulnerability. The present study evaluated the impact of prepubertal adverse childhood experiences (ACEs) on reward processing in women taking HC (N = 541) compared to naturally cycling women (N = 488). Participants completed an online survey assessing current and past HC use and exposure to 10 different adverse childhood experiences (ACEs) before puberty (ACE Questionnaire), with participants categorized into groups of low (0-1) versus high (≥2) prepubertal ACE exposure. Participants then completed a reward task rating their expected and experienced valence for images that were either erotic, pleasant (non-erotic), or neutral. Significant interactions emerged between prepubertal ACE exposure and HC use on expected (p = 0.028) and experienced (p = 0.025) valence ratings of erotic images but not pleasant or neutral images. Importantly, follow-up analyses considering whether women experienced HC-induced decreases in sexual desire informed the significant interaction for expected valence ratings of erotic images. For current HC users, prepubertal ACEs interacted with HC-induced decreased sexual desire (p = 0.008), such that high ACE women reporting decreased sexual desire on HC showed substantially decreased ratings for anticipated erotic images compared to both high prepubertal ACE women without decreased sexual desire (p < 0.001) and low prepubertal ACE women also reporting decreased sexual desire (p = 0.010). The interaction was not significant in naturally cycling women reporting previous HC use, suggesting that current HC use could be impacting anticipatory reward processing of sexual stimuli among certain women (e.g., high prepubertal ACE women reporting HC-induced decreases in sexual desire). The study provides rationale for future randomized, controlled trials to account for prepubertal ACE exposure to promote contraceptive selection informed by behavioral evidence.
This dataset contains EEG, ECG and audio recordings of 11 individual musicians playing emotional music on their instrument. The dataset consists of two parts: Experiment: Musicians’ self-reported ratings, audio recordings, and physiological recordings where the 11 expert musicians were asked to play at least 4~2-minute unfamiliar (non-popularly known) musical pieces. Participants were asked to play at least once one of the following emotions: happiness, sadness, relaxation, and anger. For the physiological recordings EEG, ECG, and GSR signals were recorded. Each musican’s data is denoted by MS_ followed by the order in which they were recorded. Self-report questionnaire: A self-assessment questionnaire and their answers where 11 expert musicians were asked to rate musical pieces recorded based on: Objective valence & arousal they felt the piece had; Felt valence & arousal during playing. For a more detailed explanation of the dataset, its recording procedure, and its contents, see L. Turchet, B. O'Sullivan, R. Ortner & C. Gugher (2024). Emotion Recognition of Playing Musicians from EEG, ECG, and Acoustic Signals. IEEE Transactions on Human-Machine Systems. File Listing The following files are available (each explained in more detail below): Name Format Contents EEG_ECG_data_for_each_musician mat This folder contains the raw EEG, ECG, & GSR data for all 11 musicians for each piece they played as well as a resting state recording, which was recorded while a neutral audio stimulus was played. audio_data_for_each_musician wav, JSON, csv This folder contains 3 subfolders: 1) wav_audio_files_original: This is the raw audio data recorded for each musician and includes all 56 pieces included in the paper reported above, please see below regarding rejected trials. 2) wav_audio_files (original split_into_3_parts): Here, the 56 pieces are appropriately split into 3 separate parts.3) analysis_audio_files: This folder contains the acoustic features extracted, 1714 acoustic features were extracted from each split trial. Each result is stored in.JSON, however, the collated results can be seen in all_results.csv. self_report_questionnaire pdf, xls Two files exist in this folder:1. The self-reported questionnaire given to the musicians of the questionnaire during the experiment.2. File Details EEG_ECG_data_for_each_musician These are the original raw data recordings. EEG data were recorded using a g.GAMMAcap2 by g.tec Medical Engineering, a 64-channel cap with g.SCARABEO active electrodes, with two g.GAMMAsys reference active ear clip electrodes. Two g.GAMMAbox electrode connector boxes were used to connect the active electrodes to two g.USBamp biosignal amplifiers with a sampling frequency of 256 Hz.The following 31 EEG channels were used: Fp1, Fp2, AFz, AF3, AF4, AF7, AF8, Fz, F3, F4, F7, F8, Cz, C3, C4, CP3, CP4, CP5, CP6, P1, P2, P3, P4, P5, P6, P7, P8, PO7, PO8, O1, and O2. AFz was used as a ground electrode and Cz was used as a re-reference electrode. The right-side ear clip electrode was used as a reference electrode. ECG data were recorded using a single g.GAMMAclip active electrode clip connected directly to the g.GAMMAbox, sharing the same ground electrode with the EEG cap and placed on position V4 of the subjects.GSR data was recorded using the g.GSRsensor² box which contains two small dry electrodes placed underneath the participant’s toes (due to the amount of hand movement required for playing an instrument). The g.GSRsensor² was connected directly to a g.GAMMAbox, using jumper cables connected to different ports in the g.USBAMPs to share the same reference and ground electrodes as the EEG cap. The locations of the channels and their corresponding number in the raw data is as follows: Channel Number Channel Name 1 Time series 2 AF3 3 AF4 4 AF7 5 AF8 6 CP3 7 CP4 8 CP5 9 CP6 10 P1 11 P2 12 P5 13 P6 14 P7 15 P8 16 O1 17 O2 18 Cz 19 Fp1 20 Fp2 21 F3 22 F4 23 F7 24 F8 25 C3 26 C4 27 P3 28 P4 29 PO7 30 PO8 31 AFz 32 GSR 33 ECG audio_data_for_each_musician/wav_audio_files_original This folder contains all of the raw audio files recorded during the experiment. All pieces were recorded using the software Audacity and exported as WAV files encoded with a bit depth of 32-bits and a sampling rate of 44.1 kHz. 56 of the included trials are present in this folder. audio_data_for_each_musician/wav_audio_files (original split_into_3_parts) This folder contains the above mentioned 56 raw audio pieces separated into 3 appropriately sized recordings. audio_data_for_each_musician/analysis_audio_filesThis folder contains the 1714 acoustic features from each of the 3 separated trials from 56 accepted pieces, these are denoted by MS_ followed by the order in which the musicians were recorded and the order in which the trails were split, the individual features are in JSON format. Within this folder, we have collated all results in a csv file (all_results.csv) where the columns show the intended emotion, trial name and number, followed by the names of the acoustic features. self_report_questionnaire This pdf is the questionnaire which each musician was given following each trial relating to the emotions communicated and felt during each trial. Most* questions in the questionnaire were multiple-choice and speak pretty much for themselves. The answers for which were collated and are described below. self_report_questionnaire_anwsers This.xls file contains the results of all the musicians self-reported ratings from the above described questionnaire. Column name Description Subject The subject code of the musician, denoted by MS_ recording_filename The name of the trial denoted by MS_01_ followed by the trial number. intended_emotion The intended emotion the musician was instructed to communicate. The emotions are as follows: Angry Sad Relaxed Happy They are intended to be reported by using the valence-arousal space. valence_communicated The valence rating (integer between 1 and 5), that participants were asked to objectively rate how they thought the music played would be perceived. arousal_communicated As above, but relating to arousal rather than valence. valence_felt The valence rating (integer between 1 and 5), that participants were asked to rate how they felt while playing the piece. arousal_felt As above, but relating to arousal rather than valence. instrument_played The instrument played for each trial.
Abstract Introduction During laboratory-based, voluntary exposure to total sleep deprivation (TSD), relatively large decreases in positive mood and relatively small increases in negative mood have been observed, with the magnitudes of change varying across individuals. People differ in their trait emotion reactivity, i.e., their sensitivity to emotion-evoking stimuli or events. As higher reactivity is associated with greater difficulty regulating emotion, we hypothesized this trait could moderate mood changes during TSD. Methods N=96 healthy adults (ages 21-38, 47% female) participated in one of three 4-day/3-night in-laboratory TSD studies. In each study, after 10h baseline sleep, participants were exposed to 38h TSD, followed by 10h recovery sleep. Participants completed the Emotion Reactivity Scale (ERS), and their self-reported affect was assessed every 2-4h during wake using the Positive and Negative Affect Schedule (PANAS). The 11 test bouts during TSD shared across studies were included in analyses (day 2: 09:00, 13:00, 21:00, 23:00; day 3: 01:00, 03:00, 05:00, 07:00, 09:00, 13:00, 21:00). Positive and negative affect ratings were analyzed with linear mixed-effects regression with fixed effects of test bout (categorical), ERS score (continuous), and their interaction, covariates for study, sex, and age, and a random intercept over participants. Results Participants showed the expected decrease in positive affect as a function of test bout (time awake and time of day, p&lt; 0.001), and a smaller, non-significant increase in negative affect (p=0.17). Higher ERS scores were associated with increased negative affect during TSD (p=0.002), but this effect was small, and negative affect showed a floor effect. Emotion reactivity was not a significant predictor of decreased positive affect during TSD (p=0.97). Conclusion As hypothesized, emotion reactivity predicted individual variability in negative affect during TSD, but unexpectedly it did not predict the decrease in positive affect associated with TSD. This suggests that the trait measured by the ERS may be biased toward negative affect, and/or that positive affective changes may be more difficult to predict during the low arousal state induced by TSD. Studies inducing greater variability in negative affect during TSD (e.g., through exposure to a stressor) are needed to confirm our findings. Support (if any) NIH CA167691, ONR N00014-13-1-0302, and CDMRP W81XWH-16-1-0319 and W81XWH-20-1-0442.
Abstract Introduction “Sleep to forget and sleep to remember” model postulates that the content of emotional memory is strengthened during sleep while the arousal component is attenuated. Experimental sleep deprivation studies in healthy individuals have supported the model, but it is not known if sleep in individuals with insomnia, a condition of chronic poor sleep, provides the same effect on emotional memory. Methods Individuals reporting insomnia and good sleepers (n = 76 and n = 83 respectively; 83% female, mean age of 43.2 years) completed a picture-rating task to assess a) emotional valence ratings pre and post sleep and b) recognition accuracy post-sleep. In addition, an autobiographical memory task was completed during the day to assess a) the number of recalled memories in a 2-minute period for good and bad past events respectively, and b) emotional intensity ratings of the retrieved memories at the time the event occurred and how they feel now. Results Wilcoxon rank-sum test was used to assess group differences. Sleep diary variables on the experimental night showed that individuals with insomnia experienced a significantly more disrupted sleep than good sleepers (SOL: r = 0.537; WASO: r = 0.620; TST: r = 0.669). In the picture-rating task, there was no significant group difference for recognition accuracy (neutral stimuli: r = 0.016; negative stimuli: r = 0.018) or change in emotional valence pre-to-post sleep (neutral stimuli: r = 0.003; negative stimuli: r = 0.008). In the autobiographical memory task, individuals with insomnia remembered significantly fewer good days (r = 0.245) and rated their bad days at the time of its occurrence as significantly worse than good sleepers (r = 0.235). There was no significant group difference in the extent to which emotional intensity faded more for bad days than good days (r = 0.113). Conclusion We have shown that autobiographical memory is altered in insomnia which may have important influence on mental health. Our findings also suggest that overnight memory consolidation for negative stimuli is not affected in insomnia versus good sleepers. This may be a reflection of our one-night protocol; future work is needed across consecutive days. Support (if any) Dr-Mortimer-&-Theresa-Sackler-Foundation
There is mounting evidence suggesting that the effectiveness of positive psychological interventions can be influenced by a variety of factors, including one’s cultural context. Identifying target variables that can effectively serve to improve individual well-being under these boundary conditions is a crucial step when developing viable interventions. To this end, we examined how gratitude disposition, self-esteem, and optimism relate to the subjective (SWB) and psychological well-being (PWB) of Japanese individuals. Multivariate regression analysis revealed that while self-esteem was predominantly more strongly associated with SWB compared to gratitude disposition, the latter was more strongly associated with the PWB dimensions, particularly personal growth, positive relations with others and purpose in life. These results were largely corroborated by a second expanded dataset. In addition, we used experience sampling over a four-week period to examine how momentary affect ratings related to overall daily evaluations. We found that both gratitude disposition and self-esteem moderated the association between momentary positive affect and “good day” evaluations but in opposite ways; increasing gratitude disposition strengthened the association, while increasing self-esteem weakened it. All in all, the current results suggest that while gratitude, self-esteem, and optimism influence individual well-being as a whole, they each likely play distinct roles as enablers of SWB and PWB in the examined cohort.
Goal pursuits have long been associated with how good or bad people feel. The exact affective influence of approaching the end of a deadline-bound task, however, remains unclear. People's effort level usually increases as they approach a deadline, yet while some theories predict this pattern will correspond to an increase in positive affective state, others may suggest that affective state will follow an opposite trend. Here, participants (n = 22; 14 women) performed a deadline-bound spaceship game paradigm designed to measure U-shape patterns in effort allocation by the frequency of keypresses. Participants vocally rated how they feel on a scale ranging from -5 ("very bad") to +5 (very good) every 30 seconds through the game. We found that both effort level and affective ratings followed a U-shape pattern through the task. This result supports an opportunity cost model of goal gradients which suggests that people postpone competing activities in favor of performing a focal task when approaching a deadline.
MOBA as one of the many popular subgenres today, Mobile Legend, Arena of Valor, League of Legend and Lokapala based on the large number of downloads can be mentioned as popular MOBA games, but the rating of these four applications is below 4.0 on the Google Play platform Store, this happens because some users may think that this mobile MOBA game has several advantages, but also some disadvantages that affect ratings. This study aims to determine the results of sentiment analysis on mobile MOBA games using Google Play Store reviews. Then it is processed using Python programming to create a model with the linear kernel Support Vector Machine (SVM) algorithm to classify the dataset. From the results of the classification model test using 19,579 data, where there were 10,017 positive sentiment data and 9,562 negative sentiment data and the distribution of train data and test data was 70%: 30%, obtained an accuracy of 82.64% and then re-evaluated using the Cross Validation method using 5 times so that an accuracy of 83.38% is obtained Keywords: MOBA, Sentiment Analysis, Machine Learning, SVM, Cross Validation
The article focuses on the analysis of anglicisms in the professional speech of language teachers on the basis of educational and methodological materials posted on the popular educational platforms “Vseosvita” and “Na urok”. The main focus is on the compounds with the element “бук”, as the number of such titles has increased significantly in recent years. The speech of teachers of language and literature abounds in the following anglicisms: буктрейлер, буккросинг, артбук, воркбук, лепбук, скрапбук. It is noteworthy that most of these new words are recorded only in the Dictionary of Modern Anglicisms, which was published in 2022. Other lexicographical sources do not list them, and also do not contain the component “бук” either as part of other compounds or as an independent word. Anglosurzhyk words appear in teachers’ speech due to various factors: learning from the professional experience of fellow teachers from English-speaking countries, using working methods from such fields as economics, business and management in the educational process, translation difficulties, the desire to make it concise, the desire to embellish their speech with unusual words and evoke interest of the students. Functioning in speech, these new words play the role of professionalisms. The latter are known for their stylistically colouring, they exist beyond the linguistic norm and may disappear over time. However, the widespread usage of these names in educational and methodological works indicates the desire to make these words normative terms. Compound words of English origin with the component “бук” used by teachers of language and literature do not denote new concepts, and therefore need to be replaced by those with more transparent semantics. The teachers frequently use foreign words to denote concepts that already have or may have Ukrainian names. Transliteration of English words does not enrich our language, but contaminates it and leads to the formation of another type of language mix – anglosurzhyk.
I delivered (24.03.2026) a seminar to over 100 teachers from Greece’s Model–Experimental Schools, introducing Helex Kids 2 alongside the Word Tool a structured lexical database I developed to support learners with reading and spelling difficulties. The session focused on translating research into classroom practice, equipping educators with evidence-informed strategies and practical digital resources to enhance literacy instruction. The strong engagement and discussion that followed highlighted both the demand for targeted literacy support tools and the value of bridging innovation with everyday teaching practice.<br/><br/>Feedback received: Participant feedback highlighted the clarity and accessibility of the seminar, with teachers noting the use of simple, comprehensible language, concise and well-structured delivery, and clear explanations of complex concepts such as neuronal processes in spelling and the involvement of multiple brain centres in reading and writing; they particularly valued how the session illuminated the difficulties faced by children while addressing a wide range of issues in a focused and highly effective way.<br/>Professor Anthi Revithiadou (Linguistics, Aristotle University of Thessalonica) and event coordinator described the seminar as “outstanding,” adding, “I truly have no words to thank you, excellent in every respect. I say this with complete sincerity. Thank you once again.”<br/>
We propose a method for the classification of objects that are structured as random trees. Our aim is to model a distribution over the node label assignments in settings where the tree data structure is associated with node attributes (typically high dimensional embeddings). The tree topology is not predetermined and none of the label assignments are present during inference. Other methods that produce a distribution over node label assignment in trees (or more generally in graphs) either assume conditional independence of the label assignment, operate on a fixed graph topology, or require part of the node labels to be observed. Our method defines a Markov Network with the corresponding topology of the random tree and an associated Gibbs distribution. We parameterize the Gibbs distribution with a Graph Neural Network that operates on the random tree and the node embeddings. This allows us to estimate the likelihood of node assignments for a given random tree and use MCMC to sample from the distribution of node assignments. We evaluate our method on the tasks of node classification in trees on the Stanford Sentiment Treebank dataset. Our method outperforms the baselines on this dataset, demonstrating its effectiveness for modeling joint distributions of node labels in random trees.
Among the goals of digital transformation, there is a transformation of the model from a traditional one into a model of continuing education, and under these circumstances, school becomes the stage where, among others, the basic skill of a modern person is formed — the ability to learn. Teaching to learn is one of the most important ideas of current federal state educational standards. The authors propose to consider a modern foreign language lesson in the unity of aspects-components: communicative-interactive, logistical, procedural, instrumental. Each of the conditional components, representing an independent whole, is determined by and determines itself the normal functioning of content and activity, internal and external entities, as regards the lesson as a form of organization of the educational process at school. It is suggested that the complete replacement of teaching tools and means with digital analogues is impractical, due to the need to take into account certain patterns of mastering linguistic norms, a foreign language communication, as well as compliance with the health care rules for children and adolescents – students of educational organizations implementing general education programs in the country. The purpose of the study is to reveal the prospects of digitalization within the framework of the modification of components-aspects of a foreign language lesson, as well as to analyze the basic principles and conditions required for the transition to one of the levels of digital transformation of education as proposed by modern researchers. The authors come to conclusion about the need of a multi-level approach to change, to digitalization of each of the components-aspects of the lesson, about the potential choice of the desired depth and intensity of digital transformations in each case.
A method for predicting adversarial instances in a language includes producing graphs for each sentence that capture the structure. The graph comprises nodes and edges, a node label represents a perturbed word, and the edge labeled notes possible typos of perturbations. Language labels are predicted for the unlabeled movie review which involves propagating labels through the graph. The adversarial instances appear nearly identical to the given input, but their equivalent prediction can be misinterpreted. In the past few years, assessing adversarial instances has become a decorum to estimate the robustness of deep learning models. While most of the literature emphasizes continuous features (i.e., image). However, some deep learning models also became successful on discrete features, and thus demand a model to understand underlying vulnerability and robustness. In this study, all the one-edit perturbations of a specific word in a review sentence of the Stanford Sentiment Treebank textual dataset which swaps the sentiment from positive to negative are analysed. Classifier's accuracy is also observed by inducing different typos and validated results indicate that graph structure alone is sufficient to predict whether an instance can fool the classifier or not.
Abstract In this chapter, we describe how we incorporate the Prosodic Model ( Brentari 1998 ), especially the concept of a hierarchy of nodes with features, into a template for feature coding of sign entries in a lexical database, and how we make use of the features to organize learner’s dictionaries of the respective signed languages. Besides coding prosodic features, other information recorded includes sign category (monomorphemic or polymorphemic), sign type (whether the sign is one-handed or two-handed), the country of origin, gloss and grammatical category. Beginning with Hong Kong Sign Language, Asia SignBank has archived signs from Indonesian Sign Languages, Sri Lanka Sign Languages, Japanese Sign Languages, Ho Chi Ming Sign Languages, and Myanmar Sign Languages. Data input and feature coding were done by Deaf signers trained in sign language analysis and dictionary compilation. Archiving phonological features of individual sign entries also facilitates identification and comparison of phonological and morphological features of signs.
Purpose The purpose of this paper is to describe a new approach to sentence representation learning leading to text classification using Bidirectional Encoder Representations from Transformers (BERT) embeddings. This work proposes a novel BERT-convolutional neural network (CNN)-based model for sentence representation learning and text classification. The proposed model can be used by industries that work in the area of classification of similarity scores between the texts and sentiments and opinion analysis. Design/methodology/approach The approach developed is based on the use of the BERT model to provide distinct features from its transformer encoder layers to the CNNs to achieve multi-layer feature fusion. To achieve multi-layer feature fusion, the distinct feature vectors of the last three layers of the BERT are passed to three separate CNN layers to generate a rich feature representation that can be used for extracting the keywords in the sentences. For sentence representation learning and text classification, the proposed model is trained and tested on the Stanford Sentiment Treebank-2 (SST-2) data set for sentiment analysis and the Quora Question Pair (QQP) data set for sentence classification. To obtain benchmark results, a selective training approach has been applied with the proposed model. Findings On the SST-2 data set, the proposed model achieved an accuracy of 92.90%, whereas, on the QQP data set, it achieved an accuracy of 91.51%. For other evaluation metrics such as precision, recall and F1 Score, the results obtained are overwhelming. The results with the proposed model are 1.17%–1.2% better as compared to the original BERT model on the SST-2 and QQP data sets. Originality/value The novelty of the proposed model lies in the multi-layer feature fusion between the last three layers of the BERT model with CNN layers and the selective training approach based on gated pruning to achieve benchmark results.
A key method for understanding the evolution of languages is to look for words which share etymological roots across languages. These words, called lexical cognates, allow historical linguists to group languages into families and study their structural history. Existing research on cognate clustering is chiefly based on analyses performed on lexical databases which limits the scope for exploring new cognates. Additionally, while the Indo-European family is a transcontinental one, its cognate analyses are largely limited to a number of European languages. In this research work we build a new dataset using word embeddings to search for cognates using phonetic matching between translations of words in different languages across the Indo-European family. The context-dependent positioning of words in word embeddings allows for comparisons with contextually similar words and hence clustering using unsupervised learning algorithms. We successfully find significant distinctions between these “contextually close clusters of cognates” among languages across the Indo-Iranian, Romance, and Germanic families. We release this novel method to demonstrate which words a language is more likely to lend to another.
Abstract The labor market is a key part of an economy. Several existing online platforms allow the upload of resumes and the search for a job. One of their limitations, however, is that obtaining the best opportunity can be hard because certain jobs need some experiences, abilities, and features that an applicant might not know. The recent diffusion and employment of conversational agents definitely have proven to benefit this kind of issue. For example, ChatGPT has shown impressive outcomes in different domains and for a variety of tasks. It has weaknesses, although, related to the veracity of the responses it generates, which might deceive the user interacting with it. The usage of external domain knowledge is the direction we suggest in this chapter. Several lexical databases and taxonomies have already been collected and designed by different organizations. We illustrate a list of such resources and provide a solution that integrates conversational agents with relevant information extracted from one of such resources showing the benefits and the impact that our proposal can generate.
The internal representations of large language models (LLMs) remain largely opaque, hindering interpretability, alignment, and cross-lingual transfer. A core obstacle is polysemanticity—the phenomenon whereby individual neurons activate for semantically unrelated concepts—which arises from the superposition of more features than there are physical dimensions. Existing approaches, such as Sparse Autoencoders (SAEs), address this by decomposing embeddings into millions of monosemantic features, yet they recover no macroscopic coordinate system that organizes these features. In this work, we introduce the Atlas Autoencoder (Atlas AE), a topology-preserving autoencoder trained on approximately 100,000 English dictionary definitions, which compresses 768-dimensional transformer embeddings onto a 128-dimensional latent manifold. Linear probing within this latent space recovers eleven interpretable cognitive axes—Concreteness, Agency, Sentiment, Intentionality, Sociality, Power, Temporality, Dynamicity, Negation, Quantity, and Perceptuality—each grounded in established psycholinguistic norms (AUROC 0.75–0.96; all permutation ). Residual subspace analysis confirms that the semantic space saturates at exactly eleven stable dimensions. To assess universality, we project embeddings from five independent transformer models spanning five typologically distinct language families—Indo-European (English, French), Uralic (Finnish), Altaic (Turkish), and Japonic (Japanese)—through the same frozen, English-trained manifold. The resulting coordinate systems exhibit striking geometric isomorphism: core structural axes such as Power and Agency show cross-model variance, and all five languages agree on directional polarity for nine of eleven axes. Inter-axis correlation analysis reveals that the recovered basis is structurally oblique, with stable covariances (e.g., Power–Sociality ) that challenge the widespread orthogonality assumption in representation learning. These findings establish the Universal Semantic Manifold as a compact, interpretable coordinate system that bridges connectionist representations and symbolic cognition, offering a rigorous geometric framework for cross-lingual alignment, controlled semantic steering, and the quantitative study of conceptual structure.
This article in addition to introducing and defining taboos, examines the existing and well-known taboos in the collection of short stories” To whom should I greet” by the famous Iranian novelist SiminDaneshvar, during which, it refers to the use of taboo words by the characters of gender(female/male) in social situations. The authors have tried to include social behaviors such as: good/bad, holy/ unholy, polite/ impolite, etc. in social relations and dialogues of the work, based on the accepted norms of the Iranian society in front of the readers. In adition to be able to explain the relationship between culture and language and the intractions between the two in terms of prohibition in the sociology of language based on gender differences and how they are used. Besides, to prove that in this work, Simin while paying attention to the values of the Persian society, has consciously tried to break the linguistic norms in the daily individual and social life of the characters of some of stories in many cases. This study also tried to show the results obtained from the frequency of using taboos and inconveniences in the actions and speech of speakers by gender in the form of a graph at the end of the article and between gender and taboos in terms of their use in speakers there was a significant relationship.
Abstract This chapter examines issues that arise in speech communities in response to changes in the language. Some usages are more appropriate for formal contexts than for informal ones, while usages out of place in formal contexts may be perfectly fine in informal ones. Schooling tends to accord formal styles a privileged status over others, while a great deal of change comes from the language spoken more casually. While the singular noun English encourages us to view the language as a single, unified speech form, this has never been so. A standard language is a set of linguistic norms established by some generally accepted political or social authority. But different speech communities can have different standards, leading to greater variety across English dialects. Furthermore, each dialect by itself encompasses a variety of “standards,” depending on whether we are speaking or writing and to whom. Writing is generally more conservative, while oral language changes far more readily—to the dismay of those who prefer to leave things as they are.
Individuals with depression experience more negative imagery and less vivid positive imagery, and the late positive potential (LPP) is considered as a viable biomarker for negative attentional and memory biases in depression; however, the LPP response to emotional imagery in depressed individuals remains unclear. This study aims to investigate neural response to emotional imagery in depressed individuals. ERPs were recorded from 40 depressed participants and 44 healthy controls during the encoding-imagery task. Depressed participants scored significantly lower in the valence rating of sad and neutral imagery compared to healthy participants. Importantly, the LPP amplitudes to sad imagery in depressed participants were significantly larger than healthy controls, particularly in the middle (800-1,400 ms) and late time windows(1,400-2,000 ms). Furthermore, depressed individuals exhibited significantly higher LPP amplitudes for sad imagery compared to happy imagery, whereas healthy participants showed the opposite pattern. The present study provides evidence that depressed individuals display abnormal electrophysiological reactivity to sad imagery, which offers a new perspective for understanding the mechanisms underlying depression.
Abstract In the contribution, we provide a theory-based and corpus-verified description of expressions for measure in Czech. We demonstrate that the measure expressions may modify quantity of entities ( approximately ten boys ), internal characteristics of events ( he works a lot ), properties ( very big ) and relations ( completely without sound ). We distinguish between the measure expressions that are an answer to the question To what extent? (Extent-modifiers) and expressions that modify an answer to the question How many? (Quantity-modifiers). The Extent-modifiers are formally, structurally and semantically more diverse than the Quantity-modifiers. For the Quantity-modifiers a list of forms and functions is provided. Theoretical knowledge stemming from the analysis will subsequently be used to improve the annotation in the Prague Dependency Treebanks. It can be also useful for other semantically-oriented descriptions of language.
Sociolects, as social varieties of language, should be classified and described on the basis of both linguistic and sociological criteria. In practice, however, most often only either the first one (e.g. according to Antoni Furdal and Danuta Buttler) or the second one is used (e.g. according to Aleksander Wilkoń and Tomasz Piekot). Both these criteria are less frequently combined (Stanisław Grabias). The article proposes those aspects of the sociological and linguistic functioning of language varieties that should be used by sociolinguistics especially in the characterisation of sociolects, but also in their typology. The sociological criteria include: the type of contacts in the group (including contacts made via the Internet), the degree of group formalisation, the durability of the group and the type of bond that binds the group together. Among the linguistic criteria the following were distinguished: nomination and expressiveness, the method of acquiring a sociolect and the degree of its codification and the rank of the linguistic norm. The article also clarifies the concept of the distinguishing features of the sociolectic vocabulary.
Focusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet.In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g.elements of wordnet taxonomy, quantifier phrases, certain collocations).In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches.We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one).Language model also proves to be better than a feature-based logistic regression model.
The paper examines team building in multicultural and multiethnic work environments – particularly in NGOs – by linking classic accounts of effective teams with sociolinguistic and socio-emotional perspectives. It conceptualises workplace teams as speech communities in which members share, negotiate, and sometimes contest linguistic norms, expectations, and emotional display rules. Drawing on management and organisational studies, the text outlines key structural conditions for effective teams, including clear goals, appropriate leadership, resource allocation, and mechanisms for accountability and cooperation. These insights are then integrated with research on emotions, culture, and communication, which highlights the role of emotional attachment, trust, and creative problem-solving in sustaining team cohesion. The paper argues that effective team building in such contexts requires not only formal structures but also deliberate cultivation of shared communicative practices and affective climates that support participation, innovation, and mutual understanding across cultural and linguistic differences.
As the most widely studied and spoken language worldwide, English is a medium of communication between speakers of varied language backgrounds. Global English users need not converge on any one variety; rather, they need adaptive, flexible skills that support communication with speakers of diverse World English (WE) varieties (Kirkpartick 2007). Nevertheless, many English learners and teachers around the world continue to adhere to ‘native speaker’1 models that prioritize language varieties from English-dominant nations, such as the United States or the United Kingdom (Tseng 2019). We have encountered this perspective first-hand in our work in Indonesia. Hanung has taught English and trained teachers at an Islamic University in Central Java for over twenty years. Tabitha has nearly twenty years of experience in ELT, including three years as a visiting instructor at Hanung’s institution. Among the students and teachers we have worked with in Central Java, the dominant language learning model is the ‘native English speaker’. For instance, on a survey given to our first-year English-major students, 81 percent reported wanting to ‘sound like a native speaker’. To challenge this tendency, Simanjuntak and Lien (2021) encourage the use of materials from global and local language communities. Using diverse listening materials can help students understand a wide variety of language varieties and accents and adapt to dialects they have never encountered before. We found podcasts, free audio recordings which are automatically delivered to users’ devices, to be a particularly rich and convenient source of listening materials featuring WE speakers. In the sections below, we describe our teaching context and use of podcasts. We close by discussing how our approach could transfer to other contexts where students and teachers continue to aspire to ‘native speaker’ models, and offer recommendations for teachers interested in shifting to a WE approach and exposing students to varied linguistic norms through podcasts. We collaborated to design lessons for ‘Listening for General Communication’, a course for first-year English majors. The students met weekly from September to December, 2021. Because of COVID-19, the course was taught asynchronously online, with the exception of in-person sessions in weeks 13 and 14. Podcasts were well suited to asynchronous online teaching because their easy accessibility facilitated students’ independent listening. Teachers can create their own podcasts (for a discussion of teacher-created podcasts, see Ingham 2022), but we found several high-quality podcasts online which featured WE speakers. Class activities revolved around seven episodes of the 22.33 podcast, which features participants in exchange programmes sponsored by the US Department of State. We used this podcast because it features speakers with diverse accents, language use, and perspectives, and focuses on intercultural encounters. With the wide variety of podcasts available online, teachers can find podcasts featuring diverse speakers which match their course objectives, contexts, and aims. In Indonesia, English is compulsory at the secondary and tertiary levels, so our students had studied English for approximately six years. Many students at Islamic universities such as ours come from secondary schools whose curriculum focuses primarily on the teaching of Islam. English instruction at these schools focuses on receptive skills, to allow students to access texts about Islam and Muslims written for international audiences. Students typically also study Arabic, but have stronger English skills due to English’s status as the first foreign language in Indonesia, and its use in mass media. The language model used in that media is most commonly standard American English. Students in our English Department take courses in language, linguistics, and teaching methods, and most enter with language skills at the B1 or B2 level. After graduation, most students gain employment as English teachers. At the beginning of the semester, we asked students to reflect upon their future use of English. Although most hoped their own language use would come to match ‘native speaker’ norms, students also acknowledged that they would probably use English to communicate with people using a variety of language norms. They expressed interest in listening to WE speakers and learning about people from around the world. We introduced the podcasts and explained that they featured WE speakers, thereby exposing students to diverse language varieties. In the first asynchronous session for each podcast, students completed pre-listening activities. Podcasts featuring WE speakers typically discussed contexts unfamiliar to our students, so these activities were essential to support comprehension. When students had some relevant prior knowledge, we activated that knowledge through reflection, opinion sharing, and hypothetical situations. When significant content was new to students, short readings, visuals, and online research helped build background knowledge. Taking the Observing Ramadan podcast as an example, students shared their own memories of observing Ramadan, considered which aspects of Ramadan might be surprising to a non-Muslim, and looked up new vocabulary, such as ‘muscle through’ and ‘famished’. Students were expected to complete while-listening activities between the two sessions. A benefit of podcasts is that students can adjust the audio speed and listen to the podcast repeatedly, if needed. To further support students’ comprehension of the WE speakers, we offered instructional scaffolds. For some lessons, students received a graphic organizer to track the various voices and topics. In other lessons, questions were divided into sections corresponding to timestamps in the podcast, helping students know when to listen intensively. For Observing Ramadan, students completed a chart with the names and nationalities of the six speakers by entering each speakers’ perspective on challenges related to fasting and non-Muslims’ awareness of Ramadan. In the second asynchronous session, students completed post-listening activities to build on and apply what they had learnt from listening. These discussions aimed to help students make personal connections and identify universal human values across the WE speakers’ varied contexts. For example, the application questions for Observing Ramadan asked students to reflect on what non-Muslims could learn from experiencing Ramadan. Students responded favourably to the use of podcasts featuring WE speakers, and felt their language skills had improved. One student explained, ‘My understanding of vocabulary and pronunciation increased, without me realizing it.’ On a survey the end of the semester, 89 percent of students said they felt their vocabulary had increased, and 76 percent reported that they appreciated the opportunity to hear good pronunciation. Students were also positively inclined toward the WE language models, which included speakers from Bangladesh, Canada, Ghana, Jordan, Lithuania, Saudi Arabia, the United States, and Yemen. Students said the speakers’ accents and vocabulary were not as difficult to understand as they had anticipated. In fact, some preferred listening to WE speakers, as the following comment reveals: ‘For me personally, listening to speakers who use English as their additional language is more easier than native speaker because we are both learning, so we have some similarities in pronunciation, or use general words.’ Students appreciated the slower speech and simpler vocabulary used by WE speakers. Students especially appreciated podcasts focused on familiar topics. Favourite episodes included Observing Ramadan, an episode about a Muslim prison chaplain in Canada, and one about a journalist who studied in Indonesia. A student explained that she preferred these episodes because ‘the experiences of the interviewees relate to our lives’. We found that the best-received podcasts were those with topics connected to students’ experiences. We see great potential in the use of podcasts featuring WE speakers in other contexts where the ‘native speaker’ model continues to hold power. We have several recommendations based on our experiences. First, we recommend orienting students to the purpose of listening to diverse WE speakers. As mentioned above, most of our students initially wished to sound like a ‘native speaker’, so it was important to offer a rationale for our seemingly contradictory WE paradigm. We believe that helping students consider the importance of engaging with a wide variety of WE speakers resulted in higher motivation and interest. Second, we suggest finding materials that include content related to students’ lives. Although our goal was to expose students to new language varieties and perspectives, we found that students were more interested when at least some of the content was familiar. They needed to be able to make personal connections with the material. When those connections do not come naturally, teachers should use teaching strategies such as guided reflection, discussion, and hypothetical scenarios to help students see how the speakers’ experiences are similar to or different from their own. Lastly, we encourage the use of materials which not only include diverse language varieties but also offer exposure to new cultures. Podcasts from diverse cultural contexts support the development of intercultural competence in addition to language competence. Given its focus on intercultural exchange, the 22.33 podcast was a rich source of such content. Other podcasts with diverse global speakers and intercultural themes include: Global Voices, Voices of Exchange, Rough Translation, The Europeans, and Sound Africa. As other practitioners use these and other podcasts to expose students to WE speakers, we hope they will also share their experiences. Last version received November 2022 Hanung Triyoko is the Head of the Language Development Unit at UIN Salatiga, Indonesia. He is an enthusiast for mutual relationships among educators across the globe and for encouraging his students and colleagues to have everyone’s unique contribution to use language as means of peace and better understanding of life and humankind. He is known as a master trainer and developer of a massive open online course (MOOC) for various levels of English teachers in Indonesia sponsored by RELO of the US Embassy. Email: [email protected] Tabitha Kidwell is a Professorial Lecturer in the TESOL programme at American University, and was previously a visiting lecturer at UIN Salatiga, Indonesia. She teaches academic writing, applied linguistics, and TESOL methods courses, and has conducted professional development for language teachers around the world. Her research focuses on innovative pedagogies, intercultural teaching approaches, and language teacher education. Email: [email protected]
Foregrounding denotes a linguistic phenomenon which allows a text to stand out against the backdrop of the text that conforms with the linguistic norms. The article sheds the light on its origin, the tendency of investigations of the theory of foregrounding as well as a wide range of scientific approaches deployed in an attempt to investigate the notion “foregrounding” and enlarges on what part it plays in generating meaning in a text. In the course of the research the following conclusions were drawn: a text that comprise foregrounded language means stimulates mental work and is perceived to be intricate and thought-provoking by a reader. Foregrounding gives a rise to the creation of a special perception of the object, the creation of a ‘vision’ of it, and not ‘recognition’”. Foregrounding in a language is normally realized through the usage of stylistic devices as well as consistent and systematic nature of actualization. It enables a writer to fulfill writer’s specific aims and motivations through the text.
Word embeddings trained on the lemmatised TOROT Treebank, using Word2Vec and the following parameters: sg = True min_count = <1,3,5> window = <3,5> vector_size = <100,200,300> epochs = 5 One model was trained for each combination of the parameters enclosed in angled brackets (< >). The release contains both the full models (.model) and the plain vector files (_vectors.txt). The models are named according to the parameters they were trained with. Note that these are the result of very preliminary experiments and no systematic evaluation of their quality was carried out, so use with caution.
Neste trabalho, realizamos uma descrição linguística e relatamos o processo de anotação do pronome -se no treebank PetroGold (v3). A atenção especial ao pronome -se se justifica pela necessidade de anotar corretamente os casos em que o pronome indica indeterminação do sujeito, voz passiva sintética ou verbo pronominal, reconhecendo sua importância para diversas tarefas de PLN. Como resultados, discriminamos as 1.960 ocorrências do "se" no corpus por classe sintática e apresentamos os verbos que se associam a cada um (ou mais de um) dos tipos do pronome -se.
We present new supertaggers trained on English grammar-based treebanks and test the effects of the best tagger on parsing speed and accuracy. The treebanks are produced automatically by large manually built grammars and feature high-quality annotation based on a well-developed linguistic theory (HPSG). The English Resource Grammar treebanks include diverse and challenging test datasets, beyond the usual WSJ section 23 and Wikipedia data. HPSG supertagging has previously relied on MaxEnt-based models. We use SVM and neural CRF- and BERT-based methods and show that both SVM and neural supertaggers achieve considerably higher accuracy compared to the baseline and lead to an increase not only in the parsing speed but also the parser accuracy with respect to gold dependency structures. Our fine-tuned BERT-based tagger achieves 97.26\% accuracy on 950 sentences from WSJ23 and 93.88% on the out-of-domain technical essay The Cathedral and the Bazaar (cb). We present experiments with integrating the best supertagger into an HPSG parser and observe a speedup of a factor of 3 with respect to the system which uses no tagging at all, as well as large recall gains and an overall precision gain. We also compare our system to an existing integrated tagger and show that although the well-integrated tagger remains the fastest, our experimental system can be more accurate. Finally, we hope that the diverse and difficult datasets we used for evaluation will gain more popularity in the field: we show that results can differ depending on the dataset, even if it is an in-domain one. We contribute the complete datasets reformatted for Huggingface token classification.
<em>This article touches upon the problems of improving the speech culture of students. It is known that one of the indicators of the level of culture, thinking, intellect of a person is his speech, which must comply with linguistic norms. It is in elementary school that children begin to master the norms of oral and written literary language, learn to use language tools in different communication conditions in accordance with the goals and objectives of speech.</em>
A dictionary’s basic function is to provide the meaning for all of the words contained within it. However, in addition to this key role, dictionaries also serve as a model for linguistic correctness, especially with reference to spelling norms and, consequently, they afford users an important guide to current orthographic standards. This chapter offers a series of reflections concerning a dictionary’s dual normative and semantic functions. First, we examine the extent to which linguistic norms should constitute part of lexicographic entries and then focus on the dictionary’s scope as a prescriptive reference tool that affirms and promotes a particular set of linguistic norms. Second, we debate the question of whether or not a dictionary should resolve any type of doubt that exists with respect to linguistic correctness, given the fact that other prescriptive tools, such as reference grammars, carry out the same prescriptive language function. We end this section with a reflection on the interdependence between language use and variation and how this interplay should be reflected in the formation of a set of ideal pan-Hispanic spelling norms. The second part of the chapter deals with the normative role played by the Royal Spanish Academy in conjunction with the Association of Spanish Academies in determining the nature and spread of orthographic standards for the Spanish language.
Nouns are one of the most frequently used word classes in English. As one variety of “World Englishes”, China English has received much concern in recent years. However, most research on nouns is still at the cognitive aspect or in the traditional qualitative way. As a new quantitative method, dependency grammar reveals the internal relations among the words in the sentence. Therefore, the syntactic features of nouns in China News English can provide a new perspective. The raw material of China English in this study is randomly crawled from China Daily, and the News part in the Corpus of Contemporary American English (COCA) from 2013 to 2016is selected as the reference corpus. By using the Jupyter Notebook based on Python, the dependency treebanks are generated. After comparing the proportion of the word classes, dependency direction, and dependency distance of different dependents of nouns in China Daily and COCA, this paper finds that 1) On the whole, the categories of dependents in China English and American English are similar. Determiners, adjectives, nouns, prepositions, verbs, numbers, proper nouns, conjunctions, pronouns and particles are the top ten frequently-used modifiers of nouns. Compared with American English, China English tends to use more adjectives and nouns, while the determiners especially the articles and pronouns are relatively less used. 2) Both in China English and American English, the percentage of HI constructions and HF constructions is close to 50%. China English tends to be more right-embedded than American English. Compared with AmE, in ChiE, prepositions, conjunctions and participles are more frequently used as post-modifiers. 3) The mean dependency distance value of China English (2.68) is larger than American English (2.61). Both in China English and American English, the mean dependency distance of pre-modifiers is shorter than in post-modifiers. The longer the mean dependency distance value is, the more difficult is in processing the information. It indicates that in the News domain, China English is more difficult to be processed in expression. This research provides a new perspective to studying English variants and enriches the current research on nouns.
This article is determined to analyze the way Russian political and ideological system is using linguistics in the means of its informational and psychological aggression against Ukraine. The war launched by Russian Federation has become the culmination of its prolonged aggressive actions on many fronts: ideological, economical, cultural, linguistic and so on. Current cremlin hosts are trying to regain full control over Ukraine in every imaginable way, drawing on the previous political regimes' experience. This article investigates one of the definitions of chauvinism as harassment of so-called "small nations" on the domestic and international levels that is shaping into political oppression and assimilation of languages. Have it that cremlin ideologues have always exploited any opportunity to diminish Ukrainians' self-consciousness, in particular by means of different quasi-scientific theories. Here we explored one example of the forementioned such as the creation of fake exception to the rule of spelling of prepositions "in" and "on" with administrative geographical names by Russian philologers that only applied to place-name "Ukraine". In contradiction to linguistic norms of Russian language Russians use "on Ukraine" in accusative and locative cases. We respectively analyzed arguments of Soviet linguists D. Rozental and K. Bilinskyy, and also modern Russian geographers and their theory of geo concept that basically comes to one statement: such version has been made up through history and is backed up by the expression "on the outskirts". The Ukrainian linguist I. Ohiyenko's work, in which he explained that the expression "in Ukraine" positioned it as a separate state, is mentioned in the article. After the collapse of the Soviet Union the logical norm "in Ukraine" was being used in the official documents of Russian Federation for a while, but after the cremlin tacked its course to re-establish its dominance over the former Soviet republics, they returned to the previous version. The sources that have been studied in the article point out that the political and ideological system is of a paramount importance for Russian linguistic science and is using it as a non-lethal weapon against Ukraine.
Capturing drivers’ affective responses given driving context and driver-pedestrian interactions remains a challenge for designing in-vehicle, empathic interfaces. To address this, we conducted two lab-based studies using camera and physiological sensors. Our first study collected participants’ (N = 21) emotion self-reports and physiological signals (including facial temperatures) toward non-verbal, pedestrian crossing videos from the Joint Attention for Autonomous Driving dataset. Our second study increased realism by employing a hybrid driving simulator setup to capture participants’ affective responses (N = 24) toward enacted, non-verbal pedestrian crossing actions. Key findings showed: (a) non-positive actions in videos elicited higher arousal ratings, whereas different in-video pedestrian crossing actions significantly influenced participants’ physiological signals. (b) Non-verbal pedestrian interactions in the hybrid simulator setup significantly influenced participants’ facial expressions, but not their physiological signals. We contribute to the development of in-vehicle empathic interfaces that draw on behavioral and physiological sensing to in-situ infer driver affective responses during non-verbal pedestrian interactions.
Aesthetic processing has profound implications for everyday life. Although liking and beauty judgements are outcomes of aesthetic processing and derive from a common hedonic value, there may be some differences in how they engage working memory. This study used maintenance and aesthetic judgement tasks to examine whether liking and beauty judgements make different demands on domain-specific working memory resources. Sixty participants (30 males) were instructed to rate picture for liking or beauty while maintaining the subjective affect or brightness of the presented pictures. Results indicated that liking judgements selectively impaired participants' performance in the affect maintenance task, and beauty judgements selectively impaired their performance in the brightness maintenance task. In addition, maintaining affect and brightness feelings in the mind increased image ratings on beauty but not on liking. Our findings provide evidence that liking judgements draw more on affective working memory resources than beauty judgements, and beauty judgements draw more on visual working memory resources than liking judgements.
Semantic relations between words in a sentence are identified using a technique called dependency parsing. Algorithms called dependency parsers are used to map the words in a sentence to the said semantic roles and to identify the syntactic relations between words. Transition-based dependencies on Indian languages have been relatively less explored, partially attributed to lack of high quality treebanks. Graph Neural Network based parsing techniques have shown relatively high accuracy levels on English language and languages with similar syntactic structures (eg. Spanish), and have potential to show good accuracy with other language syntaxes. Hindi is a language with an extremely rich morphology and a FreeWord order. Such a language, also referred to as a Morphologically Rich, FreeWord Order (MoR-FWO) language, is immensely difficult to parse using traditional methods. Using Graph Neural Networks, we present a state-of-the-art dependency parser for Hindi. We compare performance and efficiency of two approaches in this project- a Machine Learning model and a Neural Network model.
OBJECTIVE: Hoarding behaviour is a common but poorly characterised problem in real-world clinical practice. Although hoarding behaviour is the key component of Hoarding Disorder (HD), there are people who exhibit hoarding behaviour but do not suffer from HD. The aim of the present study was to characterise a clinical sample of patients with clinically relevant hoarding behaviour and evaluate the differential characteristics between patients with and without HD. METHODS: This study included patients who received treatment at the home visitation program in Barcelona (Spain) from January 2013 through December 2020, and scored ≥ 4 on the Clutter Image Rating scale. Sociodemographic, DSM-5 diagnosis, clinical data and differences between patients with and without an HD diagnosis were assessed. RESULTS: A total of 243 subjects were included. Hoarding behaviour had been unnoticed in its early stages and the median length in the sample was 10 years (IQR 15). 100% of the cases had hoarding-related complications. HD was the most common diagnosis in 117 patients (48.1%). CONCLUSIONS: The study found several differential characteristics between patients with and without HD diagnosis. Alcohol use disorder could play an important role among those without HD diagnosis. Home visitation programs could improve earlier detection, preventing hoarding-related complications.
This paper provides a comparative analysis of the patterns of formation and qualities of a modern linear text and an Internet text. The article is the result of the study of Internet text stylistics, based mainly on Russian-language texts from Russia and Ukraine. The paper considers the process of formation of a new type of text - the Internet text. Being essentially different from the classical linear text, the Internet text does not lend itself well to the description based on the classical text theory. Thus, the Internet text is not complete, vectorial, not united by the completeness of thought expression. An important characteristic of the Internet text is its interactivity, which means that the roles of addressee and addressee are constantly changing. In addition, it is difficult, or even impossible, to define the boundaries of the Internet text due to its hypertextuality, which has become habitual intertextuality. All the above-mentioned aspects make up the pragmatics of the Internet text as a subspecies of the media text. Another crucial problem addressed in the paper is the study of the regularities of Internet communication in general and the stylistics of the Internet text. In the course of the research it became obvious that speech aggression and violation of norms of speech culture are stylistic dominants of online communication. This influenced the formation of other stylistic dominants such as, hate speech, fake, hype, clickbait, etc. Internet style is clearly characterized by being provocative, aggressive, hostile. The problems of bullying, humiliation of human dignity, invective and obscenity are actively studied from the standpoint of linguoecology, because, according to most researchers, the constant neglect of communicative and linguistic norms leads to the degradation of the national language style and literary norms. Despite the fact that there is still a division between public and interpersonal online communication, i.e. formal and informal, the problems of speech behavior of Internet users are becoming increasingly relevant.