Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
While educators may be well positioned to support unaccompanied immigrant youth, there is limited interdisciplinary research focused on understanding the complexity of youth’s experiences in US schools. The purpose of this qualitative, interview-based study was to better understand how youth’s transnational experiences pre-, during, and post-migration affected their school-based experiences, and to explore how schools supported them. Participants included ten unaccompanied immigrant youths from Central America and six key informants who worked with youth in a professional capacity. Findings indicate that youth experienced multiple challenges including stressful and traumatic events, barriers to mental health and legal services, and unfamiliar cultural and linguistic norms that sometimes were not recognized or understood by their teachers and schools. The youth also brought important resources, such as high expectations and aspirations and strong connections to family and community. School-based experiences that built from youth’s resources and motivations (e.g., through school-community partnerships and responsive classroom practices) had the potential to enhance belonging, community connections, and wellness. More interdisciplinary research is needed to develop and support school-based practices and partnerships in consultation with youth that build from knowledge of their particular resources and challenges.
Dictionary-based methods in sentiment analysis have received scholarly attention recently, the most comprehensive examples of which can be found in English.However, many other languages lack polarity dictionaries, or the existing ones are small in size as in the case of Senti-TurkNet, the first and only polarity dictionary in Turkish.Thus, this study aims to extend the content of SentiTurkNet by comparing the two available WordNets in Turkish, namely KeNet and TR-wordnet of BalkaNet.To this end, a current Turkish polarity dictionary has been created relying on 76,825 synsets matching KeNet, where each synset has been annotated with three polarity labels, which are positive, negative and neutral.Meanwhile, the comparison of KeNet and TR-wordnet of BalkaNet has revealed their weaknesses such as the repetition of the same senses, lack of necessary merges of the items belonging to the same synset and the presence of redundant narrower versions of synsets, which are discussed in light of their potential to the improvement of the current lexical databases of Turkish.
OBJECTIVE: Nonsuicidal self-injury (NSSI) is often cited as a key risk factor for future suicidal behavior. Capability for suicide has been repeatedly cited as an important mechanism that can account for this association. Despite this, direct tests of this hypothesis have been rare and methodologically constrained. In the present study, we conducted a direct test of this hypothesis while addressing several constraints of prior literature. METHOD: In a large sample of suicidal and self-injuring adults (n = 1,020), we tested whether changes in fearlessness about death (FAD), a core facet of the capability for suicide, accounted for the relationship between NSSI and future suicide attempts at 28-day and 2-year follow-up. FAD was assessed using the gold-standard self-report form (ACSS-FAD), an implicit test of suicide-related affect (affect misattribution paradigm-Suicide), and explicit affective ratings of suicide-relevant images. Mediation with bootstrapping was implemented to test our main hypotheses. RESULTS: As anticipated, lifetime NSSI frequency was significantly associated with suicide attempt frequency at follow-up; however, FAD failed to consistently mediate this association. Results were largely consistent across all three measures of FAD. Post hoc power analyses indicated sufficient power to detect small effects. CONCLUSIONS: Taken together, these results fail to support the hypothesis that capability for suicide explains the link between NSSI and future suicidal behavior. We discuss the implications of our results for research and theory, situating our findings in the context of recent advances in the understanding of suicide risk more broadly. (PsycInfo Database Record (c) 2021 APA, all rights reserved).
Morphological tagging of code-switching (CS) data becomes more challenging especially when language pairs composing the CS data have different morphological representations. In this paper, we explore a number of ways of implementing a language-aware morphological tagging method and present our approach for integrating language IDs into a transformerbased framework for CS morphological tagging. We perform our set of experiments on the Turkish-German SAGT Treebank. Experimental results show that including language IDs to the learning model significantly improves accuracy over other approaches.
OBJECTIVE: To evaluate remote testing as a tool for measuring emotional responses to non-speech sounds. DESIGN: Participants self-reported their hearing status and rated valence and arousal in response to non-speech sounds on an Internet crowdsourcing platform. These ratings were compared to data obtained in a laboratory setting with participants who had confirmed normal or impaired hearing. STUDY SAMPLE: Adults with normal and impaired hearing. RESULTS: In both settings, participants with hearing loss rated pleasant sounds as less pleasant than did their peers with normal hearing. The difference in valence ratings between groups was generally smaller when measured in the remote setting than in the laboratory setting. This difference was the result of participants with normal hearing rating sounds as less extreme (less pleasant, less unpleasant) in the remote setting than did their peers in the laboratory setting, whereas no such difference was noted for participants with hearing loss. Ratings of arousal were similar from participants with normal and impaired hearing; the similarity persisted in both settings. CONCLUSIONS: In both test settings, participants with hearing loss rated pleasant sounds as less pleasant than did their normal hearing counterparts. Future work is warranted to explain the ratings of participants with normal hearing.
OBJECTIVES: Age differences in affective experience across adulthood are widely documented. According to the circumplex model of affect consists of 2 aspects-valence (positive vs negative) and arousal (low activation vs high activation). Prior research on age differences has primarily focused on the valence aspect. However, little is known about age differences in daily affect of high and low arousal. METHOD: The present study examined age differences in daily dynamics (i.e., mean levels, variability, and inertia) of negative affect (NA) and positive affect (PA) of high and low arousal in a sample of 492 adults aged 21-91. Participants completed daily affect ratings for 21 consecutive days. RESULTS: Age was negatively and linearly related to mean levels of both high-arousal and low-arousal NA. Both high-arousal and low-arousal PA mean levels showed increases after middle age. Further, age was related to lower variability in both NA and PA regardless of arousal. Additionally, high-arousal NA inertia showed a linear decrease with age, whereas low-arousal PA inertia showed an inverted-U pattern with age. After controlling for mean levels of affect, the associations between age and affect variability remained significant, whereas the associations between age and affect inertia did not. DISCUSSION: The affective profile of older age is characterized by lower mean levels of NA, higher mean levels of PA, lower affect variability, and less persistence in high-arousal NA and low-arousal PA in daily life. Our results contribute to a nuanced understanding of which affective processes improve with age and which do not.
We propose two fast neural combinatory models for constituency parsing: binary and multibranching. Our models decompose the bottomup parsing process into 1) classification of tags, labels, and binary orientations or chunks and 2) vector composition based on the computed orientations or chunks. These models have theoretical sub-quadratic complexity and empirical linear complexity. The binary model achieves an F1 score of 92.54 on Penn Treebank, speeding at 1327.2 sents/sec. Both the models with XLNet provide near state-of-theart accuracies for English. Syntactic branching tendency and headedness of a language are observed during the training and inference processes for Penn Treebank, Chinese Treebank, and Keyaki Treebank (Japanese).
In this study, the affective explicit and implicit attitudes toward electric and gasoline cars are investigated. One hundred sixty-five participants (103 cisgender women, 62 cisgender men) completed an explicit and implicit affective rating task toward pictures of electric and gasoline cars, measurements of sustainability, future and past behaviors, and mindfulness. The results showed a positive emotional attitude for the electric cars compared with the gasoline cars only for the explicit rating but not for the implicit one. Furthermore, factors that correlated to the attitudes were investigated: explicit ratings in car owners correlated with age, degree, sustainability in general, and the expressed intention to purchase an electric car in the future. Implicit attitudes in car owners correlated with the overall score of mindfulness and the dimension of "non-reactivity." For the non-car owners, explicit attitudes correlated with the expressed intention to purchase an electric car in the future and the mindfulness dimension of "describing". In this group, the implicit attitude correlated negatively with the mindfulness intention of acting with awareness. This indicates that several different factors should be considered in the development of promotion campaigns for the advantage of sustainable mobility behavior.
People use their previous experience to predict present affective events. Since we live in ever-changing environments, affective predictions must generalize from past contexts (from which they are implicitly learned) to new, potentially ambiguous contexts. This study investigated how past (un)certain relationships influence subjective experience following new ambiguous cues, and whether past relationships can be learned implicitly. Two S1-S2 paradigms were employed as learning and test phases in two experiments. S1s were colored circles, S2s negative or neutral affective pictures. Participants (N = 121, 116) were assigned to the certain (CG) or uncertain group (UG), and they were presented with 100% (CG) or 50% (UG) S1-S2 congruency during an uninstructed (Experiment 1) or implicit (Experiment 2) learning phase. During the test phase both groups were presented with a new 75% S1-S2 paradigm, and ambiguous (Experiment 1) or unambiguous (Experiment 2) S1s. Participants were asked to rate the expected valence of upcoming S2s (expectancy ratings), or their experienced valence and arousal (valence and arousal ratings). In Experiment 1 ambiguous cues elicited less negative expectancy ratings, and less unpleasant valence ratings, independently from prior experience. In Experiment 2, participants in the CG reported more negative expectancy ratings after the S1s previously paired with negative stimuli. Overall, we found that in the presence of ambiguous cues subjective affective experience is dampened, and we confirmed that people are able to infer probabilistic relationships from the environment (and to use them later) at an implicit level.
STUDY OBJECTIVES: Sleep plays a pivotal role in the off-line processing of emotional memory. However, much remains unknown for its immediate vs. long-term influences. We employed behavioral and electrophysiological measures to investigate the short- and long-term impacts of sleep vs. sleep deprivation on emotional memory. METHODS: Fifty-nine participants incidentally learned 60 negative and 60 neutral pictures in the evening and were randomly assigned to either sleep or sleep deprivation conditions. We measured memory recognition and subjective affective ratings in 12- and 60-h post-encoding tests, with EEGs in the delayed test. RESULTS: In a 12-h post-encoding test, compared to sleep deprivation, sleep equally preserved both negative and neutral memory, and their affective tones. In the 60-h post-encoding test, negative and neutral memories declined significantly in the sleep group, with attenuated emotional responses to negative memories over time. Furthermore, two groups showed spatial-temporally distinguishable ERPs at the delayed test: while both groups showed the old-new frontal negativity (300-500 ms, FN400), sleep-deprived participants additionally showed an old-new parietal, Late Positive Component effect (600-1000 ms, LPC). Multivariate whole-brain ERPs analyses further suggested that sleep prioritized neural representation of emotion over memory processing, while they were less distinguishable in the sleep deprivation group. CONCLUSIONS: These data suggested that sleep's impact on emotional memory and affective responses is time-dependent: sleep preserved memories and affective tones in the short term, while ameliorating affective tones in the long term. Univariate and multivariate EEG analyses revealed different neurocognitive processing of remote, emotional memories between sleep and sleep deprivation groups.
Recurrent neural networks are efficient ways of training language models, and various RNN networks have been proposed to improve performance. However, with the increase of network scales, the overfitting problem becomes more urgent. In this paper, we propose a framework-G2Basy-to speed up the training process and ease the overfitting problem. Instead of using predefined hyperparameters, we devise a gradient increasing and decreasing technique that changes the parameters training batch size and input dropout simultaneously by a user-defined step size. Together with a pretrained word embedding initialization procedure and the introduction of different optimizers at different learning rates, our framework speeds up the training process dramatically and improves performance compared with a benchmark model of the same scale. For the word embedding initialization, we propose the concept of "artificial features" to describe the characteristics of the obtained word embeddings. We experiment on two of the most often used corpora-the Penn Treebank and WikiText-2 datasets-and both outperform the benchmark results and show potential towards further improvement. Furthermore, our framework shows better results with the larger and more complicated WikiText-2 corpus than with the Penn Treebank. Compared with other state-of-the-art results, we achieve comparable results with network scales hundreds of times smaller and within fewer training epochs.
BACKGROUND: Youth with anxiety disorders struggle with managing emotions relative to peers, but the neural basis of this difference has not been examined. METHODS: = 13.6; range = 8-17) with (n = 37) and without (n = 24) anxiety disorders completed a cognitive reappraisal task while undergoing functional magnetic resonance imaging. Emotional reactivity and regulation, functional activation, and beta-series connectivity were compared across groups. RESULTS: Groups did not differ on emotional reactivity or regulation. However, fronto-limbic activation after viewing aversive imagery with and without regulation, as well as affect ratings without regulation, were higher for anxious youth. Neither group demonstrated age-related changes in regulation, though anxious youth became less reactive with age. Stronger amygdala-ventromedial prefrontal cortex connectivity related to greater anxiety in control youth, but less anxiety in anxious youth. CONCLUSION: Anxious youth regulated when instructed, but regulation ability did not relate to age. Viewing aversive imagery related to heightened fronto-limbic activation even after reappraisal. Emotion dysregulation in youth anxiety disorders may stem from heightened emotionality and potent bottom-up neurobiological responses to aversive stimuli. Findings suggest the importance of treatments focused on both reducing initial emotional reactivity and bolstering regulatory capacity.
State-of-the-art neural language models represented by Transformers are becoming increasingly complex and expensive for practical applications. Low-bit deep neural network quantization techniques provides a powerful solution to dramatically reduce their model size. Current low-bit quantization methods are based on uniform precision and fail to account for the varying performance sensitivity at different parts of the system to quantization errors. To this end, novel mixed precision DNN quantization methods are proposed in this paper. The optimal local precision settings are automatically learned using two techniques. The first is based on a quantization sensitivity metric in the form of Hessian trace weighted quantization perturbation. The second is based on mixed precision Transformer architecture search. Alternating direction methods of multipliers (ADMM) are used to efficiently train mixed precision quantized DNN systems. Experiments conducted on Penn Treebank (PTB) and a Switchboard corpus trained LF-MMI TDNN system suggest the proposed mixed precision Transformer quantization techniques achieved model size compression ratios of up to 16 times over the full precision baseline with no recognition performance degradation. When being used to compress a larger full precision Transformer LM with more layers, overall word error rate (WER) reductions up to 1.7% absolute (18% relative) were obtained.
A growing body of research analyzing musical scores suggests mode’s relationship with other expressive cues has changed over time. However, to the best of our knowledge, the perceptual implications of these changes have not been formally assessed. Here, we explore how compositional choices of 17th- and 19th-century composers (J. S. Bach and F. Chopin, respectively) differentially affect emotional communication. This novel exploration builds on our team’s previous techniques using commonality analysis to decompose intercorrelated cues in unaltered excerpts of influential compositions. In doing so, we offer an important naturalistic complement to traditional experimental work—often involving tightly controlled stimuli constructed to avoid the intercorrelations inherent to naturalistic music. Our data indicate intriguing changes in cues’ effects between Bach and Chopin, consistent with score-based research suggesting mode’s “meaning” changed across historical eras. For example, mode’s unique effect accounts for the most variance in valence ratings of Chopin’s preludes, whereas its shared use with attack rate plays a more prominent role in Bach’s. We discuss the implications of these findings as part of our field’s ongoing effort to understand the complexity of musical communication—addressing issues only visible when moving beyond stimuli created for scientific, rather than artistic, goals.
Using multiple treebanks to improve parsing performance has shown positive results. However, to what extent similar, yet competing annotation decisions play in parser behavior is unclear. We investigate this within a multi-task learning (MTL) dependency parser setup on two parallel treebanks, UD and SUD, which, while possessing similar annotation schemes, differ in specific linguistic annotation preferences. We perform a set of experiments with different MTL architectural choices, comparing performance across various input embeddings. We find languages tend to pattern in loose typological associations, but generally the performance within an MTL setting is lower than single model baseline parsers for each annotation scheme. The main contributing factor seems to be the competing syntactic annotation information shared between treebanks in an MTL setting, which is shown in experiments against differently annotated treebanks. This suggests that the impact of how the signal is encoded for annotations and its influence on possible negative transfer is more important than that of the input embeddings in an MTL setting.
Ukrainian literary language is a complex communicative system, synchronous-diachronic section of which is informative regarding the state of development of national consciousness, intellectual level of society, and in a broader sense – concerning inscribing of the language in the history of national and world culture. This article is devoted to tracing the temporal and conceptual dynamics of the stylistic norm in the work of Ukrainian writers of the 20th century. The methodological basis for the study was made by such methods as comparison and analysis of literary works of Ukrainian writers of the XX century. As a result of the study, the authors concluded that the linguistic norm is a cognitive reference point for the scientific parameterization of the style norm of artistic discourse. The authors also emphasize that the dynamism of the stylistic artistic norm lies in the ability of the poetic language to respond to the development of artistic and linguistic consciousness, thinking of both the author and the reader in various manifestations.
Research has shown that patients with a social anxiety disorder (SAD) show social performance deficits. These deficits are a maintaining factor in SAD, as mending social behavior improves interpersonal judgments and reduces social anxiety. Thus finding ways to enhance social behavior is evidently of importance in the treatment of SAD. This double-blind, placebo-controlled study investigated the effect of an intranasal administration of the hormone oxytocin (24 IU) on social behavior and anxious appearance in SAD patients (N = 40) and healthy controls (N = 39). Forty minutes after oxytocin administration participants were submitted to two live social situations (i.e., a waiting room situation and a getting acquainted task). The participants ('self-rated') and observers ('observer-rated') scored participants' social behavior and anxious appearance. Participants also rated their positive and negative affect. Confirming the social performance deficits in SAD, observers regarded SAD patients as more anxious and less socially skilled than healthy controls. Results indicated oxytocin-induced improvement of observer-rated social behavior in SAD patients compared to placebo but only in the getting acquainted task. This effect was not perceived as such by patients themselves and did not improve their affect ratings. In conclusion, this study found support for the idea that oxytocin helps SAD patients to perform better in social interactions, although this improvement seemed context-dependent (i.e., only present in the getting-acquainted task) and 'not perceived by the patient.
Homomorphic encryption (HE) and garbled circuit (GC) provide the protection for users' privacy. However, simply mixing the HE and GC in RNN models suffer from long inference latency due to slow activation functions. In this paper, we present a novel hybrid structure of HE and GC gated recurrent unit (GRU) network, CRYPTOGRU, for low-latency secure inferences. CRYPTOGRU replaces computationally expensive GC-based tanh with fast GC-based ReLU, and then quantizes sigmoid and ReLU to smaller bit-length to accelerate activations in a GRU. We evaluate CRYP-TOGRU with multiple GRU models trained on 4 public datasets. Experimental results show CRYPTOGRU achieves top-notch accuracy and improves the secure inference latency by up to 138 over one of the state-of-the-art secure networks on the Penn Treebank dataset.
The racist ideology of traditional, ruling‐class Brazilian nationalism, which denies the existence of racial divisions, is inherently anti‐Black. For example, in 2020 the Federal Government of Brazil revoked affirmative action programs for graduate degrees in universities. These anti‐black ideologies also influence linguistics in Brazil. In the twenty‐first century, one of the high‐profile representatives of this initiative is the Educated Urban Linguistic Norm Project that chose the urban speaker, who is mostly white, as the norm for all the speakers. Similarly, a series of daily online lectures hosted by ABRALIN, the national professional association for linguistics, beginning in May 2020, was without Black Brazilian speakers over the first month and a half of the schedule. In this work we seek to provoke discussions towards rethinking the role of whiteness in Brazilian linguistics moving from the Black‐as‐theme to Black‐as‐life framework.
Este trabalho descreve a criação do PetroGold, um treebank padrão ouro para o domínio do óleo & gás. O material é composto por teses, dissertações e monografias, contém 9.127 frases (253.640 tokens) e conta com anotação morfossintática de dependências segundo a abordagem Universal Dependencies. Detalhamos alguns dos desafios linguísticos do domínio para a anotação sintática e verificamos a qualidade do material produzido por meio de uma avaliação intrínseca: utilizando um modelo criado pela ferramenta UDPipe, o corpus leva a 90,65%, 88,53% e 82,88% de acertos conforme as medidas UAS, LAS e CLAS, respectivamente.
BACKGROUND: The strong and long lockdown adopted by the Italian government to limit COVID-19 spreading represents the first threat-related mass isolation in history that can be studied in depth by scientists to understand individuals' emotional response to a pandemic. METHODS: We investigated the effects on individuals' mental wellbeing of this long-term isolation by means of an online survey on 71 Italian volunteers. They completed the Positive and Negative Affect Schedule and Fear of COVID-19 Scale and judged valence, arousal, and dominance of words either related or unrelated to COVID-19, as identified by Google search trends. RESULTS: Emotional judgments changes from normative data varied depending on word type and individuals' emotional state, revealing early signals of individuals' mental distress to COVID-19 confinement. All individuals judged COVID-19-related words to be less positive and dominant. However, individuals with more negative feelings and COVID-19 fear also judged COVID-19-unrelated words to be less positive and dominant. Moreover, arousal ratings increased for all words among individuals with more negative feelings and COVID-19 fear but decreased among individuals with less negative feelings and COVID-19 fear. DISCUSSION: Our results show a rich picture of emotional reactions of Italians to tight and 2-month long confinement, identifying early signals of mental health distress. They are an alert to the need for intervention strategies and psychological assessment of individuals potentially needing mental health support following the COVID-19 situation.
INTRODUCTION: The addition of graphic health warnings to cigarette packets can facilitate smoking cessation, primarily through their ability to elicit a negative affective response. Smoking has been linked to COVID-19 mortality, thus making it likely to elicit a strong affective response in smokers. COVID-19-related health warnings (C19HW) may therefore enhance graphic health warnings compared to traditional health warnings (THW). Further, because impulsivity influences smoking behaviors, we also examined whether these affective responses were associated with delay discounting. METHODS: In a between-subjects design, 240 smokers rated the valence and arousal elicited by tobacco packaging that contained either a C19HW or THW (both referring to death). Participants also completed questionnaires to quantify delay discounting, and attitudes towards COVID-19 and smoking (eg, health risks, motivation to quit). RESULTS: There were no differences between the two health warning types on either valence or arousal, nor any secondary outcome variables. There was, however, a significant interaction between health warning type and delay discounting on arousal ratings. Specifically, in smokers who exhibit low delay discounting, C19HWs elicited significantly greater subjective arousal rating than did THWs, whereas there was no significant effect of health warning type on arousal in smokers who exhibited high delay discounting. CONCLUSION: The results suggest that in smokers who exhibit low impulsivity (but not high impulsivity) C19HWs may be more arousing than THWs. Future work is required to explore the long-term utility of C19HWs, and to identify the specific mechanism by which delay discounting moderates the efficacy of tobacco health warnings. IMPLICATIONS: The study is the first to explore the impact of COVID-19-related health warnings on cigarette packaging. The results suggest that COVID-19-related warnings elicit a similar level of negative emotional arousal, relative to traditional warnings. However, COVID-19 warnings, specifically, elicit especially strong emotional responses in less impulsive smokers, who report low delay discounting. Therefore, there is preliminary evidence supporting COVID-19 related warnings for tobacco products to aid smoking cessation. Additionally, there is novel evidence that, for some warnings, high impulsiveness may be a factor in reduced warning efficacy, which may explain poorer cessation success in this population.
ABSTRACT: Pain-related learning mechanisms likely play a key role in the development and maintenance of chronic pain. Previous smaller-scale studies have suggested impaired pain-related learning in patients with chronic pain, but results are mixed, and chronic back pain (CBP) particularly has been poorly studied. In a differential conditioning paradigm with painful heat as unconditioned stimuli, we examined pain-related acquisition and extinction learning in 62 patients with CBP and 61 pain-free healthy male and female volunteers using valence and contingency ratings and skin conductance responses. Valence ratings indicate significantly reduced threat and safety learning in patients with CBP, whereas no significant differences were observed in contingency awareness and physiological responding. Moreover, threat learning in this group was more impaired the longer patients had been in pain. State anxiety was linked to increased safety learning in healthy volunteers but enhanced threat learning in the patient group. Our findings corroborate previous evidence of altered pain-related threat and safety learning in patients with chronic pain. Longitudinal studies exploring pain-related learning in (sub)acute and chronic pain are needed to further unravel the role of aberrant pain-related learning in the development and maintenance of chronic pain.
We propose the Recursive Non-autoregressive Graph-to-Graph Transformer architecture (RNGTr) for the iterative refinement of arbitrary graphs through the recursive application of a non-autoregressive Graph-to-Graph Transformer and apply it to syntactic dependency parsing. We demonstrate the power and effectiveness of RNGTr on several dependency corpora, using a refinement model pre-trained with BERT. We also introduce Syntactic Transformer (SynTr), a non-recursive parser similar to our refinement model. RNGTr can improve the accuracy of a variety of initial parsers on 13 languages from the Universal Dependencies Treebanks, English and Chinese Penn Treebanks, and the German CoNLL2009 corpus, even improving over the new state-of-the-art results achieved by SynTr, significantly improving the state-of-the-art for all corpora tested.
Abstract This paper presents an open source and extendable Morphological Analyser cum Generator (MAG) for Tamil named Thamizhi Morph. Tamil is a low-resource language in terms of NLP processing tools and applications. In addition, most of the available tools are neither open nor extendable. A morphological analyser is a key resource for the storage and retrieval of morphophonological and morphosyntactic information, especially for morphologically rich languages, and is also useful for developing applications within Machine Translation. This paper describes how Thamizhi Morph is designed using a Finite-State Transducer (FST) and implemented using Foma. We discuss our design decisions based on the peculiarities of Tamil and its nominal and verbal paradigms. We specify a high-level meta-language to efficiently characterise the language’s inflectional morphology. We evaluate Thamizhi Morph using text from a Tamil textbook and the Tamil Universal Dependency treebank version 2.5. The evaluation and error analysis attest a very high performance level, with the identified errors being mostly due to out-of-vocabulary items, which are easily fixable. In order to foster further development, we have made our scripts, the FST models, lexicons, Meta-Morphological rules, lists of generated verbs and nouns, and test data sets freely available for others to use and extend upon.
In this paper, the process of creating a Dependency Treebank for tweetsin Urdu,a morphologically rich and less-resourced languageis described. The 500 Urdu tweets treebank iscreated by manually annotating the treebank withlemma, POS tags, morphological and syntacticrelations using the Universal Dependencies annotation scheme, adopted to the peculiarities of Urdu social media text. annotation process is evaluated through Inter-annotator agreement for dependency relations and total agreement of 94.5% and resultant weighted Kappa = 0.876was observed. The treebank is evaluated through 10-fold cross validation using Maltparserwith various feature settings. Results show average UAS score of 74%, LAS score of 62.9% and LA score of 69.8%.
The study aimed to investigate if hentai consumers differed from other pornography consumers regarding their attachment style, attraction to, and desire for romantic relationships with anime characters and humans. Pornography consumers were categorized into three groups. The first group consumed both hentai and human pornography (hentai consumers), the second consumed human pornography but not hentai (non-hentai), and the third did not consume hentai or human pornography (non-porn). Two hundred and eight participants completed an online study that involved self-report surveys and an image rating task. The results revealed that hentai consumers did not differ from non-hentai or non-porn consumers on avoidant attachment. However, among females, hentai consumers were higher on anxious attachment compared to non-porn consumers. For the image rating task, hentai consumers rated anime characters more attractive than non-hentai and non-porn consumers. However, there were no group differences for the image ratings of real people. Hentai consumers indicated stronger romantic desire towards anime characters compared to non-hentai and non-porn consumers; there were no group differences in romantic desire for humans. The findings highlight the importance of differentiating individuals who consume hentai and those who do not.
Text discourse parsing weighs importantly in understanding information flow and argumentative structure in natural language, making it beneficial for downstream tasks. While previous work significantly improves the performance of RST discourse parsing, they are not readily applicable to practical use cases: (1) EDU segmentation is not integrated into most existing tree parsing frameworks, thus it is not straightforward to apply such models on newly-coming data. (2) Most parsers cannot be used in multilingual scenarios, because they are developed only in English. (3) Parsers trained from single-domain treebanks do not generalize well on out-of-domain inputs. In this work, we propose a document-level multilingual RST discourse parsing framework, which conducts EDU segmentation and discourse tree parsing jointly. Moreover, we propose a cross-translation augmentation strategy to enable the framework to support multilingual parsing and improve its domain generality. Experimental results show that our model achieves state-of-the-art performance on document-level multilingual RST parsing in all sub-tasks.
The new and growing field of Quantitative Dependency Syntax has emerged at the crossroads between Dependency Syntax and Quantitative Linguistics. One of the main concerns in this field is the statistical patterns of syntactic dependency structures. These structures, grouped in treebanks, are the source for statistical analyses in these and related areas; dozens of scores devised over the years are the tools of a new industry to search for patterns and perform other sorts of analyses. The plethora of such metrics and their increasing complexity require sharing the source code of the programs used to perform such analyses. However, such code is not often shared with the scientific community or is tested following unknown standards. Here we present a new open-source tool, the Linear Arrangement Library (LAL), which caters to the needs of, especially, inexperienced programmers. This tool enables the calculation of these metrics on single syntactic dependency structures, treebanks, and collection of treebanks, grounded on ease of use and yet with great flexibility. LAL has been designed to be efficient, easy to use (while satisfying the needs of all levels of programming expertise), reliable (thanks to thorough testing), and to unite research from different traditions, geographic areas, and research fields.
Chat-based counselling has become increasingly popular in the era of telecommunication. The need for accessible therapy has been exacerbated by the COVID-19 pandemic. Given its text-based nature, chat-based counselling provides an opportunity for machine-based analysis. It even has the potential to provide machine-based counselling services. However, the informational resources for machine-based analysis and interaction are rather scarce especially in a Japanese-language context. We created a Japanese dictionary for sentiment analysis, using a technique via machine-based text analysis, tailored for counselling related text. It includes 2389 words that were frequently used in chat-based counselling corpora. The following attributes were included for each word: (1) valence rating by the general public, (2) valence rating by clinical psychologists, (3) emotionality, and (4) body-relatedness.
Adapting threat-related memories towards changing environments is a fundamental ability of organisms. One central process of fear reduction is suggested to be extinction learning, experimentally modeled by extinction training that is repeated exposure to a previously conditioned stimulus (CS) without providing the expected negative consequence (unconditioned stimulus, US). Although extinction training is well investigated, evidence regarding process-related changes in neural activation over time is still missing. Using optimized delayed extinction training in a multicentric trial we tested whether: 1) extinction training elicited decreasing CS-specific neural activation and subjective ratings, 2) extinguished conditioned fear would return after presentation of the US (reinstatement), and 3) results are comparable across different assessment sites and repeated measures. We included 100 healthy subjects (measured twice, 13-week-interval) from six sites. 24 h after fear acquisition training, extinction training, including a reinstatement test, was applied during fMRI. Alongside, participants had to rate subjective US-expectancy, arousal and valence. In the course of the extinction training, we found decreasing neural activation in the insula and cingulate cortex as well as decreasing US-expectancy, arousal and negative valence towards CS+. Re-exposure to the US after extinction training was associated with a temporary increase in neural activation in the anterior cingulate cortex (exploratory analysis) and changes in US-expectancy and arousal ratings. While ICCs-values were low, findings from small groups suggest highly consistent effects across time-points and sites. Therefore, this delayed extinction fMRI-paradigm provides a solid basis for the investigation of differences in neural fear-related mechanisms as a function of anxiety-pathology and exposure-based treatment.
Abstract Although Lebanese use different languages, mainly Arabic, English, and French, in their daily interactions, the Lebanese educational system continues to adopt monolingual-oriented practices to language teaching. In English language classes, teachers and students are expected to exclusively use the target language as the language of instruction. Such a practice hinders students’ engagement in the learning process, especially during early stages, as they often face vocabulary and pronunciation challenges. Hence, there is a discrepancy between linguistic norms prevailing outside the classroom and the teaching strategies implemented inside the classroom. In an attempt to bridge the gap and maximize students’ English learning, I utilized my position as an English instructor to teach through the implementation of multilingual strategies. This article draws on my experience and on my students’ feedback to provide insights into implementing multilingual strategies to teach English language, with a focus on design and affordances of specific strategies.
Motivated by collective emotions theories that propose emotions shared between individuals predict group-level qualities, we hypothesized that co-experienced affect during interactions is associated with relationship quality, above and beyond the effects of individually experienced affect. Consistent with positivity resonance theory, we also hypothesized that co-experienced positive affect would have a stronger association with relationship quality than would co-experienced negative affect. We tested these hypotheses in 150 married couples across 3 conversational interactions: a conflict, a neutral topic, and a pleasant topic. Spouses continuously rated their individual affective experience during each conversation while watching video-recordings of their interactions. These individual affect ratings were used to determine, for positive and negative affect separately, the number of seconds of co-experienced affect and individually experienced affect during each conversation. In line with hypotheses, results from all 3 conversational topics suggest that more co-experienced positive affect is associated with greater marital quality, whereas more co-experienced negative affect is associated with worse marital quality. Individual level affect factors added little explanatory value beyond co-experienced affect. Comparing co-experienced positive affect and co-experienced negative affect, we found that co-experienced positive affect generally outperformed co-experienced negative affect, although co-experienced negative affect was especially diagnostic during the pleasant conversational topic. Findings suggest that co-experienced positive affect may be an integral component of high-quality relationships and highlight the power of co-experienced affect for individual perceptions of relationship quality. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
While high performance have been obtained for high-resource languages, performance on low-resource languages lags behind. In this paper we focus on the parsing of the low-resource language Frisian. We use a sample of code-switched, spontaneously spoken data, which proves to be a challenging setup. We propose to train a parser specifically tailored towards the target domain, by selecting instances from multiple treebanks. Specifically, we use Latent Dirichlet Allocation (LDA), with word and character N-grams. We use a deep biaffine parser initialized with mBERT. The best single source treebank (nl_alpino) resulted in an LAS of 54.7 whereas our data selection outperformed the single best transfer treebank and led to 55.6 LAS on the test data. Additional experiments consisted of removing diacritics from our Frisian data, creating more similar training data by cropping sentences and running our best model using XLM-R. These experiments did not lead to a better performance.
Abstract Despite the importance of mastering different types of formulaic sequences in a second language, little is known about the relative effect of different input modes on their acquisition. This study explores the learning of a particular type of formulaic language (binomials) in three input modes (reading-only, listening-only, and reading-while-listening) at different frequencies of exposure (2, 4, 5 and 6 occurrences). Arabic learners of English were presented with three stories, each in a different mode, that contained novel binomials (e.g., wires and pipes ) and existing binomials (e.g., brother and sister ). Two post-tests (multiple-choice and familiarity ratings) assessed learners’ knowledge of the binomials. Results showed that reading-only and reading-while-listening led to better performance on the tasks than listening-only. Frequency of exposure had an effect on the perceived familiarity of binomials.
Eesti veebipuudepanga tekstid (Muischnek et al., 2019), mis on annoteeritud käsitsi nii ortograafiliste kui süntaktiliste lausepiiridega, samuti on kontrollitud ja parandatud sõnestust. Lausete annoteerimisprotsessi kirjeldavad Sirts ja Peekman (2020), sõnestuse kontrolli kirjeldab Kairit Peekmani (2020) bakalaureusetöö. Andmete kasutamisel palume viidata Sirts ja Peekman (2020) artiklile. Muischnek, K., Müürisep, K., & Särg, D. D. (2019). CG Roots of UD Treebank of Estonian Web Language. In Proceedings of the NoDaLiDa 2019 Workshop on Constraint Grammar-Methods, Tools and Applications, 30 September 2019, Turku, Finland (No. 168, pp. 23-26). Linköping University Electronic Press. Peekman, K. (2020). Automaatse lausestamise ja sõnestamise hindamine uue meedia keele korpusel (bakalaureusetöö). Tartu Ülikool. Kättesaadav https://comserv.cs.ut.ee/ati_thesis/datasheet.php?id=69690&year=2020. Sirts, K., & Peekman, K. (2020). Evaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts. In Volume 328: Human Language Technologies – The Baltic Perspective, Frontiers in Artificial Intelligence and Applications, pages 174-181.
Para un lingüista o computólogo orientado al análisis de textos en español, analizar oraciones sintácticamente puede ser un esfuerzo de mucho tiempo cuando la cantidad de oraciones es elevada. Esta labor es más delicada si se toma en cuenta la variante del español que se emplee. Más aún, seleccionar cómo se etiquetan las palabras según la función puede complicar el proceso y el su impacto al compartir el conocimiento adquirido. Esta investigación propone realizar parte del esfuerzo en forma automática, utilizando reglas gramaticales, con el fin de analizar las oraciones sintácticamente y etiquetarlas con dependencias universales; un etiquetado estándar y nemónico, capaz de ser aplicado a diferentes idiomas.
The Mongolian written language and the traditional Mongolian script were a “pre-modern” language and script that transcended ethnicity and dialect. The Mongolian script can be read in any dialect. However, modern languages demand pronunciation norms, making the Mongolian script unsuitable as the official script of a modern nation-state. The countries and regions that used the Mongolian script changed the script in the first half of the 20th century, except for the Mongolian ethnic areas of China which continued to use the original Mongolian script. This was possible as Mongolian was a minority language in China, and the scope of its use was limited. However, Inner Mongolia, too, faced the issue of the written language not conforming with the spoken language. Therefore, in the 1930s, Mongolian literary figures and others in Manchukuo attempted to unify the Mongolian written and spoken language based on the genbun itchi movement in Japan. Genbun itchi sought to unify the Japanese written and spoken language, and was translated literally as “üge üsüg-i nigen bolcaqui” in Mongolian. This effort was only partially successful in changing the Mongolian script to match the spoken pronunciation, and systematic genbun itchi could not be achieved. Since 1945, Inner Mongolia learned from the example of the Mongolian People’s Republic and transformed the style of Mongolian script to that of modern Mongolian while still using the Mongolian script. However, the discord between the Mongolian script and the spoken pronunciation remains unresolved, and to address this, changes are being made to the Mongolian script to this day as they have been for the last eight decades. It is important to know that the Mongolian script is not only the modern script used in Chinese territory, but also the traditional script of the written language common to areas using the Mongolian written language. For this reason, the Mongolian script must be passed on to future generations. A means for unifying the Mongolian language while using the Mongolian script is to follow the Hanyu Pinyin system, which has succeeded in standardizing pronunciation while using Chinese characters—namely, enhance the Inner Mongolian “standard-sounding” writing system so that it can also function at the sentence level. The orthography of the Cyrillic alphabet of Mongolia offers an objectively good example for the development of a standard-sounding writing system in Inner Mongolia. After all, the problem of the Mongolian script reform comes down to that of genbun itchi.
PURPOSE: Medical education has been transformed during the COVID-19 pandemic, creating challenges regarding adequate training in ultrasound (US). Due to the discontinuation of traditional classroom teaching, the need to expand digital learning opportunities is undeniable. The aim of our study is to develop a tele-guided US course for undergraduate medical students and test the feasibility and efficacy of this digital US teaching method. MATERIALS AND METHODS: A tele-guided US course was established for medical students. Students underwent seven US organ modules. Each module took place in a flipped classroom concept via the Amboss platform, providing supplementary e-learning material that was optional and included information on each of the US modules. An objective structured assessment of US skills (OSAUS) was implemented as the final exam. US images of the course and exam were rated by the Brightness Mode Quality Ultrasound Imaging Examination Technique (B-QUIET). Achieved points in image rating were compared to the OSAUS exam. RESULTS: A total of 15 medical students were enrolled. Students achieved an average score of 154.5 (SD ± 11.72) out of 175 points (88.29 %) in OSAUS, which corresponded to the image rating using B-QUIET. Interrater analysis of US images showed a favorable agreement with an ICC (2.1) of 0.895 (95 % confidence interval 0.858 < ICC < 0.924). CONCLUSION: US training via teleguidance should be considered in medical education. Our pilot study demonstrates the feasibility of a concept that can be used in the future to improve US training of medical students even during a pandemic.
This article deals with a vision for the achievement of the project of a linguistic atlas of the dialects of Algeria supported by colored digital maps, showing the dialectical diversity of the selected region. And the circulation, the dialects used and current on the tongue of the inhabitants of Algeria We will focus in our research on the aspect of processing linguistic data in the manufacture of a digital linguistic atlas, in an attempt to invest computer data in describing local dialects and access to digital content, with the possibility of this content audio recordings of dialect variations, by providing our data bank through this linguistic research Also to pay tribute to the importance of the linguistic atlas in meeting the need of dialect workers for linguistic maps that identify the locations, nature and types of dialectal diversity in Algeria, The research also comes mainly to embody the digital principle and support the Arabic language by storing it digitally with the possibility of managing and printing linguistic maps according to the changes in them.
Introduction<br><br> Penn Discourse Treebank Version 2.0 - German Translation was developed at the University of Potsdam's Applied Computational Linguistics group and consists of approximately one million tokens derived from Penn Discourse Treebank Version 2.0 (LDC2008T05). This data was translated into German and annotated for shallow discourse relations in the financial news domain.<br><br> The aim of the University of Pennsylvania's Penn Discourse Treebank (PDTB) project is to annotate the Wall Street Journal text in Treebank-2 with discourse relations. PDTB 2.0 contains 40,600 tokens of annotation relations. PDTB2-German is based on a subset of PDTB2.0 used in the 2016 CoNLL Shared Task on Multilingual Shallow Discourse Parsing. Data<br><br> Data is in CoNLL format. Text was automatically translated into German with deepL, and projections of the annotations using word alignments were produced with GIZA++. See the included documentation for more information on the relation annotations.<br><br> Source text and CoNLL format annotations are each presented in their own tab separated plain text file, encoded in UTF-8. Samples<br><br> Please view this source sample and annotation sample. Updates<br><br> None at this time. Copyright Portions © 1987-1989 Dow Jones & Company, Inc., © 2008, 2012, 2021 The Penn Discourse Treebank Group, © 2021 Manfred Stede, © 1993-1995, 2008, 2012, 2021 Trustees of the University of Pennsylvania
Recent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection. Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups. We harness this genre metadata as a weak supervision signal for targeted data selection in zeroshot dependency parsing. Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations. We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios. For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection. Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages.
The linguistic worldview is a reflection of the national cognitive worldview. ‘Worldview’ is often defined as a way of perceiving the surrounding reality, yet the way people perceive their personal inner world also reflects their national self-identification. It is difficult to compare how people of different nations experience emotions and perceive such experiences because these processes are not available for direct observation and objective assessment. The most complete representation of the way a person experiences a particular emotion can be found in fiction. Contrastive analysis of how this process is reflected in different languages can be based on a comparison of a literary text with its translation into another language, since, in this case, both texts present the same character in the same situations that cause certain emotions. To exclude the influence of the translator’s personality, in our analysis we have used three different translations of selected passages from Dostoevsky’s The Idiot. A quantitative analysis of the means employed by the translators shows that representation of emotions in English does indeed reflect the way of perceiving the world that is typical of the national linguistic worldview as a whole. In all the three English texts, state predicates prevail over ac-tion predicates, and predicatively used adjectival words prevail over those used attributively. It means that emotional states are mostly perceived by English speakers as something that, while not permanent or inherent to a person, is, at the same time, static: less a process than a result of that process. In contrast, native speakers of Russian perceive emotional states as actions, and the Russian text reveals no inclination toward perceiving emotional states as personal characteristics, whether temporary or permanent. All these regularities are statistical and not absolute, which means that they reflect usage and not the linguistic norm, and, thus, the change of predicates in translation should be regarded as a way of cognitive adaptation rather than a structural transformation.
Cloud-based enterprise search services (e.g., AWS Kendra) have been entrancing big data owners by offering convenient and real-time search solutions to them. However, the problem is that individuals and organizations possessing confidential big data are hesitant to embrace such services due to valid data privacy concerns. In addition, to offer an intelligent search, these services access the user's search history that further jeopardizes his/her privacy. To overcome the privacy problem, the main idea of this research is to separate the intelligence aspect of the search from its pattern matching aspect. According to this idea, the search intelligence is provided by an on-premises edge tier and the shared cloud tier only serves as an exhaustive pattern matching search utility. We propose Smartness at Edge (SAED mechanism that offers intelligence in the form of semantic and personalized search at the edge tier while maintaining privacy of the search on the cloud tier. At the edge tier, SAED uses a knowledge-based lexical database to expand the query and cover its semantics. SAED personalizes the search via an RNN model that can learn the user's interest. A word embedding model is used to retrieve documents based on their semantic relevance to the search query. SAED is generic and can be plugged into existing enterprise search systems and enable them to offer intelligent and privacy-preserving search without enforcing any change on them. Evaluation results on two enterprise search systems under real settings and verified by human users demonstrate that SAED can improve the relevancy of the retrieved results by on average ≈24% for plain-text and ≈75% for encrypted generic datasets.
In this paper, we address the representation of coordinate constructions in Enhanced Universal Dependencies (UD), where relevant dependency links are propagated from conjunction heads to other conjuncts. English treebanks for enhanced UD have been created from gold basic dependencies using a heuristic rule-based converter, which propagates only core arguments. With the aim of determining which set of links should be propagated from a semantic perspective, we create a large-scale dataset of manually edited syntax graphs. We identify several systematic errors in the original data, and propose to also propagate adjuncts. We observe high inter-annotator agreement for this semantic annotation task. Using our new manually verified dataset, we perform the first principled comparison of rule-based and (partially novel) machine-learning based methods for conjunction propagation for English. We show that learning propagation rules is more effective than hand-designing heuristic rules. When using automatic parses, our neural graph-parser based edge predictor outperforms the currently predominant pipelines using a basic-layer tree parser plus converters.
The aim of this paper is to offer an insight into the semantic roles of adverbials. The approach is mainly construed around the theory of adverb semantics propounded by Quirk, Greenbaum, Leech, and Svartvik (1985) – grammatical functions and the realisation of semantic roles. The theoretical approach is complemented by a practical analysis of adverbial phrases occurring in social interactions (as well as script-based stage directions) from the TV series “Friends”. The main method used is corpus analysis; in addition, a semi-automated identification of adverbs was performed using both quantitative and qualitative analyses. The tools used were ConcApp software, as well as electronic dictionaries and lexical databases. A quantitative and qualitative analysis of -ly adverbials in the script was carried out to establish certain patterns of adverb occurrence in social interaction. The results reveal a large proportion of subjuncts, in particular emphasisers, intensifier subjuncts and downtoners (approximator) (in Greenbaum et al.’s taxonomy), or, in other taxonomies, speaker-oriented (Jackendoff 1972) / sentence adverbs (Swan 1988) / stance adverbs – attitude and epistemic (Biber et al. 1999). A second important finding is that the –ly adverbs used in this sitcom display high polysemy, including some novel semantic uses peculiar to present-day US English.
Labeling data can be an expensive task as it is usually performed manually by\ndomain experts. This is cumbersome for deep learning, as it is dependent on\nlarge labeled datasets. Active learning (AL) is a paradigm that aims to reduce\nlabeling effort by only using the data which the used model deems most\ninformative. Little research has been done on AL in a text classification\nsetting and next to none has involved the more recent, state-of-the-art Natural\nLanguage Processing (NLP) models. Here, we present an empirical study that\ncompares different uncertainty-based algorithms with BERT$_{base}$ as the used\nclassifier. We evaluate the algorithms on two NLP classification datasets:\nStanford Sentiment Treebank and KvK-Frontpages. Additionally, we explore\nheuristics that aim to solve presupposed problems of uncertainty-based AL;\nnamely, that it is unscalable and that it is prone to selecting outliers.\nFurthermore, we explore the influence of the query-pool size on the performance\nof AL. Whereas it was found that the proposed heuristics for AL did not improve\nperformance of AL; our results show that using uncertainty-based AL with\nBERT$_{base}$ outperforms random sampling of data. This difference in\nperformance can decrease as the query-pool size gets larger.\n