Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract Using the framework of Language Management Theory (LMT), this article seeks to analyze the ways in which non-native speakers negotiate their position in English-language online discussion forums. Based on the material collected from four discussion forums, competing opinions have been identified regarding the acceptability of “bad English” and the need for language management, i.e. acting upon a perceived lack of compliance to linguistic norms. Some users propose that compliance to communication norms should be enforced in a top-down manner or based on an explicit set of rules, whereas others hold that the community of users can deal with potential communication problems individually in an emergent manner. While the applicability of native speaker norms to the discussion forums is being questioned, non-native speakers, especially in practically oriented forums, tend to perform pre-interaction language management, using disclaimers in their posts, such as “Excuse my poor English”, to avoid potential misunderstandings and to prevent native speaker norms from being applied to them. The article argues for the use of LMT in computer mediated communication research, as it offers a dynamic view of the process in which rules, conventions and norms of online communication are being continuously discussed, negotiated and applied.
We present structured perceptron training for neural network transition-based dependency parsing. We learn the neural network representation using a gold corpus augmented by a large number of automatically parsed sentences. Given this fixed network representation, we learn a final layer using the structured perceptron with beam-search decoding. On the Penn Treebank, our parser reaches 94.26% unlabeled and 92.41% labeled attachment accuracy, which to our knowledge is the best accuracy on Stanford Dependencies to date. We also provide in-depth ablative analysis to determine which aspects of our model provide the largest gains in accuracy.
This is a bug-fix release. It corrects non-contiguous token numbering in CoNLL-X exports, as well as potentially problematic representation of whitespace in forms and lemmas in these files. It also corrects a number of instances of incorrect tokenization, where whitespace has ended up at the beginning or end of token forms.
Dependency parsers are among the most crucial tools in natural language processing as they have many important applications in downstream tasks such as information retrieval, machine translation and knowledge acquisition. We introduce the Yara Parser, a fast and accurate open-source dependency parser based on the arc-eager algorithm and beam search. It achieves an unlabeled accuracy of 93.32 on the standard WSJ test set which ranks it among the top dependency parsers. At its fastest, Yara can parse about 4000 sentences per second when in greedy mode (1 beam). When optimizing for accuracy (using 64 beams and Brown cluster features), Yara can parse 45 sentences per second. The parser can be trained on any syntactic dependency treebank and different options are provided in order to make it more flexible and tunable for specific tasks. It is released with the Apache version 2.0 license and can be used for both commercial and academic purposes. The parser can be found at https://github.com/yahoo/YaraParser.
Although the amygdala is a major locus for hedonic processing, how it encodes valence information is poorly understood. Given the hedonic potency of odor stimuli and the amygdala's anatomical proximity to the peripheral olfactory system, we combined high-resolution fMRI with pattern-based multivariate techniques to examine how valence information is encoded in the amygdala. Ten human subjects underwent fMRI scanning while smelling 9 odorants that systematically varied in perceived valence. Representational similarity analyses showed that amygdala codes the entire dimension of valence, ranging from pleasantness to unpleasantness. This unidimensional representation significantly correlated with self-reported valence ratings but not with intensity ratings. Furthermore, within-trial valence representations evolved over time, prioritizing earlier differentiation of unpleasant stimuli. Together, these findings underscore the idea that both spatial and temporal features uniquely encode pleasant and unpleasant odor valence in the amygdala. The availability of a unidimensional valence code in the amygdala, distributed in both space and time, would create greater flexibility in determining the pleasantness or unpleasantness of stimuli, providing a mechanism by which expectation, context, attention, and learning could influence affective boundaries for guiding behavior. SIGNIFICANCE STATEMENT: Our findings elucidate the mechanisms of affective processing in the amygdala by demonstrating that this brain region represents the entire valence dimension from pleasant to unpleasant. An important implication of this unidimensional valence code is that pleasant and unpleasant valence cannot coexist in the amygdale because overlap of fMRI ensemble patterns for these two valence extremes obscures their unique content. This functional architecture, whereby subjective valence maps onto a pattern continuum between pleasant and unpleasant poles, offers a robust mechanism by which context, expectation, and experience could alter the set-point for valence-based behavior. Finally, identification of spatial and temporal differentiation of valence in amygdala may shed new insights into individual differences in emotional responding, with potential relevance for affective disorders.
We consider the problem of learning general-purpose, paraphrastic sentence embeddings based on supervision from the Paraphrase Database (Ganitkevitch et al., 2013). We compare six compositional architectures, evaluating them on annotated textual similarity datasets drawn both from the same distribution as the training data and from a wide range of other domains. We find that the most complex architectures, such as long short-term memory (LSTM) recurrent neural networks, perform best on the in-domain data. However, in out-of-domain scenarios, simple architectures such as word averaging vastly outperform LSTMs. Our simplest averaging model is even competitive with systems tuned for the particular tasks while also being extremely efficient and easy to use. In order to better understand how these architectures compare, we conduct further experiments on three supervised NLP tasks: sentence similarity, entailment, and sentiment classification. We again find that the word averaging models perform well for sentence similarity and entailment, outperforming LSTMs. However, on sentiment classification, we find that the LSTM performs very strongly-even recording new state-of-the-art performance on the Stanford Sentiment Treebank. We then demonstrate how to combine our pretrained sentence embeddings with these supervised tasks, using them both as a prior and as a black box feature extractor. This leads to performance rivaling the state of the art on the SICK similarity and entailment tasks. We release all of our resources to the research community with the hope that they can serve as the new baseline for further work on universal sentence embeddings.
Arousal is essential in understanding human behavior and decision-making. In this work, we present a multimodal arousal rating framework that incorporates minimal set of vocal and non-verbal behavior descriptors. The rating framework and fusion techniques are unsupervised in nature to ensure that it can be readily-applicable and interpretable. Our proposed multimodal framework improves correlation to human judgment from 0.66 (vocal-only) to 0.68 (multimodal); analysis shows that the supervised fusion framework does not improve correlation. Lastly, an interesting empirical evidence demonstrates that the signal-based quantification of arousal achieves a higher agreement with each individual rater than the agreement among raters themselves. This further strengthens that machine-based rating is a viable way of measuring subjective humans' internal states through observing behavior features objectively.
This paper presents several modifications of the standard annotation projection algorithm for syntactic structures in crosslingual dependency parsing.Our approach reduces projection noise and includes efficient data sub-set selection techniques that have a substantial impact on parser performance in terms of labeled attachment scores.We test our techniques on data from the Universal Dependency Treebank and demonstrate the improvements on a number of language pairs.We also look at treebank translation including syntaxbased models and data combination techniques that push the performance even further.We achieve absolute improvements of up to over seven points in labeled attachment scores pushing the state-of-the art in cross-lingual dependency parsing for all language pairs tested in our experiments.
We examined the finding that aesthetic evaluations are more similar across observers for representational images than for abstract images. It has been proposed that a difference in convergence of observers' tastes is due to differing levels of shared semantic associations (Vessel & Rubin, 2010). In Experiment 1, student participants rated 20 representational and 20 abstract artworks. We found that their judgments were more similar for representational than abstract artworks. In Experiment 2, we replicated this finding, and also found that valence ratings given to associations and meanings provided in response to the artworks converged more across observers for representational than for abstract art. Our empirical work provides insight into processes that may underlie the observation that taste for representational art is shared across individual observers, while taste for abstract art is more idiosyncratic.
The increasing diversity of languages used on the web introduces a new level of complexity to Information Retrieval (IR) systems. We can no longer assume that textual content is written in one language or even the same language family. In this paper, we demonstrate how to build massive multilingual annotators with minimal human expertise and intervention. We describe a system that builds Named Entity Recognition (NER) annotators for 40 major languages using Wikipedia and Freebase. Our approach does not require NER human annotated datasets or language specific resources like treebanks, parallel corpora, and orthographic rules. The novelty of approach lies therein - using only language agnostic techniques, while achieving competitive performance. Our method learns distributed word representations (word embeddings) which encode semantic and syntactic features of words in each language. Then, we automatically generate datasets from Wikipedia link structure and Freebase attributes. Finally, we apply two preprocessing stages (oversampling and exact surface form matching) which do not require any linguistic expertise. Our evaluation is two fold: First, we demonstrate the system performance on human annotated datasets. Second, for languages where no gold-standard benchmarks are available, we propose a new method, distant evaluation, based on statistical machine translation.
Individuals with autism spectrum disorders (ASD) often have difficulty recognizing and interpreting facial expressions of emotion, which may impair their ability to navigate and communicate successfully in their social, interpersonal environments. Characterizing specific differences between individuals with ASD and their typically developing (TD) counterparts in the neural activity subserving their experience of emotional faces may provide distinct targets for ASD interventions. Thus we used functional magnetic resonance imaging (fMRI) and a parametric experimental design to identify brain regions in which neural activity correlated with ratings of arousal and valence for a broad range of emotional faces. Participants (51 ASD, 84 TD) were group-matched by age, sex, IQ, race, and socioeconomic status. Using task-related change in blood-oxygen-level-dependent (BOLD) fMRI signal as a measure, and covarying for age, sex, FSIQ, and ADOS scores, we detected significant differences across diagnostic groups in the neural activity subserving the dimension of arousal but not valence. BOLD-signal in TD participants correlated inversely with ratings of arousal in regions associated primarily with attentional functions, whereas BOLD-signal in ASD participants correlated positively with arousal ratings in regions commonly associated with impulse control and default-mode activity. Only minor differences were detected between groups in the BOLD signal correlates of valence ratings. Our findings provide unique insight into the emotional experiences of individuals with ASD. Although behavioral responses to face-stimuli were comparable across diagnostic groups, the corresponding neural activity for our ASD and TD groups differed dramatically. The near absence of group differences for valence correlates and the presence of strong group differences for arousal correlates suggest that individuals with ASD are not atypical in all aspects of emotion-processing. Studying these similarities and differences may help us to understand the origins of divergent interpersonal emotional experience in persons with ASD. Hum Brain Mapp 37:443-461, 2016. © 2015 Wiley Periodicals, Inc.
This article examines how media scholars’ attributes affect ratings of Journalism and Mass Communication Quarterly (JMCQ), based on the 2014 JMCQ readership survey. It compares the impacts of these attributes on five different types of subjective journal ratings. For instance, importance of journal impact factor in the respondent’s institution only affects the rating of the journal’s standing in the field. Attributes such as use and knowledge of the journal, research recognition received in the field, doctoral institution affiliation, and ethnicity consistently predict the rating of the journal’s standing in the field, overall ratings, and relevance to the respondents; but other attributes predict the ratings of the journal in serving the association members well and author-friendliness.
This study demonstrates the use of social media analytics in the context of network television (TV) programs. We first downloaded social media measures for 38 TV programs and their performance ratings over a period of five weeks resulting in a sample size of 165 weekly observations. Specifically we extracted the number of Twitter tweets, followers, followings, Facebook likes, and talk from the official Twitter and Facebook profiles of each TV program. Subsequently we applied OLS Regression techniques and determined that key social media measures positively affect ratings. In essence TV shows with a higher number of Twitter tweets, followers and Facebook talk are likely to associate with higher performance ratings. This study helps TV networks in realizing the pertinence of social media in garnering viewership. Consequently we also propose a social media analytics framework for businesses in identifying brands with higher social media buzz in the objective of improving future economic performance.
In this paper, the authors propose formalism for representing a knowledge base (KB) by network. The objective is to achieve a high coverage of this base. This type of network is similar to the semantic network with the difference that the arcs are quantified by a value indicating the semantic proximity between the concepts. This semantic proximity presents taxonomic relations, synonyms, and non-taxonomic relations (contextual relations). This latter are discovered based on the association rules model. This model is based on (i) indexing method (ii) the French lexical database EuroWordNet (EWNF) and (iii) the Apriori algorithm. The contextual relations are the latent relations buried in the KB, carried by the semantic context. Evaluating our representation formalism shows better result about 80% of coverage of the KB.
This communication is concerned with theoretical aspects of NLP, and with some ‘epistemological impediments ’ [1] to the development of Arabic generation and recognition programs or lexical databases. The discussion focuses on the methodological approach underlying the elaboration of specifications associated to the entries of an Arabic lexical database, in relation with the DIINAR.1 Arabic Language database and the DIINAR-MBC Euro-Mediterranean project1. Such specifications consist of morpho-semantic as well as syntactico-semantic features. Semantic aspects belong to the field of finite semantics. A definition of that field will be given in connection with the notion of ambiguity, which is revisited here in the context of formal linguistics and a cognitive approach of computational linguistics.
Re-ranking models of parse trees have been focused on re-ordering parse trees with a syntactic view. However, also a semantic view should be considered in re-ranking parse trees, because the fact that a word pair has a dependency implies that the pair has both syntactic and semantic relations. This paper proposes a re-ranking model for dependency parsing based on a combination of syntactic and semantic plausibilities of dependencies. The syntactic probability is used as a syntactic plausibility of a parse tree, and a knowledge graph embedding is adopted to represent its semantic plausibility. The knowledge graph embedding allows the semantic plausibility of parse trees to be expressed effectively with ease. The experiments on the standard Penn Treebank corpus prove that the proposed model improves the base parser regardless of the number of candidate parse trees.
Redundancy is an important psycholinguistic concept which is often used for explanations of language change, but is notoriously difficult to operationalize and measure. Assuming that the reconstruction of a syntactic structure by a parser can be used as a rough model of the understanding of a sentence by a human hearer, I propose a method for estimating redundancy. The key idea is to compare performances of a parser on a given treebank before and after artificially removing all information about a certain grammeme from the morphological annotation. The change in performance can be used as an estimate for the redundancy of the grammeme. I perform an experiment, applying MaltParser to an Old Church Slavonic treebank to estimate grammeme redundancy in Proto-Slavic. The results show that those Old Church Slavonic grammemes within the case, number and tense categories that were estimated as most redundant are those that disappeared in modern Russian. Moreover, redundancy estimates serve as a good predictor of case grammeme frequencies in modern Russian. The small sizes of the samples do not allow to make definitive conclusions for number and tense.
Paraphrase identification is a semantic text similarity task which is an important part of many natural language processing applications. Existing methods use vector space models, word co-occurrence information, lexical databases, parsers and machine translation (MT) evaluation metrics to find text similarity. However, other aspects such as negations, inverse relations and semantic roles of the sentences are also very much important in identifying paraphrases. Furthermore, the semantics of the sentences are hidden when the sentences are complex. We propose an approach to find similarity between pair of texts by considering all these factors. We have used an approach to determine set of clauses present in the texts by resolving conjunctions in complex sentences that identify hidden triples from the text. The approach extracts clause-based similarity features namely concept score, relation score, proposition score and word score from the texts. We have combined these similarity features along with MT metrics features to identify whether the texts are paraphrases or not using Support Vector Machine model. We have evaluated our methodology to measure the paraphrase similarity for Microsoft Research corpus. The statistical tests namely |$k$|-fold paired |$t$|-test and McNemar's test show that including clause-based features significantly improved the performance. Also, our approach outperforms state-of-the-art methods in terms of accuracy, |$F$|1-measure and |$f$|1-measure.
Previous studies indicate that emotion regulation may occur unconsciously, without the cost of cognitive efforts; and that conscious acceptance effectively reduces the emotional consequences of negative events. However, it has yet to be determined how conscious and unconscious acceptance strategies differ in behavioral and physiological consequences of emotion regulation. As unconscious regulation occurs with little cost of cognitive resources, the current study hypothesizes that unconscious acceptance regulates the emotional consequence of negative events more effectively compared to conscious acceptance. Subjects were randomly assigned to conscious acceptance, unconscious acceptance and control conditions. A frustrating arithmetic task was used to induce negative emotion. Emotional experiences were assessed by the positive affect and negative affect scale (PANAS) while emotion-related physiological activation was assessed by the heart-rate reactivity. The results showed that unconscious acceptance produced less reductions of positive affect ratings compared to conscious acceptance during frustration. In addition, both conscious and unconscious acceptance strategies significantly decreased emotion-related heart-rate activity (to a similar extent) in comparison with the control condition. Moreover, heart-rate reactivity showed a trend of positive correlation with negative affect rating and a trend of negative correlation with positive affect rating during frustration compared to baseline phases. Thus, unconscious acceptance is not only able to decrease emotion-related physiological activity, but also able to produce better emotional experiences compared to conscious acceptance. This suggests that it is practically important to consider unconscious acceptance for emotion regulation in real-life settings.
The spinal tree adjoining grammar (TAG) parsing model of [Carreras 08] achieves the current state-of-the-art constituent parsing accuracy on the commonly used English Penn Treebank evaluation setting. Unfortunately, the model has the serious drawback of low parsing efficiency since its Eisner-CKY style parsing algorithm needs O(n4) computation time for input length n. This paper investigates a more practical solution and presents a beam search shift-reduce algorithm for spinal TAG parsing. Since the algorithm works in O(bn) (b is beam width), it can be expected to provide a significant improvement in parsing speed. However, to achieve faster parsing, it needs to prune a large number of candidates in an exponentially large search space and often suffers from severe search errors. In fact, our experiments show that the basic beam search shift-reduce parser does not work well for spinal TAGs. To alleviate this problem, we extend the proposed shift-reduce algorithm with two techniques: Dynamic Programming of [Huang 10a] and Supertagging. The proposed extended parsing algorithm is about 8 times faster than the Berkeley parser, which is well-known to be fast constituent parsing software, while offering state-of-the-art performance. Moreover, we conduct experiments on the Keyaki Treebank for Japanese to show that the good performance of our proposed parser is language-independent.
Word order differences between source and target languages pose a serious challenge to statistical machine translation (SMT). Pre-ordering, an approach that reorders source words into a target-word-like order as a preprocessing step, has been shown effective in handling word order between different languages and improving translation performance of SMT. In this paper, we propose a novel word reordering method based on the pre-ordering framework. Instead of using a supervised parser trained on a monolingual treebank, our method extracts bilingual structural information for reordering from automatically wordaligned sentence pairs into dependency-tree-like structures, then learns a reordering model by training a dependency parser on this extracted pseudo-treebank. Experiment results show that our pre-ordering method is effective in permuting source words to resemble word order of the target language, and improving translation quality.
In communication a great deal of meaning is exchanged through body language, including gaze, posture, hand gestures and body movements. Body language is largely culture-specific, and rests, for its comprehension, on people's sharing socio-cultural and linguistic norms. In cross-cultural communication, L2 speakers' use of body language may convey meaning that is not understood or misinterpreted by the interlocutors, affecting the pragmatics of communication. In spite of its importance for cross-cultural communication, body language is neglected in ESL/EFL teaching. This paper argues that the study of body language should be integrated in the syllabus of ESL/EFL teaching and learning. This is done by: 1) reviewing literature showing the tight connection between language, speech and gestures and the problems that might arise in cross-cultural communication when speakers use and interpret body language according to different conventions; 2) reporting the data from two pilot studies showing that L2 learners transfer L1 gestures to the L2 and that these are not understood by native L2 speakers; 3) reporting an experience teaching body language in an ESL/EFL classroom. The paper suggests that in multicultural ESL/EFL classes teaching body language should be aimed primarily at raising the students' awareness of the differences existing across cultures.
In this paper we explore different statistical dependency parsers for parsing Telugu. We consider five popular dependency parsers namely, MaltParser, MSTParser, TurboParser, ZPar and Easy-First Parser. We experiment with different parser and feature settings and show the impact of different settings. We also provide a detailed analysis of the performance of all the parsers on major dependency labels. We report our results on test data of Telugu dependency treebank provided in the ICON 2010 tools contest on Indian languages dependency parsing. We obtain state-of-the art performance of 91.8% in unlabeled attachment score and 70.0% in labeled attachment score. To the best of our knowledge ours is the only work which explored all the five popular dependency parsers and compared the performance under different feature settings for Telugu.
Recently, dependency parsing has been used for development of dependency parsers. There are many parsers built in the area of NLP for grammatical information extraction. These parsers can be used to build treebanks which can serve as resources for research purposes. This paper describes the various parsers based on different parsing methodology for different languages. One of the advantages of the dependency parsing is that it resolves ambiguity. In this paper a comparative table of different parsers is proposed for better analysis.
Abstract This chapter concludes that Negritude allowed black poets to be full participants of the “aesthetic regime.” This “aesthetic regime” is a lyric regime, and the poetry of Negritude establishes itself solidly as a text-based (rather than oral) movement. Negritude poets not only adopted typographic innovations introduced by other writers; they also developed their own way of harnessing the resistant force that the printed word harbors in its material being. The poets of Negritude in this sense raced textuality. They drew on the complex specificities of the irracialization under modern capitalism to exert pressure on thematic, lexical prosodic, typographical, and rhetorical norms. Moreover, the Negritude poem offers the promise of an identity that can be performed but will never resolve into essence, the promise of an identity that acts like a resistant force of “materiality as it plays itself out in/as the work of art.”
Coordinate structures pose difficulties in dependency parsers. In this paper, we propose a set of parsing rules specifically designed to handle coordination, which are intended to be used in combination with Eisner and Satta's dependency rules. The new rules are compatible with existing similarity-based approaches to coordination structure analysis, and thus the syntactic and semantic similarity of conjuncts can be incorporated to the parse scoring function. Although we are yet to implement such a scoring function, we analyzed the time complexity of the proposed rules as well as their coverage of the Penn Treebank converted to the Stanford basic dependencies.
This paper proposes and describes a computational system for the automatic analysis of thematic structure, as defined in Systemic Functional Linguistics, in written English. The system takes an English text as input and produces as output an analysis of the thematic structure of each sentence in the text. The system is evaluated using data from The Wall Street Journal section of the Penn Treebank (Marcus et al. 1993) and the British Academic Written English corpus (Gardner & Nesi 2013). An experiment using these data shows that the system achieves a high degree of reliability in regard to both identifying theme-rheme boundaries and determining several of the linguistic properties of the identified themes, including syntactic nodes, theme function, markedness, mood types, and theme roles. To illustrate how the system is used, we describe an example application designed to compare collections of novice and expert academic writing in terms of thematic structure.
The paper presents interpretation of universal quantifier in Bulgarian language with Universal Networking Language (UNL). It analyzes semantic and linguistic properties of universal quantifier and describes the parts of speech which function as quantifiers. In UNL frameworks they are presented as a lexical database with related grammar and semantic features using idea of synonymic semantic representation. The formal approach presented may be used also to language learning and understanding.
INTRODUCTION: Although conceptual models of sexual functioning have suggested a major role for implicit cognitive processing in sexual functioning, this has thus far, only been investigated in women. AIM: The aim of this study was to investigate the role of implicit cognition in sexual functioning in men. METHODS: Men with (N = 29) and without sexual dysfunction (N = 31) were compared. MAIN OUTCOME MEASURES: Participants performed two single-target implicit association tests (ST-IAT), measuring the implicit association of visual erotic stimuli with attributes representing, respectively, valence ('liking') and motivation ('wanting'). Participants also rated the erotic pictures that were shown in the ST-IAT on the dimensions of valence, attractiveness, and sexual excitement to assess their explicit associations with these erotic stimuli. Participants completed the International Index of Erectile Functioning for a continuous measure of sexual functioning. RESULTS: Unexpectedly, compared with sexually functional men, sexually dysfunctional men were found to show stronger implicit associations of erotic stimuli with positive valence than with negative valence. Level of sexual functioning, however, was not predicted by explicit nor implicit associations. Level of sexual distress was predicted by explicit valence ratings, with positive ratings predicting higher levels of sexual distress. CONCLUSIONS: Men with and without sexual dysfunction differed significantly with regard to implicit liking. Research recommendations and implications are discussed.
Pain catastrophising is an exaggerated cognitive attitude implemented during pain or when thinking about pain. Catastrophising was previously associated with increased pain severity, emotional distress and disability in chronic pain patients, and is also a contributing factor in the development of neuropathic pain. To investigate the neural basis of how pain catastrophising affects pain observed in others, we acquired EEG data in groups of participants with high (High-Cat) or low (Low-Cat) pain catastrophising scores during viewing of pain scenes and graphically matched pictures not depicting imminent pain. The High-Cat group attributed greater pain to both pain and non-pain pictures. Source dipole analysis of event-related potentials during picture viewing revealed activations in the left (PHGL) and right (PHGR) paraphippocampal gyri, rostral anterior (rACC) and posterior cingulate (PCC) cortices. The late source activity (600-1100 ms) in PHGL and PCC was augmented in High-Cat, relative to Low-Cat, participants. Conversely, greater source activity was observed in the Low-Cat group during the mid-latency window (280-450 ms) in the rACC and PCC. Low-Cat subjects demonstrated a significantly stronger correlation between source activity in PCC and pain and arousal ratings in the long latency window, relative to high pain catastrophisers. Results suggest augmented activation of limbic cortex and higher order pain processing cortical regions during the late processing period in high pain catastrophisers viewing both types of pictures. This pattern of cortical activations is consistent with the distorted and magnified cognitive appraisal of pain threats in high pain catastrophisers. In contrast, high pain catastrophising individuals exhibit a diminished response during the mid-latency period when attentional and top-down resources are ascribed to observed pain.
With the development of internet, there are billions of short texts generated each day. However, the accuracy of large scale short text classification is poor due to the data sparseness. Traditional methods used to use external dataset to enrich the representation of document and solve the data sparsity problem. But external dataset which matches the specific short texts is hard to find. In this paper, we propose a framework to solve the data sparsity problem without using external dataset. Our framework deal with large scale short text by making the most of semantic similarity of words which learned from the training short texts. First, we learn word distributed representation and measure the word semantic similarity from the training short texts. Then, we propose a method which enrich the document representation by using the word semantic similarity information. At last, we build classifiers based on the enriched representation. We evaluate our framework on both the benchmark dataset(Standford Sentiment Treebank) and the large scale Chinese news title dataset which collected by ourselves. For the benchmark dataset, using our framework can improve 3% classification accuracy. The result we tested on the large scale Chinese news title dataset shows that our framework achieve better result with the increase of the training set size.
BACKGROUND: Previous studies reported that old adults, relative to young adults, showed improvement of emotional stability and increased experiences of positive affects. METHODS: In order to better understand the neural underpinnings behind the aging-related enhancement of positive affects, it is necessary to investigate whether old and young adults differ in the threshold of eliciting positive or negative emotional reactions. However, no studies have examined emotional reaction differences between old and young adults by manipulating the intensity of emotional stimuli to date. To clarify this issue, the present study examined the impact of aging on the brain's susceptibility to affective pictures of varying emotional intensities. We recorded event-related potentials (ERP) for highly negative (HN), mildly negative (MN) and neutral pictures in the negative experimental block; and for highly positive (HP), mildly positive (MP) and neutral pictures in the positive experimental block, when young and old adults were required to count the number of pictures, irrespective of the emotionality of the pictures. RESULTS: Event-related potentials results showed that LPP (late positive potentials) amplitudes were larger for HN and MN stimuli compared to neutral stimuli in young adults, but not in old adults. By contrast, old adults displayed larger LPP amplitudes for HP and MP relative to neutral stimuli, while these effects were absent for young adults. In addition, old adults reported more frequent perception of positive stimuli and less frequent perception of negative stimuli than young adults. The post-experiment stimulus assessment showed more positive ratings of Neutral and MP stimuli, and reduced arousal ratings of HN stimuli in old compared to young adults. CONCLUSION: These results suggest that old adults are more resistant to the impact of negative stimuli, while they are equipped with enhanced attentional bias for positive stimuli. The implications of these results to the aging-related enhancement of positive affects were discussed.
Limbic encephalitis (LE) is an autoimmune-mediated disorder that affects structures of the limbic system, in particular, the amygdala. The amygdala constitutes a brain area substantial for processing of emotional, especially fear-related signals. The amygdala is also involved in neuroendocrine and autonomic functions, including skin conductance responses (SCRs) to emotionally arousing stimuli. This study investigates behavioral and autonomic responses to discrete emotion evoking and neutral film clips in a patient suffering from LE associated with contactin-associated protein-2 (CASPR2) antibodies as compared to a healthy control group. Results show a lack of SCRs in the patient while watching the film clips, with significant differences compared to healthy controls in the case of fear-inducing videos. There was no comparable impairment in behavioral data (emotion report, valence, and arousal ratings). The results point to a defective modulation of sympathetic responses during emotional stimulation in patients with LE, probably due to impaired functioning of the amygdala.
In this work, we present a novel way of using neural network for graph-based dependency parsing, which fits the neural network into a simple probabilistic model and can be furthermore generalized to high-order parsing. Instead of the sparse features used in traditional methods, we utilize distributed dense feature representations for neural network, which give better feature representations. The proposed parsers are evaluated on English and Chinese Penn Treebanks. Compared to existing work, our parsers give competitive performance with much more efficient inference.
Categorial grammars are attractive because they have a clear account of unbounded dependencies. This accounting is especially important in Mandarin Chinese which makes extensive usage of unbounded dependencies. However, parsers trained on existing categorial grammar annotations (Tse and Curran, 2010) extracted from the Penn Chinese Treebank This work reannotates the Penn Chinese Treebank into a generalized categorial grammar which uses a larger rule set and a substantially smaller category set while retaining the capacity to model unbounded dependencies. Experimental results show a statistically significant improvement in parsing accuracy with this categorial grammar.
Dependency parsing has become an important line of research in natural language processing in recent years. This is due to its usefulness in a wide variety of real world applications. This paper presents the improvement of Vietnamese dependency parsing using distributed word representations. Our parser achieves an accuracy of 76.29% of unlabelled attachment score or 69.25% of labelled attachment score. This is the most accurate dependency parser for the Vietnamese language in comparison to others which are trained and tested on the same dependency treebank. The distributed word representations are produced by two recent unsupervised learning models, the Skip-gram model and the GloVe model. We also show that distributed representations produced by the GloVe model are better than those produced by the Skip-gram model when being used in dependency parsing. Our dependency parsing system, including software, corpus and distributed word representations, is released as an open source project, freely available for research purpose.
Following a comparison of the different views on lexical meaning conveyed by the Latin WordNet and by a treebank-based valency lexicon for Latin, the paper evaluates the degree of overlapping between a number of homogeneous lexical subsets extracted from the two resources.
The emotional habituation plays an important role in individuals' adaptation to the environment. The present study explored the brain's emotional habituation to positive and negative pictures of diverse emotional intensities. Event-related potentials (ERPs) were recorded in two different experimental sessions, for highly positive (HP), mildly positive (MP) and neutral picture and for highly negative (HN), mildly negative (MN) and neutral picture. Subjects were asked to perform a standard/deviant categorization task, irrespective of emotionality of the deviants. The behavior results showed that the arousal ratings for HP stimuli decreased significantly with stimulus repetition. In addition, the ERP results displayed earlier N1 peak latencies with stimulus repetition in the positive session. Furthermore, the size of the emotion effect, which was computed by the emotion-neutral differences, decreased significantly for HP and MP stimuli with stimulus repetition in P3 amplitudes. Conversely, the current study failed to observe an emotional habituation effect to negative stimuli in any behavioral or ERP indexes. These results suggest that the humans' emotional reactions to positive stimuli, irrespective of the emotional intensity, are susceptible to habituation, irrespective of information processing stage. However, the humans' emotional reactions to negative stimuli are resistant to habituation, irrespective of the emotional intensities of the stimuli and the information processing stage. This valence-specific habituation effect is independent of the emotional intensity of the stimuli.
Lexical semantic information plays an important role in supervised dependency parsing. In this paper, we add lexical semantic features to the feature set of a parser, obtaining improvements on the Penn Chinese Treebank. We extract semantic categories of words from HowNet, and use them as semantic information of words. Moreover, we investigate the method to compute semantic similarity between Chinese compound words, and obtain semantic information of words which did not record in HowNet. Our experiments show that unlabeled attachment scores can increase by 1.29%.
In this paper, we propose a baseline messagelevel sentiment classification method, as developed for SemEval-2015 Task 10, Subtask B. This system leverages both hand-crafted features and message-level embedding features, and uses an SVM classifier for messagelevel sentiment classification. In pre-training the embedding features, we use one million randomly-selected tweets. We present results over SemEval-2015 Task 10, Subtask B, as well as the Stanford Sentiment Treebank. Our experiments show the effectiveness of our method over both datasets.
Abstract This article measures the productivity index of the Old English suffixes -cund, -ful, and -isc as well as the prefix ful- and checks the results against the diachronic evolution of the affixes. The frameworks brought to the discussion include Type frequency measurement, as well as productivity indexes proposed by Baayen (1992, 1993, 2009) and Trips (2009). The sources are both textual (The Dictionary of Old English Corpus) and lexicographical (the lexical database of Old English Nerthus). The conclusion drawn is that Baayen's (1992, 1993, 2002) index of Global Productivity provides the most consistent results with the diachronic evolution of the affixes.
This paper presents a novel technique for empty category (EC) detection using distributed word representations. A joint model is learned from the labeled data to map both the distributed representations of the contexts of ECs and EC types to a low dimensional space. In the testing phase, the context of possible EC positions will be projected into the same space for empty category detection. Experiments on Chinese Treebank prove the effectiveness of the proposed method. We improve the precision by about 6 points on a subset of Chinese Treebank, which is a new state-ofthe-art performance on CTB.
News websites give their users the opportunity to participate in discussions about published articles, by writing comments. Typically, these comments are unstructured making it hard to understand the flow of user discussions. Thus, there is a need for organizing comments to help users to (1) gain more insights about news topics, and (2) have an easy access to comments that trigger their interests. In this work, we address the above problem by organizing comments around the entities and the aspects they discuss. More specifically, we propose an approach for entity and aspect extraction from user comments through the following contributions. First, we extend traditional Named-Entity Recognition approaches, using coreference resolution and external knowledge bases, to detect more occurrences of entities in comments. Second, we exploit part-of-speech tag, dependency tag, and lexical databases to extract explicit and implicit aspects around discussed entities. Third, we evaluate our entity and aspect extraction approach, on manually annotated data, showing that it highly increases precision and recall compared to baseline approaches.
We propose and evaluate the use of an affective-semantic model to expand the affective lexica of German, Greek, English, Spanish and Portuguese. Motivated by the assumption that semantic similarity implies affective similarity, we use word level semantic similarity scores as semantic features to estimate their corresponding affective scores. Various context-based semantic similarity metrics are investigated using contextual features that include both words and character n-grams. The model produces continuous affective ratings in three dimensions (valence, arousal and dominance) for all five languages, achieving consistent performance. We achieve classification accuracy (valence polarity task) between 85% and 91% for all five languages. For morphologically rich languages the proposed use of character n-grams is shown to improve performance.
Songs heard between the ages of 15 and 24 should be remembered better and have a stronger relationship to autobiographical memories when compared with music from other phases of life (“reminiscence bump effect”). Additionally, the proportion of music-evoked autobiographical memories (MEAMs) is at a maximum in these years of early adolescence and then declines up to the age of 60. In our study we tried both to replicate these important findings based on a German sample and to further investigate the influence of the affective characteristics of the songs on the frequency of participants’ autobiographical memories. In Experiment 1 a group of adults ( N = 48, M age = 67.1 years) listened to excerpts from 80, number-one, popular music hits from 1930 to 2010 and gave written self-reports on MEAMs. In Experiment 2 the affective characteristics were rated by another group of adults ( N = 22, M age = 66 years) and were used to predict the frequency of MEAMs. As a main result of Experiment 1, we confirmed the reminiscence bump and decline effect with a small effect size for the ratings of feelings evoked by the song and with a medium effect size for the song recognition performance of those songs released during the participants’ age range of 15 to 24 years. The total number of MEAMs was only marginally influenced by a memory bump and decline effect, and participants showed a significant proportion of MEAMs up to the fifth decade. Experiment 2 revealed that the affective ratings of the songs were unequally distributed over the two-dimensional emotion space unlike the average rate of MEAMs which was nearly equally distributed. In contrast to previous research, we therefore conclude that popular songs can be associated with autobiographical memory over five decades of life – independent of the affective character of the music.
The purpose of this paper is to present an approach to create semi-automatically ontology from Arabic texts. The whole process is supervised by a linguistic expert. Our involvement in this project focused on a lexical ontology, taking as model the WordNet ontology, and as input source, the “Arabic verbs” of a contemporary monolingual dictionary () /mζjm Alγny/ in the form a lexical database. The verb, pivot of a sentence, is our goal in creating concepts, by adopting the synset as our meaning representation model. The Markov clustering algorithm of a graph, generated by the defining verbs, obtained from the transitive closure, allowed us to detect similar verbs and to identify as well, for a given verbal entry, all of its synonyms. A tool has been implemented, and experiments have been carried out to evaluate and show efficiency of the proposed approach.
We introduce interpolation of trained MSTParser models as a resource combination method for multi-source delexicalized parser transfer. We present both an unweighted method, as well as a variant in which each source model is weighted by the similarity of the source language to the target language. Evaluation on the HamleDT treebank collection shows that the weighted model interpolation performs comparably to weighted parse tree combination method, while being computationally much less demanding.