Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This thesis details three contributions to the advancement of semantic-enriched parsing for English sentences: inventories of semantic relations covering three semantically ambiguous linguistic phenomena, large datasets annotated according to the inventories, and, finally, a suite of tools for semantically-enriched parsing built using the datasets. For the purposes of this thesis, semantically-enriched parsing is defined as the reconstruction of the underlying grammatical structure of text along with shallow semantic annotation of semantically-ambiguous structures. Ultimately, semantically-enriched parsing is one of the most critical steps in natural language understanding---the initial step in which the text is read by the machine into a knowledge representation for further processing and reasoning. ❧ The first contribution of this thesis is to advance the theoretical foundations for the interpretation of three ambiguous linguistic phenomena in English that have significant overlap in terms of the relations expressed: noun compounds, possessive constructions, and prepositions. For these, I define inventories of relations based upon extensive annotation by myself, previous work by others, and inter-annotator agreement studies. In the case of prepositions, the relations are created by refining an existing resource whereas the other two are created from scratch. In addition to mappings to prior work, mappings are provided across the different inventories in order to create a unified set of relations. ❧ Second, I produce large datasets annotated according to the aforementioned sense inventories. Such data is vital for training most automatic tools and also provides exemplars for the theory embodied in the inventories. Some of these datasets are created from scratch, including a collection of over 17,500 noun compounds and a collection of over 21,900 possessive construction examples. In the case of prepositions, an existing resource including over 24,000 annotated examples is refined. ❧ The final contribution is a suite of tools that can construct semantically-enriched parse trees. The suite is designed to work in a sequential, pipeline-like fashion and can be thought of as consisting of two subsections. The first part reconstructs the grammatical structure of the text using a dependency parser that extends the non-directional easy-first algorithm developed by Goldberg and Elhadad in order to support non-projective trees and is trained using my improved dependency tree conversion of the Penn Treebank. Second are semantic annotation modules that add shallow semantic annotation for noun compounds, preposition senses, possessives, and verbal arguments. Combined, these tools produce semantically-enriched parse trees that include both grammatical structure and shallow semantics. The core parser itself achieves state-of-the-art accuracy and can process over \parsespeed sentences per second, which is substantially faster than most of the accurate parsers available today. ❧ In conclusion, this thesis work provides significant contributions to computational linguistics, both in terms of theory and resources. It advances our understanding of the relations expressed by three semantically-ambiguous linguistic phenomena, creates large annotated datasets useful for machine learning, and produces a fast, accurate, and informative system for semantically-enriched parsing.
The goal of this work is to bring semantics into the tasks of text recognition and retrieval in natural images. Although text recognition and retrieval have received a lot of attention in recent years, previous works have focused on recognizing or retrieving exactly the same word used as a query, without taking the semantics into consideration. In this paper, we ask the following question: \emph{can we predict semantic concepts directly from a word image, without explicitly trying to transcribe the word image or its characters at any point?} For this goal we propose a convolutional neural network (CNN) with a weighted ranking loss objective that ensures that the concepts relevant to the query image are ranked ahead of those that are not relevant. This can also be interpreted as learning a Euclidean space where word images and concepts are jointly embedded. This model is learned in an end-to-end manner, from image pixels to semantic concepts, using a dataset of synthetically generated word images and concepts mined from a lexical database (WordNet). Our results show that, despite the complexity of the task, word images and concepts can indeed be associated with a high degree of accuracy
The usage of phrasemes evidences not only their variability, transformations and modifications, but also the most frequent forms of their realization (phraseme-types) and frequency (phraseme-tokens), i.e. phrasemes’ flexibility. In this paper, selected Lithuanian idiomatic predicate phrasemes are analysed in the Corpus of Contemporary Lithuanian Language, in the Phraseological Dictionary and in the lexical database of the Dictionary of Lithuanian Phrases. The results of comparison show that the corpus research can give rich evidence about the morphological flexibility of phrasemes. This information can help to improve representation of phrasemes in the phraseological dictionaries of Lithuanian, in order to make them more usage-based and more usage-oriented.
We propose a model of Tibetan syntactic parsing which is based on Tibetan syl lables instead of Tibetan words, change Tibetan syntactic Treebank use algorith m of labeling.
Recent work on joint word segmentation, POS (Part Of Speech) tagging, and dependency parsing in Chinese has two key problems: the first is that word segmentation based on character and dependency parsing based on word were not combined well in the transition-based framework, and the second is that the joint model suffers from the insufficiency of annotated corpus. In order to resolve the first problem, we propose to transform the traditional word-based dependency tree into character-based dependency tree by using the internal structure of words and then propose a novel character-level joint model for the three tasks. In order to resolve the second problem, we propose a novel semi-supervised joint model for exploiting n-gram feature and dependency subtree feature from partially-annotated corpus. Experimental results on the Chinese Treebank show that our joint model achieved 98.31%, 94.84% and 81.71% for Chinese word segmentation, POS tagging, and dependency parsing, respectively. Our model outperforms the pipeline model of the three tasks by 0.92%, 1.77% and 3.95%, respectively. Particularly, the F1 value of word segmentation and POS tagging achieved the best result compared with those reported until now.
This paper describes our system designed for the NLPCC 2015 shared task on Chinese word segmentation (WS) and POS tagging for Weibo Text. We treat WS and POS tagging as two separate tasks and use a cascaded approach. Our major focus is how to effectively exploit multiple heterogeneous data to boost performance of statistical models. This work considers three sets of heterogeneous data, i.e., Weibo ($$\textit{WB}$$, 10K sentences), Penn Chinese Treebank 7.0 ($$\textit{CTB7}$$, 50K), and People’s Daily ($$\textit{PD}$$, 280K). For WS, we adopt the recently proposed coupled sequence labeling to combine $$\textit{WB}$$, $$\textit{CTB7}$$, and $$\textit{PD}$$, boosting F1 score from $$93.76\%$$ (baseline model trained on only $$\textit{WB}$$) to $$95.58\%$$ ($$+1.82\%$$). For POS tagging, we adopt an ensemble approach combining coupled sequence labeling and the guide-feature based method, since the three datasets have three different annotation standards. First, we convert $$\textit{PD}$$ into the annotation style of $$\textit{CTB7}$$ based on coupled sequence labeling, denoted by $$\textit{PD}^{\textit{CTB}}$$. Then, we merge CTB7 and $$\textit{PD}^{\textit{CTB}}$$ to train a POS tagger, denoted by $$\textit{Tag}_{\textit{CTB7}+\textit{PD}^{\textit{CTB}}}$$, which is further used to produce guide features on $$\textit{WB}$$. Finally, the tagging F1 score is improved from 87.93% to 88.99% (+1.06%).
We describe automatic conversion of the SynTagRus dependency treebank\nof Russian to the PROIEL format (with the ultimate purpose of obtaining a single-format\ndiachronic treebank spanning more than a thousand years), focusing\non analysis of shared arguments in verbal coordinations. Whether arguments\nare shared or private is not marked in the SynTagRus native format,\nbut the PROIEL format indicates sharing by means of secondary dependencies.\nIn order to recover missing information and insert secondary dependencies\ninto the converted SynTagRus, we create a simple guessing algorithm\nbased on four probabilistic features: how likely a given argument type\nis to be shared; how likely an argument in a given position is to be shared;\nhow likely a given verb is to have a given argument; how likely a given verb\nis to have a given argument frame. Boosted with a few deterministic rules and\ntrained on a small manually annotated sample (346 sentences), the guesser\nvery successfully inserts shared subjects (F-score 0.97), which results\nin excellent overall performance (F-score 0.92). Non-subject arguments are\nshared much more rarely, and for them the results are poorer (0.31 for objects;\n0.22 for obliques). We show, however, that there are strong reasons\nto believe that performance can be increased if a larger training sample\nis used and the guesser gets to see enough positive examples. Apart from\ndescribing a useful practical solution, the paper also provides quantitative\ndata about and offers non-trivial insights into Russian verbal coordination.
We present an analysis of a treebank of spontaneous English dyadic conversations, investigating whether the degree of syntactic priming found across speakers is a function of the degrees of affective alignment and overall positivity of the speakers. We use information theory to measure the proportion of overlap between the syntactic structures of the speakers. The affective state of the speakers is indexed by aggregated measures of the affective valences of the words they use. We find that there is a positive relation between syntactic priming and affective alignment, over and above any lexical repetition effects. This constitutes evidence for the percolation of inter-speaker alignment across multiple levels of representation. This also illustrates the indexical value of syntactic alignment, as has been proposed in modern functional theories of grammar such as Dialogic Syntax.
Based on the annotation of discourse connective in Chinese Discourse Treebank, especially the annotation of the connective and its relation classification. The authors extract syntax, lexical and position features of automatic syntax tree and standard syntax tree, and use supervised method to recognize and classify connective. Experimental results show that connective recognition F1-measure is 69.2%, and connective classification accuracy is 89.1%.
This study investigated sensory characteristics and cross-cultural consumer acceptability of sweet crispy chicken (Dakgangjeong) prepared with six types of Korean-style sauces among Korean and Chinese consumers. The main ingredient(s) of each sauce was soy sauce (SOY), Japanese apricot extract and soy sauce (JASOY), gochujang (SPICY), minced garlic (GARL), and ketchup (KET-I and KET-II); KET-I and KET-II were modified to possess ethnic Korean flavors. In general, Korean and Chinese consumers preferred all types of Dakgangjeong, except for GARL and SPICY, respectively. Least preferred products of each country had the lowest familiarity rating among consumers of the respective countries. Similar to previous studies, these results showed that familiarity is an important factor affecting consumer preference in a cross-cultural context. Particularly, it was found that higher familiarity of the product was not found to influence consumer to like a product, but rather low familiarity seemed to affect consumers to reject a product.
Arabic lexical language resources in a cross-lingual perspective are a major issue in the development of NLP applications, in the context of present-day web interactive developments. Much has already been done, and some existing achievements will be mentioned. The bulk of the challenge, though, still remains ahead. The main contribution of this paper is methodological. The first section will endeavour to show some of the specific needs for an elaborated lexical database in Arabic, in relation with the structure of the writing system (‘unvowelled ’ writing and the structure of word-forms). Grammar-lexis relations are an essential part of any Arabic lexical resource. These relations operate at word-level in morphological analysers, and at sentence-level in morphosyntactic ones. Sets of finite and exhaustive morphosyntactic specifiers have already been elaborated at both word- and sentence-level, but have only been implemented at word-level. The second section is concerned with the cross-lingual aspects of lexical LR-s including Arabic. The author underlines the necessity, in the present state of the Art, of considering ‘multilingual ’ or ‘cross-lingual ’ lexical databases including Arabic as sets of oriented bilingual LR-s, in which one of the languages under consideration is the ‘source language ’ and the other, the ‘target language’. It is essential for the future of Arabic linguistic engineering that sets of bilingual resources of this type, including the morphosyntactic specifiers outlined in section 1, be elaborated and implemented in the coming years, and that a substantial number of them take Arabic as a ‘source language’.
In this paper, we analyze the errors of NP attachments that occur when they combine with NP's ending in possessor clitic. We suggest a simple pattern based detection and correction solution. We illustrate the errors with examples from Penn Treebank.
Limbic encephalitis (LE) is an autoimmune-mediated disorder that affects structures of the limbic system, in particular, the amygdala. The amygdala constitutes a brain area substantial for processing of emotional, especially fear-related signals. The amygdala is also involved in neuroendocrine and autonomic functions, including skin conductance responses (SCRs) to emotionally arousing stimuli. This study investigates behavioral and autonomic responses to discrete emotion evoking and neutral film clips in a patient suffering from LE associated with contactin-associated protein-2 (CASPR2) antibodies as compared to a healthy control group. Results show a lack of SCRs in the patient while watching the film clips, with significant differences compared to healthy controls in the case of fear-inducing videos. There was no comparable impairment in behavioral data (emotion report, valence, and arousal ratings). The results point to a defective modulation of sympathetic responses during emotional stimulation in patients with LE, probably due to impaired functioning of the amygdala.
In this work, we present a novel way of using neural network for graph-based dependency parsing, which fits the neural network into a simple probabilistic model and can be furthermore generalized to high-order parsing. Instead of the sparse features used in traditional methods, we utilize distributed dense feature representations for neural network, which give better feature representations. The proposed parsers are evaluated on English and Chinese Penn Treebanks. Compared to existing work, our parsers give competitive performance with much more efficient inference.
Categorial grammars are attractive because they have a clear account of unbounded dependencies. This accounting is especially important in Mandarin Chinese which makes extensive usage of unbounded dependencies. However, parsers trained on existing categorial grammar annotations (Tse and Curran, 2010) extracted from the Penn Chinese Treebank This work reannotates the Penn Chinese Treebank into a generalized categorial grammar which uses a larger rule set and a substantially smaller category set while retaining the capacity to model unbounded dependencies. Experimental results show a statistically significant improvement in parsing accuracy with this categorial grammar.
Dependency parsing has become an important line of research in natural language processing in recent years. This is due to its usefulness in a wide variety of real world applications. This paper presents the improvement of Vietnamese dependency parsing using distributed word representations. Our parser achieves an accuracy of 76.29% of unlabelled attachment score or 69.25% of labelled attachment score. This is the most accurate dependency parser for the Vietnamese language in comparison to others which are trained and tested on the same dependency treebank. The distributed word representations are produced by two recent unsupervised learning models, the Skip-gram model and the GloVe model. We also show that distributed representations produced by the GloVe model are better than those produced by the Skip-gram model when being used in dependency parsing. Our dependency parsing system, including software, corpus and distributed word representations, is released as an open source project, freely available for research purpose.
Following a comparison of the different views on lexical meaning conveyed by the Latin WordNet and by a treebank-based valency lexicon for Latin, the paper evaluates the degree of overlapping between a number of homogeneous lexical subsets extracted from the two resources.
The emotional habituation plays an important role in individuals' adaptation to the environment. The present study explored the brain's emotional habituation to positive and negative pictures of diverse emotional intensities. Event-related potentials (ERPs) were recorded in two different experimental sessions, for highly positive (HP), mildly positive (MP) and neutral picture and for highly negative (HN), mildly negative (MN) and neutral picture. Subjects were asked to perform a standard/deviant categorization task, irrespective of emotionality of the deviants. The behavior results showed that the arousal ratings for HP stimuli decreased significantly with stimulus repetition. In addition, the ERP results displayed earlier N1 peak latencies with stimulus repetition in the positive session. Furthermore, the size of the emotion effect, which was computed by the emotion-neutral differences, decreased significantly for HP and MP stimuli with stimulus repetition in P3 amplitudes. Conversely, the current study failed to observe an emotional habituation effect to negative stimuli in any behavioral or ERP indexes. These results suggest that the humans' emotional reactions to positive stimuli, irrespective of the emotional intensity, are susceptible to habituation, irrespective of information processing stage. However, the humans' emotional reactions to negative stimuli are resistant to habituation, irrespective of the emotional intensities of the stimuli and the information processing stage. This valence-specific habituation effect is independent of the emotional intensity of the stimuli.
Lexical semantic information plays an important role in supervised dependency parsing. In this paper, we add lexical semantic features to the feature set of a parser, obtaining improvements on the Penn Chinese Treebank. We extract semantic categories of words from HowNet, and use them as semantic information of words. Moreover, we investigate the method to compute semantic similarity between Chinese compound words, and obtain semantic information of words which did not record in HowNet. Our experiments show that unlabeled attachment scores can increase by 1.29%.
In this paper, we propose a baseline messagelevel sentiment classification method, as developed for SemEval-2015 Task 10, Subtask B. This system leverages both hand-crafted features and message-level embedding features, and uses an SVM classifier for messagelevel sentiment classification. In pre-training the embedding features, we use one million randomly-selected tweets. We present results over SemEval-2015 Task 10, Subtask B, as well as the Stanford Sentiment Treebank. Our experiments show the effectiveness of our method over both datasets.
Abstract This article measures the productivity index of the Old English suffixes -cund, -ful, and -isc as well as the prefix ful- and checks the results against the diachronic evolution of the affixes. The frameworks brought to the discussion include Type frequency measurement, as well as productivity indexes proposed by Baayen (1992, 1993, 2009) and Trips (2009). The sources are both textual (The Dictionary of Old English Corpus) and lexicographical (the lexical database of Old English Nerthus). The conclusion drawn is that Baayen's (1992, 1993, 2002) index of Global Productivity provides the most consistent results with the diachronic evolution of the affixes.
This paper presents a novel technique for empty category (EC) detection using distributed word representations. A joint model is learned from the labeled data to map both the distributed representations of the contexts of ECs and EC types to a low dimensional space. In the testing phase, the context of possible EC positions will be projected into the same space for empty category detection. Experiments on Chinese Treebank prove the effectiveness of the proposed method. We improve the precision by about 6 points on a subset of Chinese Treebank, which is a new state-ofthe-art performance on CTB.
News websites give their users the opportunity to participate in discussions about published articles, by writing comments. Typically, these comments are unstructured making it hard to understand the flow of user discussions. Thus, there is a need for organizing comments to help users to (1) gain more insights about news topics, and (2) have an easy access to comments that trigger their interests. In this work, we address the above problem by organizing comments around the entities and the aspects they discuss. More specifically, we propose an approach for entity and aspect extraction from user comments through the following contributions. First, we extend traditional Named-Entity Recognition approaches, using coreference resolution and external knowledge bases, to detect more occurrences of entities in comments. Second, we exploit part-of-speech tag, dependency tag, and lexical databases to extract explicit and implicit aspects around discussed entities. Third, we evaluate our entity and aspect extraction approach, on manually annotated data, showing that it highly increases precision and recall compared to baseline approaches.
We propose and evaluate the use of an affective-semantic model to expand the affective lexica of German, Greek, English, Spanish and Portuguese. Motivated by the assumption that semantic similarity implies affective similarity, we use word level semantic similarity scores as semantic features to estimate their corresponding affective scores. Various context-based semantic similarity metrics are investigated using contextual features that include both words and character n-grams. The model produces continuous affective ratings in three dimensions (valence, arousal and dominance) for all five languages, achieving consistent performance. We achieve classification accuracy (valence polarity task) between 85% and 91% for all five languages. For morphologically rich languages the proposed use of character n-grams is shown to improve performance.
Songs heard between the ages of 15 and 24 should be remembered better and have a stronger relationship to autobiographical memories when compared with music from other phases of life (“reminiscence bump effect”). Additionally, the proportion of music-evoked autobiographical memories (MEAMs) is at a maximum in these years of early adolescence and then declines up to the age of 60. In our study we tried both to replicate these important findings based on a German sample and to further investigate the influence of the affective characteristics of the songs on the frequency of participants’ autobiographical memories. In Experiment 1 a group of adults ( N = 48, M age = 67.1 years) listened to excerpts from 80, number-one, popular music hits from 1930 to 2010 and gave written self-reports on MEAMs. In Experiment 2 the affective characteristics were rated by another group of adults ( N = 22, M age = 66 years) and were used to predict the frequency of MEAMs. As a main result of Experiment 1, we confirmed the reminiscence bump and decline effect with a small effect size for the ratings of feelings evoked by the song and with a medium effect size for the song recognition performance of those songs released during the participants’ age range of 15 to 24 years. The total number of MEAMs was only marginally influenced by a memory bump and decline effect, and participants showed a significant proportion of MEAMs up to the fifth decade. Experiment 2 revealed that the affective ratings of the songs were unequally distributed over the two-dimensional emotion space unlike the average rate of MEAMs which was nearly equally distributed. In contrast to previous research, we therefore conclude that popular songs can be associated with autobiographical memory over five decades of life – independent of the affective character of the music.
The purpose of this paper is to present an approach to create semi-automatically ontology from Arabic texts. The whole process is supervised by a linguistic expert. Our involvement in this project focused on a lexical ontology, taking as model the WordNet ontology, and as input source, the “Arabic verbs” of a contemporary monolingual dictionary () /mζjm Alγny/ in the form a lexical database. The verb, pivot of a sentence, is our goal in creating concepts, by adopting the synset as our meaning representation model. The Markov clustering algorithm of a graph, generated by the defining verbs, obtained from the transitive closure, allowed us to detect similar verbs and to identify as well, for a given verbal entry, all of its synonyms. A tool has been implemented, and experiments have been carried out to evaluate and show efficiency of the proposed approach.
In the present study, we raised the question of whether valence information of natural emotional sounds can be extracted rapidly and unintentionally. In a first experiment, we collected explicit valence ratings of brief natural sound segments. Results showed that sound segments of 400 and 600 ms duration-and with some limitation even sound segments as short as 200 ms-are evaluated reliably. In a second experiment, we introduced an auditory version of the affective Simon task to assess automatic (i.e. unintentional and fast) evaluations of sound valence. The pattern of results indicates that affective information of natural emotional sounds can be extracted rapidly (i.e. after a few hundred ms long exposure) and in an unintentional fashion.
The purpose of this study was to reveal the effects of Westernized arrangements of traditional Korean folk music on music familiarity and preference. Two separate labs in one intact class were assigned to one of two treatment groups of either listening to traditional Korean folk songs ( n = 18) or listening to Western arrangements of the same Korean folk songs ( n = 22); a second intact class served as a control group with no listening ( n = 20). Before and after the listening treatment session, pre- and posttests were administered that included 12 music excerpts of current popular, Western classical, and traditional Korean music. Results showed that participants who listened to traditional folk songs demonstrated significant increases in both familiarity and preference ratings; however, those who listened to Westernized folk songs showed increases only in familiarity ratings but not preference ratings for the same Korean songs in traditional versions. An analysis of participants’ open-ended responses showed that affective–positive responses were used most frequently when explaining preference for traditional versions of Korean folk songs (28.1%) among the traditional Korean listening group; structural–negative reasons (47.8%) were the most frequent among the Westernized listening group.
Using functional near-infrared spectroscopy, the present study investigated how listening to differently valenced music is associated with changes in hemoglobin concentrations in the prefrontal cortex area, indicating changes in neural activity. Thirty healthy people (15 men; M age = 24.8 yr., SD = 2.4; 15 women; M age = 25.2 yr., SD = 3.1) participated. Prefrontal cortex activation, emotional responses (heart rate variability), and self-reported affective ratings were measured while listening to calm and motivational music. The songs were presented in a random counterbalanced order and separated by periods of white noise. Mixed-model repeated-measures analysis of variance (ANOVA) evaluated the relationships for main effects and interactions. The results showed that music was associated with increased activation of the prefrontal cortex area. For both sexes, listening to the motivational song was associated with higher vagal withdrawal (lower HR) than the calm song. As expected, participants rated the motivational song with greater affective valence and higher arousal. Effects persisted longer in men than in women. These findings suggest that both the characteristics of music and sex differences may significantly affect the results of emotional neuroimaging in samples of young adults.
Prior research with four-part analogies suggests that people can detect that a novel word pair (e.g., "beaver:dam") maps analogically onto a pair in memory (e.g., "robin:nest") despite being unable to retrieve the pair from memory that is driving that detection. The present study demonstrates that the same type of detection during retrieval failure can occur when a story is used at test to illustrate a common aphorism that failed to be retrieved from an earlier list (e.g., "The squeaky wheel gets the grease" or "A watched pot never boils"). Given that prior research has suggested that analogy is related to insight, the present study also examined if such analogical detection during retrieval failure is related to the sense of presque vu, which is a term used to describe the subjective sense of an impending insight or discovery. Reports of presque vu during retrieval failure were associated with higher familiarity ratings. Participants were most likely to report presque vu after failing to identify the aphorism on the first attempt but before succeeding on the second attempt (relative to succeeding on the first attempt or failing altogether on both attempts). Additionally, instances of successful identification on the second attempt after a failed first attempt were more likely when presque vu was reported than when it was not. These patterns suggest that reports of presque vu may indicate impending retrieval of as yet unretrieved relevant information. However, instances of successful second attempt identification after initial failure occurred too infrequently to fully examine whether analogical resemblance to an unretrieved studied aphorism interacted with these patterns.
Sleep supports the consolidation of declarative memory in children and adults. However, it is unclear whether sleep improves odor memory in children as well as adults. Thirty healthy children (mean age of 10.6, ranging from 8-12 yrs.) and 30 healthy adults (mean age of 25.4, ranging from 20-30 yrs.) participated in an incidental odor recognition paradigm. While learning of 10 target odorants took place in the evening and retrieval (10 target and 10 distractor odorants) the next morning in the sleep groups (adults: n = 15, children: n = 15), the time schedule was vice versa in the wake groups (n = 15 each). During encoding, adults rated odors as being more familiar. After the retention interval, adult participants of the sleep group recognized odors better than adults in the wake group. While children in the wake group showed memory performance comparable to the adult wake group, the children sleep group performed worse than adult and children wake groups. Correlations between memory performance and familiarity ratings during encoding indicate that pre-experiences might be critical in determining whether sleep improves or worsens memory consolidation.
We introduce interpolation of trained MSTParser models as a resource combination method for multi-source delexicalized parser transfer. We present both an unweighted method, as well as a variant in which each source model is weighted by the similarity of the source language to the target language. Evaluation on the HamleDT treebank collection shows that the weighted model interpolation performs comparably to weighted parse tree combination method, while being computationally much less demanding.
We examined the potential cost of practicing suppression of negative thoughts on subsequent performance in an unrelated task. Cues for previously suppressed and unsuppressed (baseline) responses in a think/no-think procedure were displayed as irrelevant flankers for neutral words to be judged for emotional valence. These critical flankers were homographs with one negative meaning denoted by their paired response during learning. Responses to the targets were delayed when suppression cues (compared with baseline cues and new negative homographs) were used as flankers, but only following direct-suppression instructions and not when benign substitutes had been provided to aid suppression. On a final recall test, suppression-induced forgetting following direct suppression and the flanker task was positively correlated with the flanker effect. Experiment 2 replicated these findings. Finally, valence ratings of neutral targets were influenced by the valence of the flankers but not by the prior role of the negative flankers.
With a dependency grammar, this study provides a unified method for calculating the syntactic complexity in linear and hierarchical dimensions. Two metrics, mean dependency distance (MDD) and mean hierarchical distance (MHD), one for each dimension, are adopted. Some results from the Czech-English dependency treebank are revealed: (1) Positive asymmetries in the distributions of the two metrics are observed in English and Czech, which indicates both languages prefer the minimalization of structural complexity in each dimension. (2) There are significantly positive correlations between sentence length (SL), MDD, and MHD. For longer sentences, English prefers to increase the MDD, while Czech tends to enhance the MHD. (3) A trade-off relationship of syntactic complexity in two dimensions is shown between the two languages. English tends to reduce the complexity of production in the hierarchical dimension, whereas Czech prefers to lessen the processing load in the linear dimension. (4) The threshold of the MDD2 and MHD2 in English and
Abstract This work presents the development and evaluation of an extended Urdu parser. It further focuses on issues related to this parser and describes the changes made in the Earley algorithm to get accurate and relevant results from the Urdu parser. The parser makes use of a morphologically rich context free grammar extracted from a linguistically-rich Urdu treebank. This grammar with sufficient encoded information is comparable with the state-of-the-art parsing requirements for the morphologically rich Urdu language. The extended parsing model and the linguistically rich extracted-grammar both provide us better evaluation results in Urdu/Hindi parsing domain. The parser gives 87% of f-score, which outperforms the existing parsing work of Urdu/Hindi based on the tree-banking approach.
Universal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008). This is the second release of UD Treebanks, Version 1.1.
This paper proposes neural networks for integrating compositional and non-compositional sentiment in the process of sentiment composition, a type of semantic composition that optimizes a sentiment objective. We enable individual composition operations in a recursive process to possess the capability of choosing and merging information from these two types of sources. We propose our models in neural network frameworks with structures, in which the merging parameters can be learned in a principled way to optimize a well-defined objective. We conduct experiments on the Stanford Sentiment Treebank and show that the proposed models achieve better results over the model that lacks this ability.
Text-based sentiment analysis is a growing research field in affective computing, driven by both commercial applications and academic interest. Continuous dimensional representations, such as valence-arousal (VA) space, can represent the affective state more precisely than discrete effective representations. In building dimensional sentiment applications, affective lexicons with valence-arousal ratings are useful resources but are still very rare. Therefore, recent studies have investigated the automatic development of VA lexicons using linear regression techniques. One of the major limitations of linear regression is the under-fitting problem which can cause a poor fit between the algorithm and the training data. To tackle this problem, this study proposes the use of a locally weighted linear regression (LWLR) model to predict the valence-arousal ratings of affective words. The locally weighted method performs a regression around the point of interest using only training data that are "local" to that point, and thus can reduce the impact of noise from unrelated training data. Experimental results show that the proposed method achieved better performance for VA word prediction.
The goal of this work is to bring semantics into the tasks of text recognition and retrieval in natural images. Although text recognition and retrieval have received a lot of attention in recent years, previous works have focused on recognizing or retrieving exactly the same word used as a query, without taking the semantics into consideration. In this paper, we ask the following question: can we predict semantic concepts directly from a word image, without explicitly trying to transcribe the word image or its characters at any point? For this goal we propose a convolutional neural network (CNN) with a weighted ranking loss objective that ensures that the concepts relevant to the query image are ranked ahead of those that are not relevant. This can also be interpreted as learning a Euclidean space where word images and concepts are jointly embedded. This model is learned in an end-to-end manner, from image pixels to semantic concepts, using a dataset of synthetically generated word images and concepts mined from a lexical database (WordNet). Our results show that, despite the complexity of the task, word images and concepts can indeed be associated with a high degree of accuracy.
This paper describes the submitted discourse parsing system of the natural language group of Soochow University (SoNLP-DP) to the CoNLL 2015 shared task. Our System classifies discourse relations into explicit and non-explicit relations and uses a pipeline platform to conduct every subtask to form an end-toend shallow discourse parser in the Penn Discourse Treebank (PDTB). Our system is evaluated on the CoNLL-2015 Shared Task closed track and achieves the 18.51% in F1-measure on the official blind test set.