Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Text-based sentiment analysis is a growing research field in affective computing, driven by both commercial applications and academic interest. Continuous dimensional representations, such as valence-arousal (VA) space, can represent the affective state more precisely than discrete effective representations. In building dimensional sentiment applications, affective lexicons with valence-arousal ratings are useful resources but are still very rare. Therefore, recent studies have investigated the automatic development of VA lexicons using linear regression techniques. One of the major limitations of linear regression is the under-fitting problem which can cause a poor fit between the algorithm and the training data. To tackle this problem, this study proposes the use of a locally weighted linear regression (LWLR) model to predict the valence-arousal ratings of affective words. The locally weighted method performs a regression around the point of interest using only training data that are "local" to that point, and thus can reduce the impact of noise from unrelated training data. Experimental results show that the proposed method achieved better performance for VA word prediction.
This paper proposes neural networks for integrating compositional and non-compositional sentiment in the process of sentiment composition, a type of semantic composition that optimizes a sentiment objective. We enable individual composition operations in a recursive process to possess the capability of choosing and merging information from these two types of sources. We propose our models in neural network frameworks with structures, in which the merging parameters can be learned in a principled way to optimize a well-defined objective. We conduct experiments on the Stanford Sentiment Treebank and show that the proposed models achieve better results over the model that lacks this ability.
Songs heard between the ages of 15 and 24 should be remembered better and have a stronger relationship to autobiographical memories when compared with music from other phases of life (“reminiscence bump effect”). Additionally, the proportion of music-evoked autobiographical memories (MEAMs) is at a maximum in these years of early adolescence and then declines up to the age of 60. In our study we tried both to replicate these important findings based on a German sample and to further investigate the influence of the affective characteristics of the songs on the frequency of participants’ autobiographical memories. In Experiment 1 a group of adults ( N = 48, M age = 67.1 years) listened to excerpts from 80, number-one, popular music hits from 1930 to 2010 and gave written self-reports on MEAMs. In Experiment 2 the affective characteristics were rated by another group of adults ( N = 22, M age = 66 years) and were used to predict the frequency of MEAMs. As a main result of Experiment 1, we confirmed the reminiscence bump and decline effect with a small effect size for the ratings of feelings evoked by the song and with a medium effect size for the song recognition performance of those songs released during the participants’ age range of 15 to 24 years. The total number of MEAMs was only marginally influenced by a memory bump and decline effect, and participants showed a significant proportion of MEAMs up to the fifth decade. Experiment 2 revealed that the affective ratings of the songs were unequally distributed over the two-dimensional emotion space unlike the average rate of MEAMs which was nearly equally distributed. In contrast to previous research, we therefore conclude that popular songs can be associated with autobiographical memory over five decades of life – independent of the affective character of the music.
We propose and evaluate the use of an affective-semantic model to expand the affective lexica of German, Greek, English, Spanish and Portuguese. Motivated by the assumption that semantic similarity implies affective similarity, we use word level semantic similarity scores as semantic features to estimate their corresponding affective scores. Various context-based semantic similarity metrics are investigated using contextual features that include both words and character n-grams. The model produces continuous affective ratings in three dimensions (valence, arousal and dominance) for all five languages, achieving consistent performance. We achieve classification accuracy (valence polarity task) between 85% and 91% for all five languages. For morphologically rich languages the proposed use of character n-grams is shown to improve performance.
News websites give their users the opportunity to participate in discussions about published articles, by writing comments. Typically, these comments are unstructured making it hard to understand the flow of user discussions. Thus, there is a need for organizing comments to help users to (1) gain more insights about news topics, and (2) have an easy access to comments that trigger their interests. In this work, we address the above problem by organizing comments around the entities and the aspects they discuss. More specifically, we propose an approach for entity and aspect extraction from user comments through the following contributions. First, we extend traditional Named-Entity Recognition approaches, using coreference resolution and external knowledge bases, to detect more occurrences of entities in comments. Second, we exploit part-of-speech tag, dependency tag, and lexical databases to extract explicit and implicit aspects around discussed entities. Third, we evaluate our entity and aspect extraction approach, on manually annotated data, showing that it highly increases precision and recall compared to baseline approaches.
This paper presents a novel technique for empty category (EC) detection using distributed word representations. A joint model is learned from the labeled data to map both the distributed representations of the contexts of ECs and EC types to a low dimensional space. In the testing phase, the context of possible EC positions will be projected into the same space for empty category detection. Experiments on Chinese Treebank prove the effectiveness of the proposed method. We improve the precision by about 6 points on a subset of Chinese Treebank, which is a new state-ofthe-art performance on CTB.
Abstract This article measures the productivity index of the Old English suffixes -cund, -ful, and -isc as well as the prefix ful- and checks the results against the diachronic evolution of the affixes. The frameworks brought to the discussion include Type frequency measurement, as well as productivity indexes proposed by Baayen (1992, 1993, 2009) and Trips (2009). The sources are both textual (The Dictionary of Old English Corpus) and lexicographical (the lexical database of Old English Nerthus). The conclusion drawn is that Baayen's (1992, 1993, 2002) index of Global Productivity provides the most consistent results with the diachronic evolution of the affixes.
In this paper, we propose a baseline messagelevel sentiment classification method, as developed for SemEval-2015 Task 10, Subtask B. This system leverages both hand-crafted features and message-level embedding features, and uses an SVM classifier for messagelevel sentiment classification. In pre-training the embedding features, we use one million randomly-selected tweets. We present results over SemEval-2015 Task 10, Subtask B, as well as the Stanford Sentiment Treebank. Our experiments show the effectiveness of our method over both datasets.
Lexical semantic information plays an important role in supervised dependency parsing. In this paper, we add lexical semantic features to the feature set of a parser, obtaining improvements on the Penn Chinese Treebank. We extract semantic categories of words from HowNet, and use them as semantic information of words. Moreover, we investigate the method to compute semantic similarity between Chinese compound words, and obtain semantic information of words which did not record in HowNet. Our experiments show that unlabeled attachment scores can increase by 1.29%.
The emotional habituation plays an important role in individuals' adaptation to the environment. The present study explored the brain's emotional habituation to positive and negative pictures of diverse emotional intensities. Event-related potentials (ERPs) were recorded in two different experimental sessions, for highly positive (HP), mildly positive (MP) and neutral picture and for highly negative (HN), mildly negative (MN) and neutral picture. Subjects were asked to perform a standard/deviant categorization task, irrespective of emotionality of the deviants. The behavior results showed that the arousal ratings for HP stimuli decreased significantly with stimulus repetition. In addition, the ERP results displayed earlier N1 peak latencies with stimulus repetition in the positive session. Furthermore, the size of the emotion effect, which was computed by the emotion-neutral differences, decreased significantly for HP and MP stimuli with stimulus repetition in P3 amplitudes. Conversely, the current study failed to observe an emotional habituation effect to negative stimuli in any behavioral or ERP indexes. These results suggest that the humans' emotional reactions to positive stimuli, irrespective of the emotional intensity, are susceptible to habituation, irrespective of information processing stage. However, the humans' emotional reactions to negative stimuli are resistant to habituation, irrespective of the emotional intensities of the stimuli and the information processing stage. This valence-specific habituation effect is independent of the emotional intensity of the stimuli.
Following a comparison of the different views on lexical meaning conveyed by the Latin WordNet and by a treebank-based valency lexicon for Latin, the paper evaluates the degree of overlapping between a number of homogeneous lexical subsets extracted from the two resources.
Dependency parsing has become an important line of research in natural language processing in recent years. This is due to its usefulness in a wide variety of real world applications. This paper presents the improvement of Vietnamese dependency parsing using distributed word representations. Our parser achieves an accuracy of 76.29% of unlabelled attachment score or 69.25% of labelled attachment score. This is the most accurate dependency parser for the Vietnamese language in comparison to others which are trained and tested on the same dependency treebank. The distributed word representations are produced by two recent unsupervised learning models, the Skip-gram model and the GloVe model. We also show that distributed representations produced by the GloVe model are better than those produced by the Skip-gram model when being used in dependency parsing. Our dependency parsing system, including software, corpus and distributed word representations, is released as an open source project, freely available for research purpose.
Categorial grammars are attractive because they have a clear account of unbounded dependencies. This accounting is especially important in Mandarin Chinese which makes extensive usage of unbounded dependencies. However, parsers trained on existing categorial grammar annotations (Tse and Curran, 2010) extracted from the Penn Chinese Treebank This work reannotates the Penn Chinese Treebank into a generalized categorial grammar which uses a larger rule set and a substantially smaller category set while retaining the capacity to model unbounded dependencies. Experimental results show a statistically significant improvement in parsing accuracy with this categorial grammar.
In this work, we present a novel way of using neural network for graph-based dependency parsing, which fits the neural network into a simple probabilistic model and can be furthermore generalized to high-order parsing. Instead of the sparse features used in traditional methods, we utilize distributed dense feature representations for neural network, which give better feature representations. The proposed parsers are evaluated on English and Chinese Penn Treebanks. Compared to existing work, our parsers give competitive performance with much more efficient inference.
Limbic encephalitis (LE) is an autoimmune-mediated disorder that affects structures of the limbic system, in particular, the amygdala. The amygdala constitutes a brain area substantial for processing of emotional, especially fear-related signals. The amygdala is also involved in neuroendocrine and autonomic functions, including skin conductance responses (SCRs) to emotionally arousing stimuli. This study investigates behavioral and autonomic responses to discrete emotion evoking and neutral film clips in a patient suffering from LE associated with contactin-associated protein-2 (CASPR2) antibodies as compared to a healthy control group. Results show a lack of SCRs in the patient while watching the film clips, with significant differences compared to healthy controls in the case of fear-inducing videos. There was no comparable impairment in behavioral data (emotion report, valence, and arousal ratings). The results point to a defective modulation of sympathetic responses during emotional stimulation in patients with LE, probably due to impaired functioning of the amygdala.
We propose a model of Tibetan syntactic parsing which is based on Tibetan syl lables instead of Tibetan words, change Tibetan syntactic Treebank use algorith m of labeling.
In this paper we explore different statistical dependency parsers for parsing Telugu. We consider five popular dependency parsers namely, MaltParser, MSTParser, TurboParser, ZPar and Easy-First Parser. We experiment with different parser and feature settings and show the impact of different settings. We also provide a detailed analysis of the performance of all the parsers on major dependency labels. We report our results on test data of Telugu dependency treebank provided in the ICON 2010 tools contest on Indian languages dependency parsing. We obtain state-of-the art performance of 91.8% in unlabeled attachment score and 70.0% in labeled attachment score. To the best of our knowledge ours is the only work which explored all the five popular dependency parsers and compared the performance under different feature settings for Telugu.
In communication a great deal of meaning is exchanged through body language, including gaze, posture, hand gestures and body movements. Body language is largely culture-specific, and rests, for its comprehension, on people's sharing socio-cultural and linguistic norms. In cross-cultural communication, L2 speakers' use of body language may convey meaning that is not understood or misinterpreted by the interlocutors, affecting the pragmatics of communication. In spite of its importance for cross-cultural communication, body language is neglected in ESL/EFL teaching. This paper argues that the study of body language should be integrated in the syllabus of ESL/EFL teaching and learning. This is done by: 1) reviewing literature showing the tight connection between language, speech and gestures and the problems that might arise in cross-cultural communication when speakers use and interpret body language according to different conventions; 2) reporting the data from two pilot studies showing that L2 learners transfer L1 gestures to the L2 and that these are not understood by native L2 speakers; 3) reporting an experience teaching body language in an ESL/EFL classroom. The paper suggests that in multicultural ESL/EFL classes teaching body language should be aimed primarily at raising the students' awareness of the differences existing across cultures.
Word order differences between source and target languages pose a serious challenge to statistical machine translation (SMT). Pre-ordering, an approach that reorders source words into a target-word-like order as a preprocessing step, has been shown effective in handling word order between different languages and improving translation performance of SMT. In this paper, we propose a novel word reordering method based on the pre-ordering framework. Instead of using a supervised parser trained on a monolingual treebank, our method extracts bilingual structural information for reordering from automatically wordaligned sentence pairs into dependency-tree-like structures, then learns a reordering model by training a dependency parser on this extracted pseudo-treebank. Experiment results show that our pre-ordering method is effective in permuting source words to resemble word order of the target language, and improving translation quality.
The spinal tree adjoining grammar (TAG) parsing model of [Carreras 08] achieves the current state-of-the-art constituent parsing accuracy on the commonly used English Penn Treebank evaluation setting. Unfortunately, the model has the serious drawback of low parsing efficiency since its Eisner-CKY style parsing algorithm needs O(n4) computation time for input length n. This paper investigates a more practical solution and presents a beam search shift-reduce algorithm for spinal TAG parsing. Since the algorithm works in O(bn) (b is beam width), it can be expected to provide a significant improvement in parsing speed. However, to achieve faster parsing, it needs to prune a large number of candidates in an exponentially large search space and often suffers from severe search errors. In fact, our experiments show that the basic beam search shift-reduce parser does not work well for spinal TAGs. To alleviate this problem, we extend the proposed shift-reduce algorithm with two techniques: Dynamic Programming of [Huang 10a] and Supertagging. The proposed extended parsing algorithm is about 8 times faster than the Berkeley parser, which is well-known to be fast constituent parsing software, while offering state-of-the-art performance. Moreover, we conduct experiments on the Keyaki Treebank for Japanese to show that the good performance of our proposed parser is language-independent.
Previous studies indicate that emotion regulation may occur unconsciously, without the cost of cognitive efforts; and that conscious acceptance effectively reduces the emotional consequences of negative events. However, it has yet to be determined how conscious and unconscious acceptance strategies differ in behavioral and physiological consequences of emotion regulation. As unconscious regulation occurs with little cost of cognitive resources, the current study hypothesizes that unconscious acceptance regulates the emotional consequence of negative events more effectively compared to conscious acceptance. Subjects were randomly assigned to conscious acceptance, unconscious acceptance and control conditions. A frustrating arithmetic task was used to induce negative emotion. Emotional experiences were assessed by the positive affect and negative affect scale (PANAS) while emotion-related physiological activation was assessed by the heart-rate reactivity. The results showed that unconscious acceptance produced less reductions of positive affect ratings compared to conscious acceptance during frustration. In addition, both conscious and unconscious acceptance strategies significantly decreased emotion-related heart-rate activity (to a similar extent) in comparison with the control condition. Moreover, heart-rate reactivity showed a trend of positive correlation with negative affect rating and a trend of negative correlation with positive affect rating during frustration compared to baseline phases. Thus, unconscious acceptance is not only able to decrease emotion-related physiological activity, but also able to produce better emotional experiences compared to conscious acceptance. This suggests that it is practically important to consider unconscious acceptance for emotion regulation in real-life settings.
Paraphrase identification is a semantic text similarity task which is an important part of many natural language processing applications. Existing methods use vector space models, word co-occurrence information, lexical databases, parsers and machine translation (MT) evaluation metrics to find text similarity. However, other aspects such as negations, inverse relations and semantic roles of the sentences are also very much important in identifying paraphrases. Furthermore, the semantics of the sentences are hidden when the sentences are complex. We propose an approach to find similarity between pair of texts by considering all these factors. We have used an approach to determine set of clauses present in the texts by resolving conjunctions in complex sentences that identify hidden triples from the text. The approach extracts clause-based similarity features namely concept score, relation score, proposition score and word score from the texts. We have combined these similarity features along with MT metrics features to identify whether the texts are paraphrases or not using Support Vector Machine model. We have evaluated our methodology to measure the paraphrase similarity for Microsoft Research corpus. The statistical tests namely |$k$|-fold paired |$t$|-test and McNemar's test show that including clause-based features significantly improved the performance. Also, our approach outperforms state-of-the-art methods in terms of accuracy, |$F$|1-measure and |$f$|1-measure.
Redundancy is an important psycholinguistic concept which is often used for explanations of language change, but is notoriously difficult to operationalize and measure. Assuming that the reconstruction of a syntactic structure by a parser can be used as a rough model of the understanding of a sentence by a human hearer, I propose a method for estimating redundancy. The key idea is to compare performances of a parser on a given treebank before and after artificially removing all information about a certain grammeme from the morphological annotation. The change in performance can be used as an estimate for the redundancy of the grammeme. I perform an experiment, applying MaltParser to an Old Church Slavonic treebank to estimate grammeme redundancy in Proto-Slavic. The results show that those Old Church Slavonic grammemes within the case, number and tense categories that were estimated as most redundant are those that disappeared in modern Russian. Moreover, redundancy estimates serve as a good predictor of case grammeme frequencies in modern Russian. The small sizes of the samples do not allow to make definitive conclusions for number and tense.
Re-ranking models of parse trees have been focused on re-ordering parse trees with a syntactic view. However, also a semantic view should be considered in re-ranking parse trees, because the fact that a word pair has a dependency implies that the pair has both syntactic and semantic relations. This paper proposes a re-ranking model for dependency parsing based on a combination of syntactic and semantic plausibilities of dependencies. The syntactic probability is used as a syntactic plausibility of a parse tree, and a knowledge graph embedding is adopted to represent its semantic plausibility. The knowledge graph embedding allows the semantic plausibility of parse trees to be expressed effectively with ease. The experiments on the standard Penn Treebank corpus prove that the proposed model improves the base parser regardless of the number of candidate parse trees.
This paper addresses Chinese discourse segmentation based on punctuation mark. Particularly, we propose various kinds of lexical, syntactic, position and punctuation features to train classifiers for Chinese discourse segmentation. Experimental results on CDTB (Chinese Discourse Treebank) show that our method based on punctuation mark is appropriate for Chinese discourse segmentation with 89.2% in accuracy.
This communication is concerned with theoretical aspects of NLP, and with some ‘epistemological impediments ’ [1] to the development of Arabic generation and recognition programs or lexical databases. The discussion focuses on the methodological approach underlying the elaboration of specifications associated to the entries of an Arabic lexical database, in relation with the DIINAR.1 Arabic Language database and the DIINAR-MBC Euro-Mediterranean project1. Such specifications consist of morpho-semantic as well as syntactico-semantic features. Semantic aspects belong to the field of finite semantics. A definition of that field will be given in connection with the notion of ambiguity, which is revisited here in the context of formal linguistics and a cognitive approach of computational linguistics.
In this paper, the authors propose formalism for representing a knowledge base (KB) by network. The objective is to achieve a high coverage of this base. This type of network is similar to the semantic network with the difference that the arcs are quantified by a value indicating the semantic proximity between the concepts. This semantic proximity presents taxonomic relations, synonyms, and non-taxonomic relations (contextual relations). This latter are discovered based on the association rules model. This model is based on (i) indexing method (ii) the French lexical database EuroWordNet (EWNF) and (iii) the Apriori algorithm. The contextual relations are the latent relations buried in the KB, carried by the semantic context. Evaluating our representation formalism shows better result about 80% of coverage of the KB.
Arabic lexical language resources in a cross-lingual perspective are a major issue in the development of NLP applications, in the context of present-day web interactive developments. Much has already been done, and some existing achievements will be mentioned. The bulk of the challenge, though, still remains ahead. The main contribution of this paper is methodological. The first section will endeavour to show some of the specific needs for an elaborated lexical database in Arabic, in relation with the structure of the writing system (‘unvowelled ’ writing and the structure of word-forms). Grammar-lexis relations are an essential part of any Arabic lexical resource. These relations operate at word-level in morphological analysers, and at sentence-level in morphosyntactic ones. Sets of finite and exhaustive morphosyntactic specifiers have already been elaborated at both word- and sentence-level, but have only been implemented at word-level. The second section is concerned with the cross-lingual aspects of lexical LR-s including Arabic. The author underlines the necessity, in the present state of the Art, of considering ‘multilingual ’ or ‘cross-lingual ’ lexical databases including Arabic as sets of oriented bilingual LR-s, in which one of the languages under consideration is the ‘source language ’ and the other, the ‘target language’. It is essential for the future of Arabic linguistic engineering that sets of bilingual resources of this type, including the morphosyntactic specifiers outlined in section 1, be elaborated and implemented in the coming years, and that a substantial number of them take Arabic as a ‘source language’.
This study investigated sensory characteristics and cross-cultural consumer acceptability of sweet crispy chicken (Dakgangjeong) prepared with six types of Korean-style sauces among Korean and Chinese consumers. The main ingredient(s) of each sauce was soy sauce (SOY), Japanese apricot extract and soy sauce (JASOY), gochujang (SPICY), minced garlic (GARL), and ketchup (KET-I and KET-II); KET-I and KET-II were modified to possess ethnic Korean flavors. In general, Korean and Chinese consumers preferred all types of Dakgangjeong, except for GARL and SPICY, respectively. Least preferred products of each country had the lowest familiarity rating among consumers of the respective countries. Similar to previous studies, these results showed that familiarity is an important factor affecting consumer preference in a cross-cultural context. Particularly, it was found that higher familiarity of the product was not found to influence consumer to like a product, but rather low familiarity seemed to affect consumers to reject a product.
Based on the annotation of discourse connective in Chinese Discourse Treebank, especially the annotation of the connective and its relation classification. The authors extract syntax, lexical and position features of automatic syntax tree and standard syntax tree, and use supervised method to recognize and classify connective. Experimental results show that connective recognition F1-measure is 69.2%, and connective classification accuracy is 89.1%.
We present an analysis of a treebank of spontaneous English dyadic conversations, investigating whether the degree of syntactic priming found across speakers is a function of the degrees of affective alignment and overall positivity of the speakers. We use information theory to measure the proportion of overlap between the syntactic structures of the speakers. The affective state of the speakers is indexed by aggregated measures of the affective valences of the words they use. We find that there is a positive relation between syntactic priming and affective alignment, over and above any lexical repetition effects. This constitutes evidence for the percolation of inter-speaker alignment across multiple levels of representation. This also illustrates the indexical value of syntactic alignment, as has been proposed in modern functional theories of grammar such as Dialogic Syntax.
We describe automatic conversion of the SynTagRus dependency treebank\nof Russian to the PROIEL format (with the ultimate purpose of obtaining a single-format\ndiachronic treebank spanning more than a thousand years), focusing\non analysis of shared arguments in verbal coordinations. Whether arguments\nare shared or private is not marked in the SynTagRus native format,\nbut the PROIEL format indicates sharing by means of secondary dependencies.\nIn order to recover missing information and insert secondary dependencies\ninto the converted SynTagRus, we create a simple guessing algorithm\nbased on four probabilistic features: how likely a given argument type\nis to be shared; how likely an argument in a given position is to be shared;\nhow likely a given verb is to have a given argument; how likely a given verb\nis to have a given argument frame. Boosted with a few deterministic rules and\ntrained on a small manually annotated sample (346 sentences), the guesser\nvery successfully inserts shared subjects (F-score 0.97), which results\nin excellent overall performance (F-score 0.92). Non-subject arguments are\nshared much more rarely, and for them the results are poorer (0.31 for objects;\n0.22 for obliques). We show, however, that there are strong reasons\nto believe that performance can be increased if a larger training sample\nis used and the guesser gets to see enough positive examples. Apart from\ndescribing a useful practical solution, the paper also provides quantitative\ndata about and offers non-trivial insights into Russian verbal coordination.
This paper describes our system designed for the NLPCC 2015 shared task on Chinese word segmentation (WS) and POS tagging for Weibo Text. We treat WS and POS tagging as two separate tasks and use a cascaded approach. Our major focus is how to effectively exploit multiple heterogeneous data to boost performance of statistical models. This work considers three sets of heterogeneous data, i.e., Weibo ($$\textit{WB}$$, 10K sentences), Penn Chinese Treebank 7.0 ($$\textit{CTB7}$$, 50K), and People’s Daily ($$\textit{PD}$$, 280K). For WS, we adopt the recently proposed coupled sequence labeling to combine $$\textit{WB}$$, $$\textit{CTB7}$$, and $$\textit{PD}$$, boosting F1 score from $$93.76\%$$ (baseline model trained on only $$\textit{WB}$$) to $$95.58\%$$ ($$+1.82\%$$). For POS tagging, we adopt an ensemble approach combining coupled sequence labeling and the guide-feature based method, since the three datasets have three different annotation standards. First, we convert $$\textit{PD}$$ into the annotation style of $$\textit{CTB7}$$ based on coupled sequence labeling, denoted by $$\textit{PD}^{\textit{CTB}}$$. Then, we merge CTB7 and $$\textit{PD}^{\textit{CTB}}$$ to train a POS tagger, denoted by $$\textit{Tag}_{\textit{CTB7}+\textit{PD}^{\textit{CTB}}}$$, which is further used to produce guide features on $$\textit{WB}$$. Finally, the tagging F1 score is improved from 87.93% to 88.99% (+1.06%).
Recent work on joint word segmentation, POS (Part Of Speech) tagging, and dependency parsing in Chinese has two key problems: the first is that word segmentation based on character and dependency parsing based on word were not combined well in the transition-based framework, and the second is that the joint model suffers from the insufficiency of annotated corpus. In order to resolve the first problem, we propose to transform the traditional word-based dependency tree into character-based dependency tree by using the internal structure of words and then propose a novel character-level joint model for the three tasks. In order to resolve the second problem, we propose a novel semi-supervised joint model for exploiting n-gram feature and dependency subtree feature from partially-annotated corpus. Experimental results on the Chinese Treebank show that our joint model achieved 98.31%, 94.84% and 81.71% for Chinese word segmentation, POS tagging, and dependency parsing, respectively. Our model outperforms the pipeline model of the three tasks by 0.92%, 1.77% and 3.95%, respectively. Particularly, the F1 value of word segmentation and POS tagging achieved the best result compared with those reported until now.
The usage of phrasemes evidences not only their variability, transformations and modifications, but also the most frequent forms of their realization (phraseme-types) and frequency (phraseme-tokens), i.e. phrasemes’ flexibility. In this paper, selected Lithuanian idiomatic predicate phrasemes are analysed in the Corpus of Contemporary Lithuanian Language, in the Phraseological Dictionary and in the lexical database of the Dictionary of Lithuanian Phrases. The results of comparison show that the corpus research can give rich evidence about the morphological flexibility of phrasemes. This information can help to improve representation of phrasemes in the phraseological dictionaries of Lithuanian, in order to make them more usage-based and more usage-oriented.
The goal of this work is to bring semantics into the tasks of text recognition and retrieval in natural images. Although text recognition and retrieval have received a lot of attention in recent years, previous works have focused on recognizing or retrieving exactly the same word used as a query, without taking the semantics into consideration. In this paper, we ask the following question: \emph{can we predict semantic concepts directly from a word image, without explicitly trying to transcribe the word image or its characters at any point?} For this goal we propose a convolutional neural network (CNN) with a weighted ranking loss objective that ensures that the concepts relevant to the query image are ranked ahead of those that are not relevant. This can also be interpreted as learning a Euclidean space where word images and concepts are jointly embedded. This model is learned in an end-to-end manner, from image pixels to semantic concepts, using a dataset of synthetically generated word images and concepts mined from a lexical database (WordNet). Our results show that, despite the complexity of the task, word images and concepts can indeed be associated with a high degree of accuracy
This thesis details three contributions to the advancement of semantic-enriched parsing for English sentences: inventories of semantic relations covering three semantically ambiguous linguistic phenomena, large datasets annotated according to the inventories, and, finally, a suite of tools for semantically-enriched parsing built using the datasets. For the purposes of this thesis, semantically-enriched parsing is defined as the reconstruction of the underlying grammatical structure of text along with shallow semantic annotation of semantically-ambiguous structures. Ultimately, semantically-enriched parsing is one of the most critical steps in natural language understanding---the initial step in which the text is read by the machine into a knowledge representation for further processing and reasoning. ❧ The first contribution of this thesis is to advance the theoretical foundations for the interpretation of three ambiguous linguistic phenomena in English that have significant overlap in terms of the relations expressed: noun compounds, possessive constructions, and prepositions. For these, I define inventories of relations based upon extensive annotation by myself, previous work by others, and inter-annotator agreement studies. In the case of prepositions, the relations are created by refining an existing resource whereas the other two are created from scratch. In addition to mappings to prior work, mappings are provided across the different inventories in order to create a unified set of relations. ❧ Second, I produce large datasets annotated according to the aforementioned sense inventories. Such data is vital for training most automatic tools and also provides exemplars for the theory embodied in the inventories. Some of these datasets are created from scratch, including a collection of over 17,500 noun compounds and a collection of over 21,900 possessive construction examples. In the case of prepositions, an existing resource including over 24,000 annotated examples is refined. ❧ The final contribution is a suite of tools that can construct semantically-enriched parse trees. The suite is designed to work in a sequential, pipeline-like fashion and can be thought of as consisting of two subsections. The first part reconstructs the grammatical structure of the text using a dependency parser that extends the non-directional easy-first algorithm developed by Goldberg and Elhadad in order to support non-projective trees and is trained using my improved dependency tree conversion of the Penn Treebank. Second are semantic annotation modules that add shallow semantic annotation for noun compounds, preposition senses, possessives, and verbal arguments. Combined, these tools produce semantically-enriched parse trees that include both grammatical structure and shallow semantics. The core parser itself achieves state-of-the-art accuracy and can process over \parsespeed sentences per second, which is substantially faster than most of the accurate parsers available today. ❧ In conclusion, this thesis work provides significant contributions to computational linguistics, both in terms of theory and resources. It advances our understanding of the relations expressed by three semantically-ambiguous linguistic phenomena, creates large annotated datasets useful for machine learning, and produces a fast, accurate, and informative system for semantically-enriched parsing.
This article explains why XML format has become established as the standard format for multilevel hierarchical structuring of linguistic databases and how an XML Schema can be used to manage the formal structure and content of elements in a dictionary database. Various aspects that must be taken into account when structuring complex dictionary databases in XML format are presented: the lexicographic or content aspect, the practical aspect, and the technical aspect. Decision-making is illustrated with the example of designing an XML Schema for the Dictionary of Slovenian Synonyms.
In this article, we present an incremental dependency parsing algorithm with an arc-eager variant of the left-corner parsing strategy. Our algorithm’s stack depth captures the center-embeddedness of the recognized dependency structure. A higher stack depth occurs only when processing deeper center-embedded sentences in which people find difficulty in comprehension. We examine whether our algorithm can capture the syntactic regularity that universally exists in languages through two kinds of experiments across treebanks of 19 languages. We first show through oracle parsing experiments that our parsing algorithm consistently requires less stack depth to recognize annotated trees relative to other algorithms across languages. This result also suggests the existence of a syntactic universal by which deeper center-embedding is a rare construction across languages, a result that has yet to be quantitatively cross-linguistically examined. We further investigate the above claim through supervised parsing experiments and show that our proposed parser is consistently less sensitive to constraints on stack depth bounds when decoding across languages, while the performance of other parsers such as the arc-eager parser is largely affected by such constraints. We thus conclude that the stack depth of our parser represents a more meaningful measure for capturing syntactic regularity in languages than those of existing parsers.
The article contributes to the discussion on the interpretation of processes taking place in the contemporary Polish language from two perspectives: diachronic and synchronic (ahistorical, to be precise). The author draws attention to the problem of not accounting for the achievements of diachronic linguistics, which is visible in the studies devoted to the analysis of the language used by the contemporary Poles. This leads to a situation in which the same research questions are not infrequently examined and interpreted in completely different ways in synchronic studies and those dealing with the history of language. Sometimes the differences are so vast that coherent scientific discourse becomes disabled. Enriching the investigation of the processes occurring in the contemporary Polish language with conclusions drawn from analysing historical sources is thus an important claim. Considering the diachronic perspective could be particularly beneficial for discussing the issue of the norm. Since a number of problems with language correctness result from tendencies and processes which have been happening in language for centuries, only becoming acutely aware of them can bring a change in understanding codification and the linguistic norm.
Large-scale data resources needed for progress toward natural language understanding are not yet widely available and typically require considerable expense and expertise to create. This paper addresses the problem of developing scalable approaches to annotating semantic frames and explores the viability of crowdsourcing for the task of frame disambiguation. We present a novel supervised crowdsourcing paradigm that incorporates insights from human computation research designed to accommodate the relative complexity of the task, such as exemplars and real-time feedback. We show that non-experts can be trained to perform accurate frame disambiguation, and can even identify errors in gold data used as the training exemplars. Results demonstrate the efficacy of this paradigm for semantic annotation requiring an intermediate level of expertise. 1 The semantic bottleneck Behind every great success in speech and language lies a great corpus—or at least a very large one. Advances in speech recognition, machine translation and syntactic parsing can be traced to the availability of large-scale annotated resources (Wall Street Journal, Europarl and Penn Treebank, respectively) providing crucial supervised input to statistically learned models. Semantically annotated resources have been comparatively harder to come by: representing meaning poses myriad philosophical, theoretical and practical challenges, particularly for general purpose resources that can be applied to diverse domains. If these challenges can be addressed, however, semantic resources hold significant potential for fueling progress beyond shallow syntax and toward deeper language understanding. This paper explores the feasibility of developing scalable methodologies for semantic annotation, inspired by three strands of work. First, frame semantics, and its instantiation in the Berkeley FrameNet project (Fillmore and Baker, 2010), offers a principled approach to representing meaning. FrameNet is a lexicographic resource that captures syntactic and semantic generalizations that go beyond surface form and part of speech, famously including the relationships among words like buy, sell, purchase and price. These rich structural relations provide an attractive foundation for work in deeper natural language understanding and inference, as attested by the breadth of applications at the Workshop in Honor of Chuck Fillmore at ACL 2014 (Petruck and de Melo, 2014). But FrameNet was not designed to support scalable language technologies; indeed, it is perhaps a paradigm example of a hand-curated knowledge resource, one that has required significant expertise, training, time and expense to create and that remains under development. Second, the task of automatic semantic role labeling (ASRL) (Gildea and Jurafsky, 2002) serves as an applied counterpart to the ideas of frame semantics. Recent progress has demonstrated the viability of training automated models using frameannotated data (Das et al., 2013; Das et al., 2010; Johansson and Nugues, 2006). Results based on FrameNet data have been limited by its incomplete
Quantifying image quality through subjective evaluation is very critical to image quality evaluation. Using the image quality ruler method, an average score per stimulus can be easily obtained in the unit of Just Noticeable Differences (JNDs). However, it requires a large number of subjects, since pure averaging does not consider the different judging quality of different subjects. In this paper, we propose an image quality evaluation framework using the image quality ruler method with a statistical model. By incorporating this model, we consider the quality score, the expertise of the subjects, and the difficulty of image rating task as three hidden variables. Then we use expectation-maximization (EM) to estimate these hidden variables. From our experimental results, we show that our method provides reliable results without using a large number of subjects. Preliminary results also demonstrate that the estimates of the parameters can guide us to better distribute the valuable human resources used to conduct psychophysical experiments.
This study examines what effect mindfulness has on anxiety and memory levels in comparison to a suppression and control group. Participants underwent natural, suppression, and mindful conditions while being shown a series of positive and negative images and rating how happy/unhappy (valence) they felt and how excited/calm (arousal). After answering a series of questionnaires, participants were tested on their recall levels. The results revealed that arousal and valence ratings of the pictures were not significantly different across conditions, and neither was the recall of the participants. The results of this experiment do not align with previous research and may be due to a few limitations within the study. Therefore, more studies will need to be conducted and further research will need to be completed.
BACKGROUND: Bipolar Spectrum Disorder (BPSD) is associated with changes in self-related processing and affect, yet the relationship between self-image and affect in the BPSD phenotype is unclear. METHODS: 47 young adults were assessed for hypomanic experiences (BPSD phenotype) using the Mood Disorders Questionnaire. Current and future self-images (e.g. I am… I will be…) were generated and rated for emotional valence, stability, and (for future self-images only) certainty. The relationship between self-image ratings and measures of affect (depression, anxiety and mania) were analysed in relation to the BPSD phenotype. RESULTS: The presence of the BPSD phenotype significantly moderated the relationship between (1) affect and stability ratings for negative self-images, and (2) affect and certainty ratings for positive future self-images. Higher positivity ratings for current self-images were associated with lower depression and anxiety scores. LIMITATIONS: This was a non-clinical group of young adults sampled for hypomanic experiences, which limits the extension of the work to clinical levels of psychopathology. This study cannot address the causal relationships between affect, self-images, and BPSD. Future work should use clinical samples and experimental mood manipulation designs. CONCLUSIONS: BPSD phenotype can shape the relationship between affect and current and future self-images. This finding will guide future clinical research to elucidate BPSD vulnerability mechanisms and, consequently, the development of early interventions.