Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We unify recent neural approaches to one-shot learning with older ideas of associative memory in a model for metalearning. Our model learns jointly to represent data and to bind class labels to representations in a single shot. It builds representations via slow weights, learned across tasks through SGD, while fast weights constructed by a Hebbian learning rule implement one-shot binding for each new task. On the Omniglot, Mini-ImageNet, and Penn Treebank one-shot learning benchmarks, our model achieves state-of-the-art results.
We propose Efficient Neural Architecture Search (ENAS), a faster and less expensive approach to automated model design than previous methods. In ENAS, a controller learns to discover neural network architectures by searching for an optimal path within a larger model. The controller is trained with policy gradient to select a path that maximizes the expected reward on the validation set. Meanwhile the model corresponding to the selected path is trained to minimize the cross entropy loss. On the Penn Treebank dataset, ENAS can discover a novel architecture thats achieves a test perplexity of 57.8, which is state-of-the-art among automatic model design methods on Penn Treebank. On the CIFAR-10 dataset, ENAS can design novel architectures that achieve a test error of 2.89%, close to the 2.65% achieved by standard NAS (Zoph et al., 2017). Most importantly, our experiments show that ENAS is more than 10x faster and 100x less resource-demanding than NAS.
Cognitive variation due to language and culture has been shown in a range of domains, including visual perception,emotions, theory of mind, economic strategies, decision making, and categorization. While such patterns are robust,individuals within a given culture are affected by these cultural patterns differentially. One possible cause for theseindividual differences is personality (e.g. extroversion or agreeableness). The personality traits of individuals will affecthow they interact with and adopt cultural patterns. To explore this possibility, we perform analyses on online data fromindividuals with self-identified Myers-Briggs personality types (a popularized personality measure that is widely self-reported in social media). In particular, we examine how personality type predicts the rate at which individuals adopt novellexical items and conform to the linguistic norms of their surrounding community. The results make explicit predictionsabout which individuals will be more affect by cultural and linguistic patterns.
Film clips are proven to be one of the most efficient techniques in emotional induction. However, there is scant literature on the effect of this procedure in older adults and, specifically, the effect of using different positive stimuli. Thus, the aim of the present study was to examine emotional differences between young and older adults and to know how a set of film clips works as mood induction procedure in older adults, especially, when trying to elicit attachment-related emotions. To this end, we use this procedure to analyze differences in subjective emotional response between young and older adults. A sample of 57 older adults and 83 young adults watched a film set previously validated in young population. Their responses were studied in an individual laboratory session to elicit 6 target emotions (disgust, fear, sadness, anger, amusement and tenderness) and neutral state. Self-reported emotional experience was measured using the Self-Assessment Manikin (SAM). Our results show that film clips are capable of evoking positive and negative emotions in older adults. Furthermore, older adults experienced more intensely negative emotions than young adults, especially in response to disgust and fear clips. They also reported higher arousal than young adults, especially in the case of sadness, anger and tenderness clips. Nevertheless, the older adults recovered more easily from the effects of the emotion induction. The young adults reported higher arousal ratings than older adults in response to amusement film clips. On the other hand, this study reflects the importance of controlling the baseline state to study the real strength of mood induction. Overall, current data suggests significant differences occur in emotional response in adult age and that film clips are an effective tool for studying positive and negative emotions in aging research.
<h3>Introduction</h3><br> DEFT Spanish Treebank was developed by the Linguistic Data Consortium (LDC) and the <a href="http://clic.ub.edu/">Language and Computation Center (CLiC), University of Barcelona</a>. It contains treebank annotation of international Spanish newswire text and Latin American Spanish discussion forum data created for the DARPA Deep Exploration and Filtering of Text (DEFT) program. <br> DEFT aimed to improve state-of-the-art capabilities in automated deep natural language processing with a particular focus on technologies dealing with inference, casual relationships and anomaly detection across several languages. DEFT Spanish Treebank supported the program's goal of deep natural language understanding. <br> <h3>Data</h3><br> Newswire source files were selected from Spanish Gigaword Third Edition (<a href="../../../ldc2011t12">LDC2011T12</a>) and were manually sentence-segmented for DEFT. Discussion forum source files were selected from Spanish discussion forum source data collected by LDC, consisting of continuous multi-posts of 100-1000 words. <br> This release contains 114 files (54,394 tokens) of newswire data and 60 files (55,307 tokens) of discussion forum data all of which were annotated with constituents and syntactic functions. The annotation guidelines for DEFT Spanish Treebank are included in the documentation accompanying this release. <br> Source documents are presented as plain text files with one sentence unit per line. Treebank annotation files are in xml. <br> <h3>Samples</h3><br> Please view this <a href="desc/addenda/LDC2018T01.txt">source sample</a> and <a href="desc/addenda/LDC2018T01.xml">treebank sample</a>. <br> <h3>Updates</h3><br> None at this time. <br> <h3>Acknowledgement</h3><br> This material is based on research sponsored by Air Force Research Laboratory and Defense Advance Research Projects Agency under agreement number FA8750-13-2-0045. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of Air Force Research Laboratory and Defense Advanced Research Projects Agency or the U.S. Government. </br> Portions © 1994-2001, 2004-2009 The Associated Press, © 2002, 2005, 2007, 2009-2010 Xinhua News Agency, © 2006, 2009, 2011, 2018 Trustees of the University of Pennsylvania
Artificially created treebank of elliptical constructions (gapping), in the annotation style of Universal Dependencies. Data taken from UD 2.1 release, and from large web corpora parsed by two parsers. Input data are filtered, sentences are identified where gapping could be applied, then those sentences are transformed, one or more words are omitted, resulting in a sentence with gapping. Details in Droganova et al.: Parse Me if You Can: Artificial Treebanks for Parsing Experiments on Elliptical Constructions, LREC 2018, Miyazaki, Japan.
We demonstrate that replacing an LSTM encoder with a self-attentive architecture can lead to improvements to a state-of-the-art discriminative constituency parser. The use of attention makes explicit the manner in which information is propagated between different locations in the sentence, which we use to both analyze our model and propose potential improvements. For example, we find that separating positional and content information in the encoder can lead to improved parsing accuracy. Additionally, we evaluate different approaches for lexical representation. Our parser achieves new state-of-the-art results for single models trained on the Penn Treebank: 93.55 F1 without the use of any external data, and 95.13 F1 when using pre-trained word representations. Our parser also outperforms the previous best-published accuracy figures on 8 of the 9 languages in the SPMRL dataset.
BACKGROUND: Life events (LEs) are associated with future physical and mental health. They are crucial for understanding the pathways to mental disorders as well as the interactions with biological parameters. However, deeper insight is needed into the complex interplay between the type of LE, its subjective evaluation and accompanying factors such as social support. The "Stralsund Life Event List" (SEL) was developed to facilitate this research. METHODS: The SEL is a standardized interview that assesses the time of occurrence and frequency of 81 LEs, their subjective emotional valence, the perceived social support during the LE experience and the impact of past LEs on present life. Data from 2265 subjects from the general population-based cohort study "Study of Health in Pomerania" (SHIP) were analysed. Based on the mean emotional valence ratings of the whole sample, LEs were categorized as "positive" or "negative". For verification, the SEL was related to lifetime major depressive disorder (MDD; Munich Composite International Diagnostic Interview), childhood trauma (Childhood Trauma Questionnaire), resilience (Resilience Scale) and subjective health (SF-12 Health Survey). RESULTS: The report of lifetime MDD was associated with more negative emotional valence ratings of negative LEs (OR = 2.96, p < 0.0001). Negative LEs (b = 0.071, p < 0.0001, β = 0.25) and more negative emotional valence ratings of positive LEs (b = 3.74, p < 0.0001, β = 0.11) were positively associated with childhood trauma. In contrast, more positive emotional valence ratings of positive LEs were associated with higher resilience (b = - 7.05, p < 0.0001, β = 0.13), and a lower present impact of past negative LEs was associated with better subjective health (b = 2.79, p = 0.001, β = 0.05). The internal consistency of the generated scores varied considerably, but the mean value was acceptable (averaged Cronbach's alpha > 0.75). CONCLUSIONS: The SEL is a valid instrument that enables the analysis of the number and frequency of LEs, their emotional valence, perceived social support and current impact on life on a global score and on an individual item level. Thus, we can recommend its use in research settings that require the assessment and analysis of the relationship between the occurrence and subjective evaluation of LEs as well as the complex balance between distressing and stabilizing life experiences.
People frequently engage in conversation about shared autobiographical events from their lives, particularly those with emotional significance. The pervasiveness of this practice raises the question whether shared memory reconstruction has the power to influence the memory and emotions associated with such events. We developed a novel paradigm that combined the strengths of the methods from autobiographical and collaborative memory research traditions to examine such consequences. We selected a shared, real-life autobiographical event of an exam, and asked students to recall their memory of taking a recent exam where they provided a group and/or personal narrative of this autobiographical event. Students first recalled the event either collaboratively (C) or individually (I), followed by a final individual (I) recall by all. Valence ratings as well as the emotional tone of the narratives converged to show that prior collaborative remembering down-regulated negative emotion and enhanced the positive emotional tone of the memories. The recalled detail in the narratives indicated that at initial recall members of collaborative groups reported fewer internal details than those who recalled alone, and reported more external details in a later recall when working alone. Earlier collaboration also increased collective memory such that more of these details were shared among prior group members in their later individual recall compared with those who did not collaborate before. We discuss the influence of collaborative remembering on shaping memory and emotion for autobiographical events as well as the potential mechanisms that promote collective autobiographical memory. (PsycINFO Database Record (c) 2018 APA, all rights reserved).
Our study aims to explore how much information about areal patterns of colexification we can gain from lexical databases such as CLICS and ASJP. We adopt a bottom-up (rather than hypothesis-driven) approach, identifying areal patterns in three steps: (i) determine spatial autocorrelations in the data, (ii) identify clusters as candidates for convergence areas and (iii) test the clusters resulting from the second step controlling for genealogical relatedness. Moreover, we identify a (genealogical) diversity index for each cluster. This approach yields promising results, which we regard as a proof of concept, but we also point out some drawbacks of the use of major lexical databases.
We know very little about how neural language models (LM) use prior linguistic context. In this paper, we investigate the role of context in an LSTM LM, through ablation studies. Specifically, we analyze the increase in perplexity when prior context words are shuffled, replaced, or dropped. On two standard datasets, Penn Treebank and WikiText-2, we find that the model is capable of using about 200 tokens of context on average, but sharply distinguishes nearby context (recent 50 tokens) from the distant history. The model is highly sensitive to the order of words within the most recent sentence, but ignores word order in the long-range context (beyond 50 tokens), suggesting the distant past is modeled only as a rough semantic field or topic. We further find that the neural caching model (Grave et al., 2017b) especially helps the LSTM to copy words from within this distant context. Overall, our analysis not only provides a better understanding of how neural LMs use their context, but also sheds light on recent success from cache-based models.
Modern solutions for implicit discourse relation recognition largely build universal models to classify all of the different types of discourse relations. In contrast to such learning models, we build our model from first principles, analyzing the linguistic properties of the individual top-level Penn Discourse Treebank (PDTB) styled implicit discourse relations: Comparison, Contingency and Expansion. We find semantic characteristics of each relation type and two cohesion devices---topic continuity and attribution---work together to contribute such linguistic properties. We encode those properties as complex features and feed them into a NaiveBayes classifier, bettering baselines(including deep neural network ones) to achieve a new state-of-the-art performance level. Over a strong, feature-based baseline, our system outperforms one-versus-other binary classification by 4.83% for Comparison relation, 3.94% for Contingency and 2.22% for four-way classification.
The transfer or share of knowledge between languages is a popular solution to resource scarcity in NLP. However, the effectiveness of cross-lingual transfer can be challenged by variation in syntactic structures. Frameworks such as Universal Dependencies (UD) are designed to be cross-lingually consistent, but even in carefully designed resources trees representing equivalent sentences may not always overlap. In this paper, we measure cross-lingual syntactic variation, or anisomorphism, in the UD treebank collection, considering both morphological and structural properties. We show that reducing the level of anisomorphism yields consistent gains in cross-lingual transfer tasks. We introduce a source language selection procedure that facilitates effective cross-lingual parser transfer, and propose a typologically driven method for syntactic tree processing which reduces anisomorphism. Our results show the effectiveness of this method for both machine translation and cross-lingual sentence similarity, demonstrating the importance of syntactic structure compatibility for boosting cross-lingual transfer in NLP.
long-span dependencies between discourse units is crucial to improve discourse parsing performance. Most existing approaches design sophisticated features or exploit various off-the-shelf tools, but achieve little success. In this paper, we propose a new transition-based discourse parser that makes use of memory networks to take discourse cohesion into account. The automatically captured discourse cohesion benefits discourse parsing, especially for long span scenarios. Experiments on the RST discourse treebank show that our method outperforms traditional featured based methods, and the memory based discourse cohesion can improve the overall parsing performance significantly 1.
We compare and analyze sequential, random access, and stack memory architectures for recurrent neural network language models. Our experiments on the Penn Treebank and Wikitext-2 datasets show that stack-based memory architectures consistently achieve the best performance in terms of held out perplexity. We also propose a generalization to existing continuous stack models (Joulin & Mikolov,2015; Grefenstette et al., 2015) to allow a variable number of pop operations more naturally that further improves performance. We further evaluate these language models in terms of their ability to capture non-local syntactic dependencies on a subject-verb agreement dataset (Linzen et al., 2016) and establish new state of the art results using memory augmented language models. Our results demonstrate the value of stack-structured memory for explaining the distribution of words in natural language, in line with linguistic theories claiming a context-free backbone for natural language.
Świgra is a parser of Polish generating constituency trees using a DCG style grammar stemming from Marek Świdzinski’s grammar “Gramatyka formalna jezyka polskiego” (1992). The grammar was heavily rewritten for the purpose of annotating the Skladnica treebank. The structure of trees was simplified with respect to Świdzinski’s version, many new types of constructions were included (in particular various forms of coordinated structures), a statistical disambiguating component was added. Moreover, the Clarin version of Świgra uses the valency dictionary Walenty developed within Clarin.
International audience
Previous studies have demonstrated differential perception of body expressions between males and females. However, only two recent studies (Kret et al., 2011; Krüger et al., 2013) explored the interaction effect between observer gender and subject gender, and it remains unclear whether this interaction between the two gender factors is gender-congruent (i.e., better recognition of emotions expressed by subjects of the same gender) or gender-incongruent (i.e., better recognition of emotions expressed by subjects of the opposite gender). Here, we used event-related potentials (ERPs) to investigate the recognition of fearful and angry body expressions posed by males and females. Male and female observers also completed an affective rating task (including valence, intensity, and arousal ratings). Behavioral results showed that male observers reported higher arousal rating scores for angry body expressions posed by females than males. ERP data showed that when recognizing angry body expressions, female observers had larger P1 for male than female bodies, while male observers had larger P3 for female than male bodies. These results indicate gender-incongruent effects in early and later stages of body expression processing, which fits well with the evolutionary theory that females mainly play a role in care of offspring while males mainly play a role in family guarding and protection. Furthermore, it is found that in both angry and fearful conditions male observers exhibited a larger N170 for male than female bodies, and female observers showed a larger N170 for female than male bodies. This gender-incongruent effect in the structural encoding stage of processing may be due to the familiarity of the body configural features of the same gender. The current results provide insights into the significant role of gender in body expression processing, helping us understand the issue of gender vulnerability associated with psychiatric disorders characterized by deficits of body language reading.
Implicit discourse relation recognition is a challenging task as the relation prediction without explicit connectives in discourse parsing needs understanding of text spans and cannot be easily derived from surface features from the input sentence pairs. Thus, properly representing the text is very crucial to this task. In this paper, we propose a model augmented with different grained text representations, including character, subword, word, sentence, and sentence pair levels. The proposed deeper model is evaluated on the benchmark treebank and achieves state-of-the-art accuracy with greater than 48% in 11-way and $F_1$ score greater than 50% in 4-way classifications for the first time according to our best knowledge.
Recent work on the problem of latent tree learning has made it possible to train neural networks that learn to both parse a sentence and use the resulting parse to interpret the sentence, all without exposure to ground-truth parse trees at training time. Surprisingly, these models often perform better at sentence understanding tasks than models that use parse trees from conventional parsers. This paper aims to investigate what these latent tree learning models learn. We replicate two such models in a shared codebase and find that (i) only one of these models outperforms conventional tree-structured models on sentence classification, (ii) its parsing strategies are not especially consistent across random restarts, (iii) the parses it produces tend to be shallower than standard Penn Treebank (PTB) parses, and (iv) they do not resemble those of PTB or any other semantic or syntactic formalism that the authors are aware of.
Released only a year ago as the outputs of a research project (``Parsing Web 2.0 Sentences'', supported in part by a TÜBİTAK 1001 grant (No. 112E276) and a part of the ICT COST Action PARSEME (IC1207)), IMST and IWT are currently the most comprehensive Turkish dependency treebanks in the literature. This article introduces the final states of our treebanks, as well as a newly integrated hierarchical categorization of the multiheaded dependencies and their organization in an exclusive deep dependency layer in the treebanks. It also presents the adaptation of recent studies on standardizing multiword expression and named entity annotation schemes for the Turkish language and integration of benchmark annotations into the dependency layers of our treebanks and the mapping of the treebanks to the latest Universal Dependencies (v2.0) standard, ensuring further compliance with rising universal annotation trends. In addition to significantly boosting the universal recognition of Turkish treebanks, our recent efforts have shown an improvement in their syntactic parsing performance (up to 77.8{\%}/82.8{\%} LAS and 84.0{\%}/87.9{\%} UAS for IMST/IWT, respectively). The final states of the treebanks are expected to be more suited to different natural language processing tasks, such as named entity recognition, multiword expression detection, transfer-based machine translation, semantic parsing, and semantic role labeling.
This paper presents results from the first statistical dependency parser for Turkish. Turkish is a free-constituent order language with complex agglutinative inflectional and derivational morphology and presents interesting challenges for statistical parsing, as in general, dependency relations are between “portions” of words – called inflectional groups. We have explored statistical models that use different representational units for parsing. We have used the Turkish Dependency Treebank to train and test our parser but have limited this initial exploration to that subset of the treebank sentences with only left-to-right non-crossing dependency links. Our results indicate that the best accuracy in terms of the dependency relations between inflectional groups is obtained when we use inflectional groups as units in parsing, and when contexts around the dependent are employed.
In literature, writers have the liberty to deviate from linguistic norms under a principle known as poetic license. Poetic license allows deviation in favour of making language inspiring. Deviation from linguistic norms often implies that writers can take liberties with word formation, thus neology in literary contexts should be addressed specifically. This article analyses the status of literary coinages in the scope of neology and describes the specific context of children’s literature. The article also offers a typology of nonce formation processes for occasionalisms, with textual analysis, from a corpus of children’s books, using J. Tournier’s matrices of lexicogenesis [2007: 51].
This paper presents a truly full character-level neural dependency parser together with a newly released character-level dependency treebank for Chinese, which has suffered a lot from the dilemma of defining word or not to model character interactions. Integrating full character-level dependencies with character embedding and human annotated character-level part-of-speech and dependency labels for the first time, we show an extra performance enhancement from the evaluation on Chinese Penn Treebank and SJTU (Shanghai Jiao Tong University) Chinese Character Dependency Treebank and the potential of better understanding deeper structure of Chinese sentences.
To approximately parse an unfamiliar language, it helps to have a treebank of a similar language. But what if the closest available treebank still has the wrong word order? We show how to (stochastically) permute the constituents of an existing dependency treebank so that its surface part-of-speech statistics approximately match those of the target language. The parameters of the permutation model can be evaluated for quality by dynamic programming and tuned by gradient descent (up to a local optimum). This optimization procedure yields trees for a new artificial language that resembles the target language. We show that delexicalized parsers for the target language can be successfully trained using such "made to order" artificial languages.
This article proposes a surface-syntactic annotation scheme called SUD that is near-isomorphic to the Universal Dependencies (UD) annotation scheme while following distributional criteria for defining the dependency tree structure and the naming of the syntactic functions. Rule-based graph transformation grammars allow for a bi-directional transformation of UD into SUD. The back-and-forth transformation can serve as an error-mining tool to assure the intralanguage and inter-language coherence of the UD treebanks.
As the amount of unstructured text data that humanity produces overall and on the Internet grows, so does the need to intelligently to process it and extract different types of knowledge from it. Convolutional neural networks (CNNs) and recurrent neural networks (RNNs) have been applied to natural language processing systems with comparative, remarkable results. The CNN is a noble approach to extract higher level features that are invariant to local translation. However, it requires stacking multiple convolutional layers in order to capture long-term dependencies, due to the locality of the convolutional and pooling layers. In this paper, we describe a joint CNN and RNN framework to overcome this problem. Briefly, we use an unsupervised neural language model to train initial word embeddings that are further tuned by our deep learning network, then, the pre-trained parameters of the network are used to initialize the model. At a final stage, the proposed framework combines former information with a set of feature maps learned by a convolutional layer with long-term dependencies learned via long-short-term memory. Empirically, we show that our approach, with slight hyperparameter tuning and static vectors, achieves outstanding results on multiple sentiment analysis benchmarks. Our approach outperforms several existing approaches in term of accuracy; our results are also competitive with the state-of-the-art results on the Stanford Large Movie Review data set with 93.3% accuracy, and the Stanford Sentiment Treebank data set with 48.8% fine-grained and 89.2% binary accuracy, respectively. Our approach has a significant role in reducing the number of parameters and constructing the convolutional layer followed by the recurrent layer as a substitute for the pooling layer. Our results show that we were able to reduce the loss of detailed, local information and capture long-term dependencies with an efficient framework that has fewer parameters and a high level of performance.
Recently, researchers have developed black-box approaches to mine design and interaction data from mobile apps. Although the data captured during this interaction mining is descriptive, it does not expose the design semantics of UIs: what elements on the screen mean and how they are used. This paper introduces an automatic approach for generating semantic annotations for mobile app UIs. Through an iterative open coding of 73k UI elements and 720 screens, we contribute a lexical database of 25 types of UI components, 197 text button concepts, and 135 icon classes shared across apps. We use this labeled data to learn code-based patterns to detect UI components and to train a convolutional neural network that distinguishes between icon classes with 94% accuracy. To demonstrate the efficacy of our approach at scale, we compute semantic annotations for the 72k unique UIs in the Rico dataset, assigning labels for 78% of the total visible, non-redundant elements.
We introduce a novel architecture for dependency parsing: stack-pointer networks (STACKPTR). Combining pointer networks The stack tracks the status of the depthfirst search and the pointer networks select one child for the word at the top of the stack at each step. The STACKPTR parser benefits from the information of the whole sentence and all previously derived subtree structures, and removes the leftto-right restriction in classical transitionbased parsers. Yet, the number of steps for building any (including non-projective) parse tree is linear in the length of the sentence just as other transition-based parsers, yielding an efficient decoding algorithm with O(n 2 ) time complexity. We evaluate our model on 29 treebanks spanning 20 languages and different dependency annotation schemas, and achieve state-of-theart performance on 21 of them.
We propose Efficient Neural Architecture Search (ENAS), a fast and inexpensive approach for automatic model design. In ENAS, a controller learns to discover neural network architectures by searching for an optimal subgraph within a large computational graph. The controller is trained with policy gradient to select a subgraph that maximizes the expected reward on the validation set. Meanwhile the model corresponding to the selected subgraph is trained to minimize a canonical cross entropy loss. Thanks to parameter sharing between child models, ENAS is fast: it delivers strong empirical performances using much fewer GPU-hours than all existing automatic model design approaches, and notably, 1000x less expensive than standard Neural Architecture Search. On the Penn Treebank dataset, ENAS discovers a novel architecture that achieves a test perplexity of 55.8, establishing a new state-of-the-art among all methods without post-training processing. On the CIFAR-10 dataset, ENAS designs novel architectures that achieve a test error of 2.89%, which is on par with NASNet (Zoph et al., 2018), whose test error is 2.65%.
This paper describes Stanford's system at the CoNLL 2018 UD Shared Task. We introduce a complete neural pipeline system that takes raw text as input, and performs all tasks required by the shared task, ranging from tokenization and sentence segmentation, to POS tagging and dependency parsing. Our single system submission achieved very competitive performance on big treebanks. Moreover, after fixing an unfortunate bug, our corrected system would have placed the 2 nd, 1 st, and 3 rd on the official evaluation metrics LAS, MLAS, and BLEX, and would have outperformed all submission systems on lowresource treebank categories on all metrics by a large margin. We further show the effectiveness of different model components through extensive ablation studies. * These authors contributed roughly equally.
International audience
When adding enhanced dependencies to an existing UD treebank, one can opt for heuristics that predict the enhanced dependencies on the basis of the UD annotation only. If the treebank is the result of conversion from an underlying treebank, an alternative is to produce the enhanced dependencies directly on the basis of this underlying annotation. Here we present a method for doing the latter for the Dutch UD treebanks. We compare our method with the UD -based approach of Schuster et al. (2018). While there are a number of systematic differences in the output of both methods, it appears these are the result of insufficient detail in the annotation guidelines and it is not the case that one approach is superior over the other in principle.
How to make the most of multiple heterogeneous treebanks when training a monolingual dependency parser is an open question. We start by investigating previously suggested, but little evaluated, strategies for exploiting multiple treebanks based on concatenating training sets, with or without fine-tuning. We go on to propose a new method based on treebank embeddings. We perform experiments for several languages and show that in many cases fine-tuning and treebank embeddings lead to substantial improvements over single treebanks or concatenation, with average gains of 2.0-3.5 LAS points. We argue that treebank embeddings should be preferred due to their conceptual simplicity, flexibility and extensibility.
The performance of a machine translation system depends on the availability of bilingual lexical dictionary and completion of its word sense disambiguation performance. Word sense disambiguation plays a vital role in several applications such as machine translation, information retrieval and many other Natural Language Processing (NLP). In oder to construct a reliable machine translation system, not only consistent lexical database like WordNet (WN), but also word sense distinguished database, and bilingual lexical machine readable dictionary (MRD) are required to achieve an accurate translation. Since WN is for the English language, there have been several approcaches building WN like database for other languages, also known as multilingual WN, by linking and extending bilingual MRDs with synsets from WN. Lexical database like WN for Myanmar languae have been proposed in recent years which are bilingual lexical databases as the synsets from WN are merged with bilingual MRD. However, there are few limitations in constructing bilingual database based on MRD where thesauries are derived from WN yet there are no directed alignemnt of meanings between English and Myanmar language. In order to design a complete lexical database with correct parallel meanings for both languages, our proposed method composes a model which accurately matches the similar meanings between English words and Myanmar words by translating and comparing the word senses for both languages. The experimental result shows that the meanings for both languages are aligned as accurate as the manual alignment of the word senses for lexical database.
This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and dependency parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin.
We know very little about how neural language models (LM) use prior linguistic context.In this paper, we investigate the role of context in an LSTM LM, through ablation studies.Specifically, we analyze the increase in perplexity when prior context words are shuffled, replaced, or dropped.On two standard datasets, Penn Treebank and WikiText-2, we find that the model is capable of using about 200 tokens of context on average, but sharply distinguishes nearby context (recent 50 tokens) from the distant history.The model is highly sensitive to the order of words within the most recent sentence, but ignores word order in the long-range context (beyond 50 tokens), suggesting the distant past is modeled only as a rough semantic field or topic.We further find that the neural caching model (Grave et al., 2017b) especially helps the LSTM to copy words from within this distant context.Overall, our analysis not only provides a better understanding of how neural LMs use their context, but also sheds light on recent success from cache-based models.
We present the Uppsala system for the CoNLL 2018 Shared Task on universal dependency parsing. Our system is a pipeline consisting of three components: the first performs joint word and sentence segmentation; the second predicts part-ofspeech tags and morphological features; the third predicts dependency trees from words and tags. Instead of training a single parsing model for each treebank, we trained models with multiple treebanks for one language or closely related languages, greatly reducing the number of models. On the official test run, we ranked 7th of 27 teams for the LAS and MLAS metrics. Our system obtained the best scores overall for word segmentation, universal POS tagging, and morphological features.
This paper presents a treebank for the healthcare domain developed at ezDI. The treebank is created from a wide array of clinical health record documents across hospitals. The data has been de-identified and annotated for constituent syntactic structure. The treebank contains a total of 52053 sentences that have been sampled for subdomains as well as linguistic variations. The paper outlines the sampling process followed to ensure a better domain representation in the corpus, the annotation process and challenges, and corpus statistics. The Penn Treebank tagset and guidelines were largely followed, but there were many syntactic contexts that warranted adaptation of the guidelines. The treebank created was used to re-train the Berkeley parser and the Stanford parser. These parsers were also trained with the GENIA treebank for comparative quality assessment. Our treebank yielded great-er accuracy on both parsers. Berkeley parser performed better on our treebank with an average F1 measure of 91 across 5-folds. This was a significant jump from the out-of-the-box F1 score of 70 on Berkeley parser’s default grammar.
Annotation corpus for discourse relations benefits NLP tasks such as machine translation and question answering. In this paper, we present SciDTB, a domain-specific discourse treebank annotated on scientific articles. Different from widely-used RST-DT and PDTB, SciDTB uses dependency trees to represent discourse structure, which is flexible and simplified to some extent but do not sacrifice structural integrity. We discuss the labeling framework, annotation workflow and some statistics about SciDTB. Furthermore, our treebank is made as a benchmark for evaluating discourse dependency parsers, on which we provide several baselines as fundamental work.
Web 2.0 has brought with it numerous user-produced data revealing one's thoughts, experiences, and knowledge, which are a great source for many tasks, such as information extraction, and knowledge base construction. However, the colloquial nature of the texts poses new challenges for current natural language processing techniques, which are more adapt to the formal form of the language. Ellipsis is a common linguistic phenomenon that some words are left out as they are understood from the context, especially in oral utterance, hindering the improvement of dependency parsing, which is of great importance for tasks relied on the meaning of the sentence. In order to promote research in this area, we are releasing a Chinese dependency treebank of 319 weibos, containing 572 sentences with omissions restored and contexts reserved.
In this paper we discuss the project of digitization of the Dictionary of the Serbo-Croatian Standard and Vernacular \nLanguage. Scanning and character recognition were a particular challenge, since various non-standard \ncharacter set encoding was used in the course of the almost 60-year long production of the dictionary. The first \naim of the project was to formalize the micro-structure of the dictionary articles in order to parse the digitized \ntext of and transform it into structured data stored in relational lexical database. This approach is compatible \nwith several standard structured forms and ontologies (TEI, LMF, Ontolex, LexInfo). A lexical database model \nwas designed in compliance with these structured forms, following mostly the lemon model. Mapping of \nthe lexical entry markers to LexInfo and TEI enabled export of the lexical data to the mentioned formats. A \nsoftware solution for the dictionary text analysis, parsing and lexical database population was developed and \ntested on the first and the last published volumes of the dictionary (which contain 27,141 articles in total). An \nevaluation of the results shows that the developed model and software solution can be successfully used for \nthe other volumes as well.
This paper describes the development of the first syntactically-annotated corpus of Breton. The corpus is part of the Universal Dependencies project. In the paper we describe how the corpus was prepared, some Breton-specific constructions that required special treatment, and in addition we give results for parsing Breton using a number of off-the-shelf data-driven parsers.
This paper describes a method of creating synthetic treebanks for cross-lingual dependency parsing using a combination of machine translation (including pivot translation), annotation projection and the spanning tree algorithm. Sentences are first automatically translated from a lesser-resourced language to a number of related highly-resourced languages, parsed and then the annotations are projected back to the lesser-resourced language, leading to multiple trees for each sentence from the lesser-resourced language. The final treebank is created by merging the possible trees into a graph and running the spanning tree algorithm to vote for the best tree for each sentence. We present experiments aimed at parsing Faroese using a combination of Danish, Swedish and Norwegian. In a similar experimental setup to the CoNLL 2018 shared task on dependency parsing we report state-of-the-art results on dependency parsing for Faroese using an off-theshelf parser.