Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
We introduce the task of word and phrase-level polarity annotation for German as part of an attempt to develop a compositional theory of clause-level polarity determination. Thus, annotations should give access to the nested building blocks, the structural strata of polarity composition. Therefore and in contrast to existing polarity-tagged corpora, we annotate not exclusively on the basis of surface strings, but argue that proper polarity annotation of complex phrases requires access to their syntactic structures. We discuss the principles of our treebank design, and present the inter-annotator agreement of our kick-off annotations on a test suite of 270 sentences that was compiled specifically to contain interesting polarity combinations.
Cataplexy is pathognomonic of narcolepsy with cataplexy, and defined by a transient loss of muscle tone triggered by strong emotions. Recent researches suggest abnormal amygdala function in narcolepsy with cataplexy. Emotion treatment and emotional regulation strategies are complex functions involving cortical and limbic structures, like the amygdala. As the amygdala has been shown to play a role in facial emotion recognition, we tested the hypothesis that patients with narcolepsy with cataplexy would have impaired recognition of facial emotional expressions compared with patients affected with central hypersomnia without cataplexy and healthy controls. We also aimed to determine whether cataplexy modulates emotional regulation strategies. Emotional intensity, arousal and valence ratings on Ekman faces displaying happiness, surprise, fear, anger, disgust, sadness and neutral expressions of 21 drug-free patients with narcolepsy with cataplexy were compared with 23 drug-free sex-, age- and intellectual level-matched adult patients with hypersomnia without cataplexy and 21 healthy controls. All participants underwent polysomnography recording and multiple sleep latency tests, and completed depression, anxiety and emotional regulation questionnaires. Performance of patients with narcolepsy with cataplexy did not differ from patients with hypersomnia without cataplexy or healthy controls on both intensity rating of each emotion on its prototypical label and mean ratings for valence and arousal. Moreover, patients with narcolepsy with cataplexy did not use different emotional regulation strategies. The level of depressive and anxious symptoms in narcolepsy with cataplexy did not differ from the other groups. Our results demonstrate that narcolepsy with cataplexy accurately perceives and discriminates facial emotions, and regulates emotions normally. The absence of alteration of perceived affective valence remains a major clinical interest in narcolepsy with cataplexy, and it supports the argument for optimal behaviour and social functioning in narcolepsy with cataplexy.
We describe the architecture we set up during the SANCL shared task for parsing usergenerated texts, that deviate in various ways from linguistic conventions used in available training treebanks. This architecture focuses in coping with such a divergence. It relies on the PCFG-LA framework (Petrov and Klein, 2007), as implemented by Attia et al. (2010). We explore several techniques to augment robustness: (i) a lexical bridge technique (Candito et al., 2011) that uses unsupervised word clustering (Koo et al., 2008); (ii) a special instanciation of self-training aimed at coping with POS tags unknown to the training set; (iii) the wrapping of a POS tagger with rulebased processing for dealing with recurrent non-standard tokens; and (iv) the guiding of out-of-domain parsing with predicted part-ofspeech tags for unknown words and unknown (word, tag) pairs. Our systems ranked second and third out of eight in the constituency parsing track of the SANCL competition. 1
This paper examines both linguistic behavior and practical implications of empty argument insertion in the Hindi PropBank. The Hindi PropBank is annotated on the Hindi Dependency Treebank, which contains some empty categories but rarely the empty arguments of verbs. In this paper, we analyze four kinds of empty arguments, *PRO*, *REL*, *GAP*, *pro*, and suggest effective ways of annotating these arguments. Empty arguments such as *PRO * and *REL * can be inserted deterministically; we present linguistically motivated rules that automatically insert these arguments with high accuracy. On the other hand, it is difficult to find deterministic rules to insert *GAP * and *pro*; for these arguments, we introduce a new annotation scheme that concurrently handles both semantic role labeling and empty category insertion, producing fast and high quality annotation. In addition, we present algorithms for finding antecedents of *REL * and *PRO*, and discuss why finding antecedents for some types of *PRO * is difficult.
The article discusses problems related to translating titles of films, literary works and academic publications from German to Polish and vice-versa. The authors emphasize aspects common to translation of all titles, and analyse typical problems occurring while rendering film and literary titles. In case of titles of academic publications formal and content-related changes are considerably smaller than, for instance, in film titles. This is due to the fact that while working on the latter translators take into account not only linguistic norms and conventions, but also references to the cultural background. A title has to be equally catchy in both source and target languages, and that is why translators perform certain operations when they are convinced that literal translation might yield poor results. Producing a title not related to the original is a borderline occurrence. The translator frequently needs to seek the golden mean between the content and form of the original and the requirements of the target language (linguistic norms, conventions, traditions).
This is an exploratory study investigating potential effects of emotional valence in images and their influence on conversation in the presence of the images. We used latent-semantic analysis to generalize valence ratings of Swedish words to a corpus of spoken conversations. Each utterance in the conversation was given a valence rating, which represented how emotionally positive or emotionally negative the utterance was. We found no effects that indicate that valenced images have an effect on conversations. However, we find that valenced images in general, and positive images in particular, were considered more helpful by the participants who engaged in the conversations. Additionally, we find no results that interlocutors align over time in their use of valenced language.
This thesis studies parsing and literature with the Data-Oriented Parsing framework, which assumes that chunks of previous experience can be exploited to analyze new sentences. As chunks we consider syntactic tree fragments. After presenting a method to efficiently extract such fragments from treebanks based on heuristics of re-occurrence, we employ them to develop a multi-lingual statistical parser. We show how a mildly context-sensitive grammar can be employed to produce discontinuous constituents, and compare this to an approximation that stays within the efficiently parsable context-free framework. We show that tree fragments allow the grammar to adequately capture the statistical regularities of non-local relations, without the need for the increased generative capacity of mildly context-sensitive grammar. The second part investigates what separates literary from other novels. We work with a corpus of novels and a reader survey with ratings of how literary they are perceived to be. The main goal is to find out the extent to which the literary ratings can be predicted from the texts. We first evaluate simple measures such as vocabulary richness, text compressibility, and the number of cliché expressions. In addition we apply more sophisticated, predictive models: a topic model, bag-of-words model, and a model based on syntactic tree fragments. We find that literary ratings are predictable from textual features to a large extent. While it is not possible to infer a causal relation, this result clearly rules out the notion that these value-judgments of literary merit were arbitrary, or predominantly determined by factors beyond the text.
A novel method for hybrid graph-based dependency parsing of natural language text is proposed. It is based on k-best maximum spanning tree dependency parsing and evaluation of the spanning trees by using a verb valency lexicon for a given language as a reranking knowledge base. The approach is compared with existing state-of-the-art transition-based and graph-based approaches to dependency parsing. As the proposed generic method was developed specifically for improving the accuracy of Croatian dependency parsing, Croatian Dependency Treebank and CROVALLEX verb valency lexicon are used in the experiment. The suggested approach scored approximately 77.21% LAS, outperforming the tested state-of-the-art approaches by at least 2.68% LAS.
Overview Since the early nineties, the on-going dramatic loss of the world’s linguistic diversity has gained attention, first by the linguists and increasingly also by the general public. As a response, the new field of language documentation emerged from around 2000 on, starting with the funding initiative ‘Dokumentation Bedrohter Sprachen’ (DoBeS, funded by the Volkswagen foundation, Germany), soon to be followed by others such as the ‘Endangered Languages Documentation Programme’ (ELDP, at SOAS, London), or, in the USA, ‘Electronic Meta-structure for Endangered Languages Documentation’ (EMELD, led by the LinguistList) and ‘Documenting Endangered Languages’ (DEL, by the NSF). From its very beginning, the new field focused on digital technologies not only for recording in audio and video, but also for annotation, lexical databases, corpus building and archiving, among others. This development not just coincides but is intrinsically interconnected with the increasing focus on digital data, technology and methods in all sciences, in particular in the humanities.
This article discusses how linguistic and translation norms, as evident in dictionaries, enforce the ideology of heteronormativity in Slovenia. The aim of this paper is to show how translation norms concerning homoeroticism were shaped in the translation of classical literature in Slovenia in the twentieth century. Translation norms are shaped in a certain period of time and in a certain environment among translators and others involved in translation according to social and cultural circumstances, expectations, and adaptations of topics to these expectations, in which the translation contrasts the initial norms with the norms of the target culture. At the same time, the linguistic norms of the standard language are created as a result of speakers’ continuous adaptation to a social and cultural environment, as a result of adapting to the current social ideal. It is assumed that translations contributed to creating a model of heteronormativity, which continues to characterize the Slovenian community today because of the limited number of new translations of classical works of literature. The paper concludes with a brief analysis of evidence of homosexuality in Slovenian translations of Shakespeare’s The Merchant of Venice.
This article attempts to prove the hypothesis on the “processual homogeneity” of the changes iR (*ŕ̥ *ir *yr) ≥ eR i iL ≥ eL. The source material for the following analysis has been excerpted from letters written in Polish between 1525–1550. Within the specified time brackets, the most dynamic is the process of the extension of the groups iR ≤ *ŕ̥, whereas the extended combinations eR ≤ *ir *yr are characterized by far lower intensity. The extensions within the groups of the type eL are not numerous and rare. Textual extensions of the combinations eR and eL show a significant convergence: forms with eL are to be found mostly in those letters that consistently demonstrate and provide evidence of the occurrence of the group eR; while they are virtually non-existent in the texts that document exclusively the combination iR. This fact confirms the hypothesis that combines both types of change into one process. The author does not associate the dependencies in the different pace of both phenomena and their different outcome with phonetic factors, but with the sensitivity to phonotactic and phonostatistical provisions and patterns, as well as with the morphological placement of the dyads under scrutiny. This interpretation, in turn, provides a new insight into the nature of the process of shaping linguistic norms: all the contributing processes have random (non-purposeful) origins, and their constraints involve a necessity to retain the necessary minimum of the internal balance.
We demonstrate that an unlexicalized PCFG with refined conjunction categories can parse much more accurately than previously shown, by making use of simple, linguistically motivated state splits, which break down false independence assumptions latent in a vanilla treebank grammar and reflect the Chinese idiosyncratic grammatical property. Indeed, its performance is the best result in the 3nd Chinese Parsing Evaluation of single model. This result has showed that refine the function words to represent Chinese subcat frame is a good method. An unlexicalized PCFG is much more compact, easier to replicate, and easier to interpret than more complex lexical models, and the parsing algorithms are simpler, more widely understood, of lower asymptotic complexity, and easier to optimize.
We investigated how shape features in natural images influence emotions aroused in human beings. Shapes and their characteristics such as roundness, angularity, simplicity, and complexity have been postulated to affect the emotional responses of human beings in the field of visual arts and psychology. However, no prior research has modeled the dimensionality of emotions aroused by roundness and angularity. Our contributions include an in-depth statistical analysis to understand the relationship between shapes and emotions. Through experimental results on the International Affective Picture System (IAPS) dataset we provide evidence for the significance of roundness-angularity and simplicitycomplexity on predicting emotional content in images. We combine our shape features with other state-of-theart features to show a gain in prediction and classification accuracy. We model emotions from a dimensional perspective in order to predict valence and arousal ratings which have advantages over modeling the traditional discrete emotional categories. Finally, we distinguish images with strong emotional content from emotionally neutral images with high accuracy.
We investigate the problem of automatically converting from a dependency representation to a phrase structure representation, a key aspect of understanding the relationship between these two representations for NLP work. We implement a new approach to this problem, based on a small number of supertags, along with an encoding of some of the underlying principles of the Penn Treebank guidelines. The resulting system significantly outperforms previous work in such automatic conversion. We also achieve comparable results to a system using a phrase-structure parser for the conversion. A comparison with our system using either the part-of-speech tags or the supertags provides some indication of what the parser is contributing. 1
We provide a model that extends the split-merge framework of Petrov et al. (2006) to jointly learn latent annotations and Tree Sub-stitution Grammars (TSGs). We then conduct a variety of experiments with this model, first inducing grammars on a portion of the Penn Treebank and the Korean Treebank 2.0, and next experimenting with grammar refinement from a single nonterminal and from the Uni-versal Part of Speech tagset. We present quali-tative analysis showing promising signs across all experiments that our combined approach successfully provides for greater flexibility in grammar induction within the structured guidance provided by the treebank, leveraging the complementary natures of these two ap-proaches. 1
Crawlers are basic entity that makes search engine to work efficiently in World Wide Web. Semantic Concept is implied into the search engine to provide precise and constricted search results which is required by end users of Internet. Search engine could be enhanced in searching mechanism through semantic Lexical Database such as WordNet, ConceptNet, YAGO, etc; Search results would be retrieved from Lexical and Semantic Knowledge Base (KB) by applying word sense and metadata technique based on the user query. The Uniform Resource Locater (URL) could be added and updated by the user to Semantic knowledge base so that crawlers can easily extract meta data and text which is available in specified web page. The proposed methodology enables web crawler to extract all meta tags and metadata from the web page which are stored in Semantic KB, hence search results are expected to be more significant and effective.
Current practices for sentiment prediction from text mostly involve words-in-a-bag approach that utilize techniques such as support vector machines or naïve Bayes. In this study, ANET (Affective Norms for English Text) sentence ratings of pleasure and arousal are compared with ANEW (Affective Norms for English Words) word ratings using regression and single layer neural networks. The sentences in ANET are decomposed into their words to obtain valence and arousal ratings from ANEW. A stop list is formed for non-words as well as words that are not found in ANEW. Then we studied whether the sentence sentiment reflected in terms of valence and arousal can be predicted from the sentiment of words in the sentence. Using linear regression, we found that approximately 35% of the variance in ANET valence and arousal ratings can be explained by ANEW valence and arousal ratings. Furthermore, Pearson correlation coefficient for ANEW and ANET ratings are similar for both valence and arousal, and close to 0.6. We also trained neural networks to investigate if non-linear approximations improved prediction of sentence sentiments from the constituent words. Out of several feedforward neural network configurations, a network with 200 hidden layer nodes turned out to be capable of identifying sentence sentiments accurately: the words' valence and arousal values explained 88% of the variance in the sentences' valence ratings and 91% of the variance in the sentences' arousal ratings. This preliminary study indicates that a proper choice of neural network might be adequate to estimate sentiments of sentences from sentiments of words.
We present a constituency parsing system for Modern Hebrew. The system is based on the PCFG-LA parsing method of Petrov et al. 2006, which is extended in various ways in order to accommodate the specificities of Hebrew as a morphologically rich language with a small treebank. We show that parsing performance can be enhanced by utilizing a language resource external to the treebank, specifically, a lexicon-based morphological analyzer. We present a computational model of interfacing the external lexicon and a treebank-based parser, also in the common case where the lexicon and the treebank follow different annotation schemes. We show that Hebrew word-segmentation and constituency-parsing can be performed jointly using CKY lattice parsing. Performing the tasks jointly is effective, and substantially outperforms a pipeline-based model. We suggest modeling grammatical agreement in a constituency-based parser as a filter mechanism that is orthogonal to the grammar, and present a concrete implementation of the method. Although the constituency parser does not make many agreement mistakes to begin with, the filter mechanism is effective in fixing the agreement mistakes that the parser does make. These contributions extend outside of the scope of Hebrew processing, and are of general applicability to the NLP community. Hebrew is a specific case of a morphologically rich language, and ideas presented in this work are useful also for processing other languages, including English. The lattice-based parsing methodology is useful in any case where the input is uncertain. Extending the lexical coverage of a treebank-derived parser using an external lexicon is relevant for any language with a small treebank.
Building large-scale semantic resource is one of the major tasks in Language Information Processing. We propose Feature Structure Theory, and apply this theory in building a large-scale Chinese semantic resource based on Penn Chinese Treebank corpus. The feature structure theory aims at addressing annotation problems from special sentence patterns, flexible word order, and serial noun phrase, etc., which are universal in Chinese. Annotation based on feature structure theory describes more semantic information than traditional approaches, and achieves higher annotating efficiency and higher accuracy.
The paper aims to represent a bilingual online dictionary as a useful tool helping preservation of the natural languages. The author focuses on the approach that was taken to develop compatible bilingual lexical database for the Bulgarian-Polish online dictionary. A formal model for the dictionary encoding is developed in accordance with the complex structures of the dictionary entries. These structures vary depending on the grammatical characteristics of Bulgarian headwords. The Webapplication for presentation of the bilingual dictionary is also describred.
This paper investigates the ways languages are used in Philadelphia Chinatown through qualitative content analysis of 330 photos. Examining the linguistic landscape of public spaces exposes issues of linguistic tensions, language vitality, and language shift in multilingual settings. While Chinese in the form of Mandarin is highly publicized, thereby placing disproportionate emphasis upon one language over others, Philadelphia Chinatown shows diversity, coexistence, and creative uses of multiple Chinese languages alongside English. The signage suggests linguistic rescaling connecting real and imagined audiences, conforming to broader ‘Chinese’ linguistic norms while localized to connect to a range of Chineses. We show how linguistic and cultural pluralism of ‘Chinese’ have always existed – and continue to exist – and the importance of developing socially sensitive literacy pedagogy, especially when there is a mismatch between the informal, community-level signage and what is formally taught in ‘Chinese’ language classrooms in the U.S. Keywords: linguistic landscape; Chinatown; Chinese languages; literacy education; heritage language; education
This paper presents a discriminative reranking model for the discourse segmentation task, the first step in a discourse parsing system. Our model exploits subtree features to rerank N-best outputs of a base segmenter, which uses syntactic and lexical features in a CRF framework. Experimental results on the RST Discourse Treebank corpus show that our model outperforms existing discourse segmenters in both settings that use gold standard Penn Treebank parse trees and Stanford parse trees. 1
We propose the first joint model for word segmentation, POS tagging, and dependency parsing for Chinese. Based on an extension of the incremental joint model for POS tagging and dependency parsing (Hatori et al., 2011), we propose an efficient character-based decoding method that can combine features from state-of-the-art segmentation, POS tagging, and dependency parsing models. We also describe our method to align comparable states in the beam, and how we can combine features of different characteristics in our incremental framework. In experiments using the Chinese Treebank (CTB), we show that the accuracies of the three tasks can be improved significantly over the baseline models, particularly by 0.6% for POS tagging and 2.4% for dependency parsing. We also perform comparison experiments with the partially joint models.
We propose a method for the extraction of a Tree Adjoining Grammar (TAG) from a dependency treebank which has some representative examples annotated with phrase structures. We show that the resulting TAG along with corresponding dependency structure can be used to convert a dependency treebank to a TAG-based phrase structure treebank.
Treebanks are language resources that provide annotations at various levels of linguistic structure starting from the word level. They typically provide syntactic constituent or dependency structures for sentences, but increasingly extend to annotation beyond syntactic structure, including semantic, pragmatic and rhetorical annotation, or go beyond a single language, as in parallel treebanks.
 Experience in building treebanks has shown that there is a close relation between formal linguistic theory and the design and practice of annotation. With increasing complexity of annotations, the design of annotation schemes becomes more and more theory-dependent. At the same time, linguistically motivated treebank annotations have become crucially important for the development of data-driven approaches to natural language processing and for linguistic research in general.
 Treebanks therefore constitute an important link between linguistic theory and computational linguistics.
 The International Workshop on Treebanks and Linguistic Theories provides a forum for researchers working on treebanks from both perspectives. The present volume presents the contents of the 10th edition of this workshop series, held in 2012 at the University of Heidelberg.
Recurrent neural network language models (RNNLMs) have recently demonstrated state-of-the-art performance across a variety of tasks. In this paper, we improve their performance by providing a contextual real-valued input vector in association with each word. This vector is used to convey contextual information about the sentence being modeled. By performing Latent Dirichlet Allocation using a block of preceding text, we achieve a topic-conditioned RNNLM. This approach has the key advantage of avoiding the data fragmentation associated with building multiple topic models on different data subsets. We report perplexity results on the Penn Treebank data, where we achieve a new state-of-the-art. We further apply the model to the Wall Street Journal speech recognition task, where we observe improvements in word-error-rate.
Individuals with autism spectrum disorders (ASD) demonstrate increased visual attention and elevated brain reward circuitry responses to images related to circumscribed interests (CI), suggesting that a heightened affective response to CI may underlie their disproportionate salience and reward value in ASD. To determine if individuals with ASD differ from typically developing (TD) adults in their subjective emotional experience of CI object images, non-CI object images and social images, 213 TD adults and 56 adults with ASD provided arousal ratings (sensation of being energized varying along a dimension from calm to excited) and valence ratings (emotionality varying along dimension of approach to withdrawal) for a series of 114 images derived from previous research on CI. The groups did not differ on arousal ratings for any image type, but ASD adults provided higher valence ratings than TD adults for CI-related images, and lower valence ratings for social images. Even after co-varying the effects of sex, the ASD group, but not the TD group, gave higher valence ratings to CI images than social images. These findings provide additional evidence that ASD is characterized by a preference for certain categories of non-social objects and a reduced preference for social stimuli, and support the dissemination of this image set for examining aspects of the circumscribed interest phenotype in ASD.
A neurocomputational model based on emergent massively overlapping neural cell assemblies (CAs) for resolving prepositional phrase (PP) attachment ambiguity is described. PP attachment ambiguity is a well-studied task in natural language processing and is a case where semantics is used to determine the syntactic structure. A large network of biologically plausible fatiguing leaky integrate-and-fire neurons is trained with semantic hierarchies (obtained from WordNet) on sentences with PP attachment ambiguity extracted from the Penn Treebank corpus. During training, overlapping CAs representing semantic similarities between the component words of the ambiguous sentences emerge and then act as categorizers for novel input. The resulting average resolution accuracy of 84.56% is on par with known machine learning algorithms.
A treebank is an important resource for developing many NLP based tools. Errors in the treebank may lead to error in the tools that use it. It is essential to ensure the quality of a treebank before it can be deployed for other purposes. Automatic (or semi-automatic) detection of errors in the treebank can reduce the manual work required to find and remove errors. Usually, the errors found automatically are manually corrected by the annotators. There is not much work reported so far on error correction tools which helps the annotators in correcting errors efficiently. In this paper, we present such an error correction tool that is an extension of the error detection method described earlier (Ambati et al., 2010; Ambati et al., 2011; Agarwal et al., 2012). Keywords:Treebank, Error Detection, Graphical User Interface
Most of the reliable language resources are developed via human supervision. Developing supervised annotated data is hard and tedious, and it will be very time consuming when it is done totally manually; as a result, various types of annotated data, including treebanks, are not available for many languages. Considering that a portion of the language is regular, we can define regular expressions as grammar rules to recognize the strings which match the regular expressions, and reduce the human effort to annotate further unseen data. In this paper, we propose an incremental bootstrapping approach via extracting grammar rules when no treebank is available in the first step. Since Persian suffers from lack of available data sources, we have applied our method to develop a treebank for this language. Our exper-iment shows that this approach significantly decreases the amount of manual effort in the annotation process while enlarging the treebank. Keywords:Treebank Development, Bootstrapping Approach, Grammar Rule Extraction, the Persian Language 1.
Learning a foreign language (FL) entails more than attaining the mastery of a system of linguistic norms or the functional and pragmatic aspects of that language. It requires learning to adapt to different cultural norms. So, the challenge is to provide FL learners with opportunities to interact with people from other cultural and linguistic realities and re-focus the aims of FL learning to the development of intercultural communicative competence (ICC). The introduction of information and communication technology (ICT) in education not only enhances the access to information but also enables intercultural contact among individuals from diverse cultural backgrounds, setting the conditions for the development of curriculum-based telecollaboration projects. The LOA eTwinning project presented in this chapter was implemented in the context of an action-research project aimed to introduce an intercultural approach to teaching English to raise pupils’ motivation and challenge them to become more creative, more collaborative, and more autonomous.
The aim here is to create a dependency treebank from a phrase-structure treebank for Arabic. Arabic has a number of characteristics,described below, which make it particularly challenging to any natural language processing (NLP) applications. We describe an encouraging semi-automatic technique for converting phrase-structure trees to dependency trees by using a head percolation table.One of the most significant challenges here is the determination of the head of each subtree. We therefore examined different versionsof the head percolation table to find the best priority list for each entry in the table. Given that there is no absolute measure of the‘correctness’ of a conversion of a phrase structure tree to dependency form, we tested the various transformations by seeing how well astate-of-the-art dependency parser learnt the generalisations that were embodied by the converted trees.
Conversion between different grammar frameworks is of great importance to comparative performance analysis of the parsers developed based on them and to discover the essential nature of languages. This paper presents an approach that converts Combinatory Categorial Grammar (CCG) derivations to Penn Treebank (PTB) trees using a maximum entropy model. Compared with previous work, the presented technique makes the conversion practical by eliminating the need to develop mapping rules manually and achieves state-of-the-art results.
Treebanks are a linguistic resource: a large database where the morphological, syntactic and lexical information for each sentence has been explicitly marked. The critical requirements of treebanks for various NLP activities (research and application) are well known. This also implies that treebanks need to be as error free as possible. However, manual validation of a treebank is very costly, both in terms of time and money. This paper describes an approach to automatically detect errors in a treebank after a complete manual annotation. Over and above improving an earlier error detection tool (Ambati et al. (2011)) for a Hindi treebank. We also present a user study to show that our system reduces the validation time significantly while detecting 81.49% of the errors at the dependency level.
This three-variable predictive model has excellent predictive ability in both the derivation cohort and the validation cohort. This model can identify women who are at high risk of non-initiating breastfeeding within the first hour after delivery.
We present a simple and effective framework for exploiting multiple monolingual treebanks with different annotation guidelines for pars-ing. Several types of transformation patterns (TP) are designed to capture the systematic an-notation inconsistencies among different tree-banks. Based on such TPs, we design quasi-synchronous grammar features to augment the baseline parsing models. Our approach can significantly advance the state-of-the-art pars-ing accuracy on two widely used target tree-banks (Penn Chinese Treebank 5.1 and 6.0) using the Chinese Dependency Treebank as the source treebank. The improvements are respectively 1.37 % and 1.10 % with automatic part-of-speech tags. Moreover, an indirect comparison indicates that our approach also outperforms previous work based on treebank conversion. 1
MT systems typically use parsers to help reorder constituents. However most languages do not have adequate treebank data to learn good parsers, and such training data is extremely time-consuming to annotate. Our earlier work has shown that a reordering model learnt from word-alignments using POS tags as features can improve MT performance (Visweswariah et al., 2011). In this paper, we investigate the effect of word-classing on reordering performance using this model. We show that unsupervised word clusters perform somewhat worse but still reasonably well, compared to a part-of-speech (POS) tagger built with a small amount of annotated data; while a richer tagset including case and gender-number-person further improves reordering performance by around 1.2 monolingual BLEU points. While annotating this richer tagset is more complicated than annotating the base tagset, it is much easier than annotating treebank data. Keywords:SMT, Reordering, POS-Tagging 1.
Syntactically annotated corpora have become important resources for natural language processing due in part to the success of corpus-based methods. Since words are often considered as primitive units of language structures, the annotation of word segmentation forms the basis of these corpora. This is also an issue for the Vietnamese Treebank (VTB), which is the first and only publicly available syntactically annotated corpus for the Vietnamese language. Although word segmentation is straight-forward for space-delimited languages like English, this is not the case for languages like Vietnamese for which a standard criterion for word segmentation does not exist. This work explores the challenges of Vietnamese word segmentation through the detection and correction of inconsistency for VTB. Then, by combining and splitting the inconsistent annotations that were detected, we are able to observe the influence of different word segmentation criteria on automatic word segmentation, and the applications of word segmentation, including text classification and English-Vietnamese statistical machine translation. The analysis and experimental results showed that our methods improved the quality of VTB, which positively affected the performance of its applications. Title and Abstract in another language, L2 (optional, and on same page) So sánh các tiêu chí tách từ khác nhau thông qua ứng dụng
Traditional Chinese grammar is represented by Li Jinxi's A New Chinese Grammar firstly.Li's grammar system,taking the sentence layout and the sentence component as its main characteristic,is called the sentence-based grammar.This paper firstly briefly reviews the development history of the Chinese grammar,and summarizes the main ideas and theoretical features of two schools: traditional grammar and structural grammar.Then the paper analyzes the advantages and disadvantages of the main grammar systems in Chinese Information Process(CIP) from the view of the Chinese treebank,and compared them to traditional grammar to reveal the necessity of applying traditional grammar to CIP field.Finally,the paper discusses some key issues to be handled in the future of application.
This paper presents an identification framework for extracting Tibetan base noun phrase (NP). The framework includes two phases. In the first phase, Chinese base NPs are extracted from all Chinese sentences in the sentence aligned Chinese-Tibetan corpus using Stanford Chinese parser. In the second phase, the Tibetan translations of those Chinese NPs are identified using four different methods, that is, word alignment, iterative re-evaluation, dictionary and word alignment, and sequence intersection method. We implemented and tested these methods on Chinese-Tibetan sentence aligned unlabelled corpus without Tibetan POS tagger and Treebank. The experimental results demonstrate these methods can get satisfactory results, and the best performance with 0.5283 precision is got using sequence intersection identification method. The identification framework can also be extended to extract Tibetan verb phrase.
This contribution explores the subgroup of text structuring expressions with the form preposition + demonstrative pronoun, thus it is devoted to an aspect of the interaction of coreference relations and relations signaled by discourse connectives (DCs) in a text. The demonstrative pronoun typically signals a referential link to an antecedent, whereas the whole expression can, but does not have to, carry a discourse meaning in sense of discourse connectives. We describe the properties of these phrases/expressions with regard to their antecedents, their position among the text-structuring language means and their features typical for the “connective function ” of them compared to their “non-connective function”. The analysis is carried out on Czech data from the approx. 50,000 sentences of the Prague Dependency Treebank 2.0, directly on the syntactic trees. We explore the characteristics of these phrases/expressions discovered during two projects: the manual annotation of coreference relations (Nedoluzhko et al. 2011) and discourse connectives, their scopes and meanings (Mladová et al. 2008).
"This paper addresses the problem of optimizing the training treebank data because the size and quality of the data has always been a bottleneck for the purposes of training. In previous studies we realized that current corpora used for training machine learning{based dependency parsers contain a significant proportion of redundant information at the syntactic structure level. Since the development of such training corpora involves a big effort, we argue that an appropriate process for selecting the sentences to be included in them can result in having parsing models as accurate as the ones given when training with bigger { non optimized corpora (or alternatively, bigger accuracy for an equivalent annotation effort). This argument is supported by the results of the study we carried out, which is presented in this paper. Therefore, this paper demonstrates that the training corpora contain more information than needed for training accurate data{driven dependency parsers."
The objective of textology (text linguistics) is to analyse and describe interpersonal communication in all its aspects, as it happens now and did in the past, in various discourse communities. The basic categories of so defined textology are: text, utterance and discourse, regarded as different perspectives on the phenomenon of interpersonal communication. This phenomenon manifests itself in established works, with their specific semantic and syntactic structure, but also in genre affinity, in interactive events fixed in a widely understood situational context, and in social, cultural and linguistic norms which regulate communication activities in individual human communities and make interpersonal communication possible. These three categories can be said to reify in definite communication phenomena: in written and recorded text, and in spoken utterances, i.e. in actualized discourses.
This article presents the phonetic and inflection linguistic changes in the three editions of Krzysztof Kluk’s The trees, garden herbs and gardens (edit. 1777, 1797, 1808). The first part of the analysis focuses on those changes that are consistent with the evaluation of the Polish nationwide standardized language at the turn of the eighteenth and nineteenth centuries. The second part of the article presents linguistic changes where the modern and correct forms are changed into the older one, consistent with the north-eastern Polish language of the Eastern Borderlands. Ascertaining why the editors of Kluk’s book used the older linguistic forms instead of the correct forms, leads to attempts to determine the standard linguistic norms in the Piarists’ printing house during the very short period at the turn of the eighteenth and the nineteenth centuries.