Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This preliminary study looks at the familiarity rating of words in Sarawak Malay Dialect (SMD). Although familiarity ratings of language items are usually utilised in psycholinguistic research, they can be very useful for studies in the area of language change. The aim of this study is twofold: (1) to compare the perceptive familiarity rating of Sarawak Malay words; and (2) to document Sarawak Malay words that are undergoing lexical change. Fifty SMD words were used in this study consisting of those with meanings that can be considered as medium to high in frequency for everyday speech. Questionnaires were designed using a 5-point Likert-type scale to rate word familiarity and distributed to 15 participants who were native SMD speakers between the ages of 20 and 25. Across word items, more than one third were found to be rated as less familiar and unknown and thus were not actively used in daily conversations. There were also a number of words perceived as familiar to highly familiar but were not widely used in everyday speech. Evidently, it is crucial to document and preserve these SMD words as they are fast becoming passive vocabulary for the young and may eventually be lost in their lexicon.
The study presented in this article is dedicated to a syntactic parser for Romanian. The central goal of the presented technique is to learn a model which is able to discriminate between probability for a word to be head of another word in a dependency structure corresponding to a sentence in the considered language. The model described in this paper was trained on a dependency treebank linguistic resource and is intended to be used in order to develop a dependency syntactic parser.
International audience
Cet article constitue une version réduite de l'article "The French Social Media Bank: a Treebank of Noisy User Generated Content" (mêmes auteurs)
In this paper we present and justify methodological principles and syntactic criteria to design an annotation scheme for a Persian Treebank. The advantages of the proposed scheme for annotation of the Persian Treebank will be discussed. At the same time, we present the way that different types of linguistic knowledge (morphological, syntactic and semantic) are encoded in the structures of the schema. We will show how this scheme can account for many of the syntactic constructions that appear to be unique to the Persian language.
Unknown words, or out of vocabulary words (OOV), cause a significant problem to morphological analysers, syntactic parses, MT systems and other NLP applications. Unknown words make up 29 % of the word types in in a large Arabic corpus used in this study. With today's corpus sizes exceeding 10 9 words, it becomes impossible to manually check corpora for new words to be included in a lexicon. We develop a finite-state morphological guesser and integrate it with a machine-learning-based pre-annotation tool in a pipeline architecture for extracting unknown words, lemmatizing them, and giving them a priority weight for inclusion in a lexical database. The processing is performed on a corpus of contemporary Arabic of 1,089,111,204 words. Our method is tested on a manually-annotated gold standard and yields encouraging results despite the complexity of the task. Our work shows the usability of a highly
State-of-the-art dependency representations such as the Stanford Typed Dependencies may represent the grammatical relations in a sentence as directed, possibly cyclic graphs. Querying a syntactically annotated corpus for grammatical structures that are represented as graphs requires graph matching, which is a non-trivial task. In this paper, we present an algorithm for graph matching that is tailored to the properties of large, syntactically annotated corpora. The implementation of the algorithm is built on top of the popular IMS Open Corpus Workbench, allowing corpus linguists to re-use existing infrastructure. An evaluation of the resulting software, CWB-treebank, shows that its performance in real world applications, such as a web query interface, compares favourably to implementations that rely on a relational database or a dedicated graph database while at the same time offering a greater expres-sive power for queries. An intuitive graphical interface for building the query graphs is available via the Treebank.info project.
This paper presents the work of the Hong Kong Polytechnic University (PolyUCOMP) team which has participated in the Semantic Textual Similarity task of SemEval-2012. The PolyUCOMP system combines semantic vectors with skip bigrams to determine sentence similarity. The semantic vector is used to compute similarities between sentence pairs using the lexical database WordNet and the Wikipedia corpus. The use of skip bigram is to introduce the order of words in measuring sentence similarity. 1
This paper describes the support for mouth activity annotation provided by the iLex annotation workbench on a holistic level connected to the lexical database, on a feature level, as well as in the context of semi-automatic annotation.
Query expansion is a crucial step in recall-oriented domains such as Patent Searching. Currently, automatic query expansion in patent search is mostly based on statistical measures. Additional query terms are extracted from the query documents based on entropy measures. To automate query expansion in patent searching, we acquire lexical knowledge from Query Logs of USPTO Patent Examiners. Results show good performance in query expansion and patent searching using the lexical database. This will help improving (semi-) automated query expansion in patent searching.
Treebanking a large corpus of relatively structured speech transcribed from various Arabic Broadcast News (BN) sources has allowed us to begin to address the many challenges of annotating and parsing a speech corpus in Arabic. The now completed Arabic Treebank BN corpus consists of 432,976 source tokens (517,080 tree tokens) in 120 files of manually transcribed news broadcasts. Because news broadcasts are predominantly scripted, most of the transcribed speech is in Modern Standard Arabic (MSA). As such, the lexical and syntactic structures are very similar to the MSA in written newswire data. However, because this is spoken news, cross-linguistic speech effects such as restarts, fillers, hesitations, and repetitions are common. There is also a certain amount of dialect data present in the BN corpus, from on-the-street interviews and similar informal contexts. In this paper, we describe the finished corpus and focus on some of the necessary additions to our annotation guidelines, along with some of the technical challenges of a treebanked speech corpus and an initial parsing evaluation for this data. This corpus will be available to the community in 2012 as an LDC publication.
Continuous self-reported emotion expressed by four pieces of music were collected on a two-dimensional (valence and arousal) emotion space in a repeated measures (test-retest conditions) design. Initial orientation time (IOT), test-retest reliability and afterglow were examined. Median IOT was 8 seconds. Valence ratings took up to 25 (median 4), and for arousal up to 35 (median 12) seconds. Slower tempi seemed to require longer IOT. Test-retest reliability examined correlation coefficients, and compared periods of sample-by-sample good agreement in response between Test and Retest condition. About 80% of responses were reliable in both the Test and Retest conditions regardless of response dimension. Pearson correlations demonstrated better test-retest reliability for arousal responses than for valence. Retest condition ratings were within 8% of Test condition rating within participant. Average standard deviations for ratings collapsed across dimension, stimulus and conditions was 12.2% of the ratings scale range. Afterglow effects – large outliers in spread of scores just after the end of a piece – were identified. The reliability of continuous emotional response is therefore considered to be quite good, but caution must be taken as to how to deal with the opening and ending of continuous emotional response data.
The lack of annotated corpora brings limitations in research of discourse classification for many languages. In this paper, we present the first effort towards recognizing ambiguities of discourse connectives, which is fundamental to discourse classification for resource-poor language such as Chinese. A language independent framework is proposed utilizing bilingual dictionaries, Penn Discourse Treebank and parallel data between English and Chinese. We start from translating the English connectives to Chinese using a bi-lingual dictionary. Then, the ambiguities in terms of senses a connective may signal are estimated based on the ambiguities of English connectives and word alignment information. Finally, the ambiguity between discourse usage and non-discourse usage were disambiguated using the co-training algorithm. Experimental results showed the proposed method not only built a high quality connective lexicon for Chinese but also achieved a high performance in recognizing the ambiguities. We also present a discourse corpus for Chinese which will soon become the first Chinese discourse corpus publicly available.
Mood is an important aspect of music and knowledge of mood can be used as a basic feature in music recommender and retrieval systems. A listening experiment was carried out establishing ratings for various moods and a number of attributes, e.g., valence and arousal. The analysis of these data covers the issues of the number of basic dimensions in music mood, their relation to valence and arousal, the distribution of moods in the valence-arousal plane, distinctiveness of the labels, and appropriate (number of) labels for full coverage of the plane. It is also shown that subject-averaged valence and arousal ratings can be predicted from music features by a linear model.
We also show ensemble dependency parsing and self training approaches applicable to under-resourced languages using our manually annotated dependency structures. We show that for an under-resourced language, the use of tuning data for a meta classifier is more effective than using it as additional training data for individual parsers. This meta-classifier creates an ensemble dependency parser and increases the dependency accuracy by 4.92% on average and 1.99% over the best individual models on average. As the data sizes grow for the the under-resourced language a meta classifier can easily adapt. To the best of our knowledge this is the first full implementation of a dependency parser for Indonesian. Using self-training in combination with our Ensemble SVM Parser we show additional improvement. Using this parsing model we plan on expanding the size of the corpus by using a semi-supervised approach by applying the parser and correcting the errors, reducing the amount of annotation time needed.
Language resources are essential for linguistic research and the development of NLP applications. Low-density languages, such as Irish, therefore lack significant research in this area. This paper describes the early stages in the development of new language resources for Irish – namely the first Irish dependency treebank and the first Irish statistical dependency parser. We present the methodology behind building our new treebank and the steps we take to leverage upon the few existing resources. We discuss language-specific choices made when defining our dependency labelling scheme, and describe interesting Irish language characteristics such as prepositional attachment, copula and clefting. We manually develop a small treebank of 300 sentences based on an existing POS-tagged corpus and report an inter-annotator agreement of 0.7902. We train MaltParser to achieve preliminary parsing results for Irish and describe a bootstrapping approach for further stages of development.
This paper describes a method to convert existing treebanks with syntactic information into banks of meaning representations. The central component is a system of evaluation for a small formal language with respect to an information state. Inputs to the evaluation system are formal language expressions obtained from the conversion of parsed representations conforming to (Penn Treebank Project) guidelines. Outputs from the evaluation system are Davidsonian (higher-order) predicate logic meaning representations. Having a system of evaluation as the basis for generating meaning representations makes possible accepting input with minimal conversion from existing treebanks and from the tools used to construct treebanks. Results of having built corresponding banks of meaning representations from available treebanks are discussed.
This article presents an online dictionary environment, with enhanced sorting and searching functionalities and a text to speech feature, for hearing the pronunciation of the words. The online dictionary environment has been developed as part of the ‘Syntychies’ research program. ‘Syntychies’ online environment is a pioneering webservice for Greek dialectal lexicography and it is the first of its kind for Cypriot Greek.
The Latvian Treebank is being developed since 2010. In this paper we describe the latest developments of this project and the problems currently faced. We examine several gaps in our annotation scheme like determinant, ellipsis and insertion annotation and describe solutions we have chosen.
We present a detailed error analysis of a transition-based dependency parser trained on a Hindi dependency treebank. Parser error analysis has not been systematically examined from the point of view of treebanking before and this work intends to contribute in this area. We address two main questions in this paper: Can the parsing of certain structures be made easier by using alternative analyses for these structures? Are there certain linguistic cues implicit (or missing) in the current treebank that can be made explicit (or added) in order to make the parsing of complex constructions easier? These questions will guide us in examining the potential benefits of parser error analysis during treebanking. Through our experiments and analysis we were able to shed light on the causes of errors and subsequently have been able to improve the performance of the parser.
Dependency parsing has attracted considerable interest from researchers and developers in natural language processing. However, to obtain a high‐accuracy dependency parser, supervised techniques require a large volume of hand‐annotated data, which are extremely expensive. This paper presents a simple and effective approach for improving dependency parsing with subtrees derived from unannotated data, which are easy to obtain. First, we use a baseline parser to parse large‐scale unannotated data. Then, we extract subtrees from dependency parse trees in the auto‐parsed data. Next, the extracted subtrees are classified into several sets according to their frequency. Finally, we design new features based on the subtree sets for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we conduct experiments on the English Penn Treebank and Chinese Penn Treebank. The results show that our approach significantly outperforms baseline systems. It also achieves the best accuracy for the Chinese data and an accuracy competitive with the best known systems for the English data.
After a period when the focus was essentially on mental architecture, the cognitive sciences are increasingly integrating the social dimension. The rise of a cognitive sociolinguistics is part of this trend. The article argues that this process requires a re-evaluation of some entrenched positions in linguistics: those that see linguistic norms as antithetical to a descriptive and variational linguistics. Once such a re-evaluation has taken place, however, the social recontextualization of cognition will enable linguistics (including sociolinguistics as an integral part), to eliminate the cracks in the foundations that were the result of suppressing the sociocultural underpinnings of linguistic facts. Structuralism, cognitivism and social constructionism introduced new and necessary distinctions, but in their strong forms they all turned into unnecessary divides. The article tries to show that an evolutionary account can reintegrate the opposed fragments into a whole picture that puts each of them in their ‘ecological position’ with respect to each other. Empirical usage facts should be seen in the context of operational norms in relation to which actual linguistic choices represent adaptations. Variational patterns should be seen in the context of structural categories without which there would be only ‘differences’ rather than variation. And emergence, individual choice, and flux should be seen in the context of the individual’s dependence on lineages of community practice sustained by collective norms.
In this article, we investigate ambiguity in syntactic annotation. The ambiguity in question is inherent in a way that even human annotators interpret the meaning differently. In our experiment, we detect potential structurally ambiguous sentences with Constraint Grammar rules. In the linguistic phenomena we investigate, structural ambiguity is primarily caused by word order. The potentially ambiguous particle or adverbial is located between the main verb and the (participial) NP. After detecting the structures, we analyze how many of the potentially ambiguous cases are actually ambiguous using the double-blind method. We rank the sentences captured by the rules on a 1 to 5 scale to indicate which reading the annotator regards as the primary one. The results indicate that 67% of the sentences are ambiguous. Introducing ambiguity in the treebank/parsebank increases the informativeness of the representation since both correct analyses are presented.
We describe a transformation-based learning method for learning a sequence of mono-lingual tree transformations that improve the agreement between constituent trees and word alignments in bilingual corpora. Using the manually annotated English Chinese Transla-tion Treebank, we show how our method au-tomatically discovers transformations that ac-commodate differences in English and Chi-nese syntax. Furthermore, when transforma-tions are learned on automatically generated trees and alignments from the same domain as the training data for a syntactic MT system, the transformed trees achieve a 0.9 BLEU im-provement over baseline trees. 1
Decoding pain in others is of high individual and social benefit in terms of harm avoidance and demands for accurate care and protection. The processing of facial expressions includes both specific neural activation and automatic congruent facial muscle reactions. While a considerable number of studies investigated the processing of emotional faces, few studies specifically focused on facial expressions of pain. Analyses of brain activity and facial responses elicited by the perception of facial pain expressions in contrast to other emotional expressions may unravel the processing specificities of pain-related information in healthy individuals and may contribute to explaining attentional biases in chronic pain patients. In the present study, 23 participants viewed short video clips of neutral, emotional (joy, fear), and painful facial expressions while affective ratings, event-related brain responses, and facial electromyography (Musculus corrugator supercilii, M. orbicularis oculi, M. zygomaticus major, M. levator labii) were recorded. An emotion recognition task indicated that participants accurately decoded all presented facial expressions. Electromyography analysis suggests a distinct pattern of facial response detected in response to happy faces only. However, emotion-modulated late positive potentials revealed a differential processing of pain expressions compared to the other facial expressions, including fear. Moreover, pain faces were rated as most negative and highly arousing. Results suggest a general processing bias in favor of pain expressions. Findings are discussed in light of attentional demands of pain-related information and communicative aspects of pain expressions.
The paper presents and evaluates an efficient algorithm for measuring semantic similarity of texts. Calculating the level of semantic similarity of texts is a very difficult task and the proposed up to now methods suffer from computational complexity. This substantially limits their application area. The proposed algorithm tries to reduce the problem by merging a computationally efficient statistical approach to text analysis with a semantic component. The semantic properties of text words are extracted from the WordNet lexical database. The approach was tested using WordNets for two languages: English and Polish. The basic properties of this approach are also studied. The paper concludes with an analysis of the performance of the proposed method on a sample database and suggests some possible application areas.
Unsupervised dependency parsing is one of the most challenging tasks in natural languages processing. The task involves finding the best possible dependency trees from raw sentences without getting any aid from annotated data. In this paper, we illustrate that by applying a supervised incremental parsing model to unsupervised parsing; parsing with a linear time complexity will be faster than the other methods. With only 15 training iterations with linear time complexity, we gain results comparable to those of other state of the art methods. By employing two simple universal linguistic rules inspired from the classical dependency grammar, we improve the results in some languages and get the state of the art results. We also test our model on a part of the ongoing Persian dependency treebank. This work is the first work done on the Persian language. 1
State of the art parsers are currently trained on converted versions of Penn Treebank into dependency representations which however don’t include null elements. This is done to facilitate structural learning and prevent the probabilistic engine to postulate the existence of deprecated null elements everywhere (see [15]). However it is a fact that in this way, the semantics of the representation used and produced on runtime is inconsistent and will reduce dramatically its usefulness in real life applications like Information Extraction, Q/A and other semantically driven fields by hampering the mapping of a complete logical form. What systems have come up with are “Quasi”-logical forms or partial logical forms mapped directly from the surface representation in dependency structure. We show the most common problems derived from the conversion and then describe an algorithm that we have implemented to apply to our converted Italian Treebank, that can be used on any CONLL-style treebank or representation to produce an “almost complete” semantically consistent dependency treebank.
A major computational burden, while performing document clustering, is the calculation of similarity measure between a pair of documents. Similarity measure is a function that assign a real number between 0 and 1 to a pair of documents, depending upon the degree of similarity between them. A value of zero means that the documents are completely dissimilar whereas a value of one indicates that the documents are practically identical. Traditionally, vector-based models have been used for computing the document similarity. The vector-based models represent several features present in documents. These approaches to similarity measures, in general, cannot account for the semantics of the document. Documents written in human languages contain contexts and the words used to describe these contexts are generally semantically related. Motivated by this fact, many researchers have proposed semantic-based similarity measures by utilizing text annotation through external thesauruses like WordNet (a lexical database). In this paper, we define a semantic similarity measure based on documents represented in topic maps. Topic maps are rapidly becoming an industrial standard for knowledge representation with a focus for later search and extraction. The documents are transformed into a topic map based coded knowledge and the similarity between a pair of documents is represented as a correlation between the common patterns. The experimental studies on the text mining datasets reveal that this new similarity measure is more effective as compared to commonly used similarity measures in text clustering.
Retrieval practice for some memory items from a given category can impair subsequent retrieval of unpracticed items from the same category (retrieval-induced forgetting, RIF). Inhibition of these items has been invoked as an explanation, and inhibition has also been proposed to cause stimulus devaluation. The present experiments investigated whether a similar devaluation effect can be observed in a RIF experiment for the unpracticed and presumably inhibited items. We report two experiments using the RIF paradigm, and both experiments yielded a RIF effect. At the same time, affective ratings of the very same items did not show signs of devaluation. These results run counter the idea that both RIF and devaluation effects are caused by a (similar) inhibitory mechanism, or at least they suggest differences between the mechanisms involved in external perceptual and internal memory selection.
We present a system for cross-lingual parse disambiguation, exploiting the assumption that the meaning of a sentence remains unchanged during translation and the fact that different languages have different ambiguities. We simultaneously reduce ambiguity in multiple languages in a fully automatic way. Evaluation shows that the system reliably discards dispreferred parses from the raw parser output, which results in a pre-selection that can speed up manual treebanking. 1
The Slovene language is often presented as a national element. Even in the 19th century, which saw the Spring of Nations and the United Slovenia project, the Slovene language was a constitutive element of the Slovene nation. In the meantime, the Slovene language was positioning itself as an all-Slovene language, trying to be supra-regional. By the end of the 19th and early 20th centuries, the Slovene written language had stabilized, while at the same time the spoken language had only begun to assert itself. During this time, the prevailing principle was to "speak the way the language is written." In the mid-20th century, the theoretical idea of a literary language that is based on the central Slovene-speech (i.e. the speech of Ljubljana) came to dominate. In the third millennium, the question is whether a regionally-defined speech can be used as the basis for a Standard language. Another central question is what this "suitable" regionally-conditioned speech would be like. The principle of how important, decision-wise, the centre of a nation is, when it comes to questions of linguistic norms, may seem very attractive and, to a certain extent, logical. However, even examples of historically and linguistically comparable languages do not support the theory of creating the norm for the Standard Slovene language, based on the contemporary speech of Ljubljana, as claimed by Toporišič in Slovenska slovnica and, later, in Slovenski pravopis. Within Slovenia, the Standard Slovene language is tied to written language, which has proven, in the past, to be a suitable way of setting the norm. Regressing back to the principles of standardising a language, based on regional variants, would be unproductive, would introduce needless discord, and would cause problems with everyday, public communication. Contemporary research of actual speech, a portion of which is also presented within this article, confirms the all-Slovene and regionally-independent character of the Slovene Standard language.
Starting from the definition of treebanks and considering that treebanks are theory dependent, we propose an annotation scheme for Romanian using several approaches ranging from phrase structure to dependency grammars and property grammars. The annotation has its starting point in a generative grammar study of the Romanian AP and validates the data of the linguistic study using an annotation scheme consisting of a constraint based approach.
The degree of translation adequacy and full-value depends on its compliance with the existing general linguistic norms. Vocabulary potential of a translator is determined by the proficiency of language to translate into. Key moments in course of transferring means of another language text are those three main features: context, word-collocations, the knowledge of ethnic specifications. Meanings of words and sentences and even whole abstracts are not autonomous, and depend on the general distributions and surroundings.
In the past few years, much attention has been paid on extending phrase-based statistical machine translation with syntactic structures. In this paper we introduce a novel syntax encapsulated phrase(SEP) model, in which treebank tag sequences are employed to decorate the bilingual phrase pairs. We use tag sequences, instead of phrase pairs, to train the lexicalized reordering model. Since the number of treebank tags is much smaller than the number of words, the tag sequence based reordering model is smaller and more accurate than the phrase based reordering model. Experiments were carried out on four types of models: the phrase model, the hierarchical phrase model, the POS tag encapsulated phrase(PTEP) model and the syntactic tag encapsulated phrase(STEP) model. The STEP model obtained higher BLEU-4 score than other models on NIST 2005 MT task.
Learning vocabulary and understanding texts present difficulty for language learners due to, among other things, the high degree of lexical ambiguity. By developing an intelligent tutoring system, this dissertation examines whether automatically providing enriched sense-specific information is effective for vocabulary learning and reading comprehension of second language learners. The system developed in this study contributes to an extended understanding of how NLP techniques can be applied more effectively in an educational environment. The system allows learners to upload texts and click on any content word in order to obtain sense-appropriate lexical information for unfamiliar or unknown words during reading. The system consists of three components: (1) the system manager controls the interaction among each learner, the NLP server, and the lexical database; (2) the NLP server converts a raw input text to a linguistically-analyzed text; (3) the lexical database is used to provide a sense-appropriate definition and example sentences of a word to the learner. To obtain the sense-appropriate information, the system first performs word sense disambiguation (WSD) on the input text. Pointing to appropriate examples tuned for language learners, however, is complicated by the fact that the database of examples is from one repository (COBUILD), while automatic WSD systems generally rely on senses from another (WordNet). The lexical database, then, is indexed by WordNet senses, each of which points to an appropriate corresponding COBUILD sense. The fact that every sense inventory has its own standards of sense distinction poses a serious problem in integrating these inventories into one. To redirect an input WordNet sense to a corresponding COBUILD sense, thus, a word sense alignment algorithm was developed, following a heuristic of favoring flatter alignment structures. With this system, an empirical study was conducted with 60 intermediate learners of English as a second language to examine whether this system can lead learners to improve their vocabulary acquisition and reading comprehension. The findings show that learners demonstrated higher performance when receiving sense-specific information. Furthermore, the qualitative examination of the effect of automatic system errors show that, although learners showed learning regardless of the appropriateness of lexical information, they still showed relatively greater learning when given appropriate lexical information.
We propose HamleDT – HArmonized Multi-LanguagE Dependency Treebank. HamleDT is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. While the license terms prevent us from directly redistributing the corpora, most of them are easily acquirable for research purposes. What we provide instead is the software that normalizes tree structures in the data obtained by the user from their original providers. Keywords:dependency treebank, annotation scheme, harmonization 1.
This article focuses on the relationship between a lexical database and a derivational map, a hierarchical representation of the lexicon that specifies lexical and morphological inheritance by means of graph theory. In order to take steps towards construing a three-dimensional lexicon, this article also puts forward the concept of semantic pole. A semantic pole is a pivot of lexical organization defined as the area of lexical space comprised of the intersection of the lexical areas of one or more derivational paradigms and the major exponents of a semantic prime. This proposal is applied to the semantic pole sōð-trēowe in Old English and two main conclusions are reached. Firstly, a semantic pole constitutes a panchronic representation of lexical relations that contributes to the development of the third-generation Internet, which aims, among other things, at compiling databases and representing contents in 3D. Secondly, the concept of semantic pole constitutes an explanatory principle of derivational morphology and lexical semantics because it explains the degree of convergence between morphological and lexical inheritance, accounts for the clustering of lexical items around certain semantic poles and predicts the rise of polysemy.
International audience
Korean is a morphologically rich language in which grammatical functions are marked by inflections and affixes, and they can indicate grammatical relations such as subject, object, predicate, etc. A Korean sentence could be thought as a sequence of eojeols. An eo- jeol is a word or its variant word form ag- glutinated with grammatical affixes, and eo- jeols are separated by white space as in En- glish written texts. Korean treebanks (Choi et al., 1994; Han et al., 2002; Korean Lan- guage Institute, 2012) use eojeol as their fun- damental unit of analysis, thus representing an eojeol as a prepreterminal phrase inside the constituent tree. This eojeol-based an- notating schema introduces various complex- ity to train the parser, for example an en- tity represented by a sequence of nouns will be annotated as two or more different noun phrases, depending on the number of spaces used. In this paper, we propose methods to transform eojeol-based Korean treebanks into entity-based Korean treebanks. The methods are applied to Sejong treebank, which is the largest constituent treebank in Korean, and the transformed treebank is used to train and test various probabilistic CFG parsers. The experi- mental result shows that the proposed transfor- mation methods reduce ambiguity in the train- ing corpus, increasing the overall F1 score up to about 9 %.
International audience
International audience
We address the issue of consuming heterogeneous annotation data for Chinese word segmentation and part-of-speech tagging. We empirically analyze the diversity between two representative corpora, i.e. Penn Chinese Treebank (CTB) and PKU’s People’s Daily (PPD), on manually mapped data, and show that their linguistic annotations are systematically different and highly compatible. The analysis is further exploited to improve processing accuracy by (1) integrating systems that are respectively trained on heterogeneous annotations to reduce the approximation error, and (2) re-training models with high quality automatically converted data to reduce the estimation error. Evaluation on the CTB and PPD data shows that our novel model achieves a relative error reduction of 11 % over the best reported result in the literature. 1