Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
SUC-CORE is a subset of Stockholm Umea Corpus 2.0 and Swedish Treebank, annotated with noun phrase coreference. While most coreference annotated corpora consist of texts of similar types within rel...
ABSTRACT Emotional content of verbal material affects the speed of visual word recognition in various cognitive tasks, independently of lexicosemantic variables. However, little is known about how the dimensions of emotional arousal and valence interact with the lexicosemantic properties of words such as age of acquisition, familiarity, and imageability, that determine word recognition performance. This study aimed to examine these relationships using English ratings for affective and lexicosemantic features. Eighty-two native English speakers rated 300 words for emotional valence, arousal, familiarity, age of acquisition, and imageability. Although both dimensions of emotion were correlated with lexicosemantic variables, a unique emotion cluster produced the strongest quadratic relationship. This finding suggests that emotion should be included in models of word recognition as it is likely to make an independent contribution.
This chapter presents the Lassy Small and Lassy Large treebanks, as well as related tools and applications. Lassy Small is a corpus of written Dutch texts (1,000,000 words) which has been syntactically annotated with manual verification and correction. Lassy Large is a much larger corpus (over 500,000,000 words) which has been syntactically annotated fully automatically. In addition, various browse and search tools for syntactically annotated corpora have been developed and made available. Their potential for applications in corpus linguistics and information extraction has been illustrated and evaluated in a series of case studies.
Persian with its about 100,000,000 speakers in the world belongs to the group of languages with less developed linguistically annotated resources and tools. The few existing resources and tools are neither open source nor freely available. Thus, our goal is to develop open source resources such as corpora and treebanks, and tools for data-driven linguistic analysis of Persian. We do this by exploring the reusability of existing resources and adapting state-of-the-art methods for the linguistic annotation. We present fully functional tools for text normalization, sentence segmentation, tokenization, part-of-speech tagging, and parsing. As for resources, we describe the Uppsala PErsian Corpus (UPEC) which is a modified version of the Bijankhan corpus with additional sentence segmentation and consistent tokenization modified for more appropriate syntactic annotation. The corpus consists of 2,782,109 tokens and is annotated with parts of speech and morphological features. A treebank is derived from UPEC with an annotation scheme based on Stanford Typed Dependencies and is planned to consist of 10,000 sentences of which 215 have already been annotated.
PURPOSE: To describe the judgement of the concreteness of a set of 162 Brazilian Portuguese words, prior to the elaboration of a speech recognition test, as well as to verify the influence of variables such as the frequency of occurrence of the words and age and undergraduate program year of the participants on the concreteness ratings. METHODS: Fifty undergraduate Speech-Language Pathology and Audiology students from a public university rated the concreteness of a set of 162 words using a seven-point scale where the lowest concreteness degree was represented by number one and the highest by number seven. Participants were free to choose any number in the scale. RESULTS: The results showed a tri-modal distribution of values, suggesting the classification of three categories, according to the concreteness rating. The low concreteness category ranged from 1.76 to 3.45; the medium concreteness category, from 3.46 to 4.95; and the high concreteness rating, from 4.96 to 6.70. Positive correlation was found between the concreteness rating and the coefficient of variation, whereby the higher the rating attributed to a word, the lesser variation in the responses. No significant correlation was found between concreteness ratings and the frequency of occurrence of words. The influence of age and undergraduate year was significant for some correlations. CONCLUSION: Results showed three concreteness categories, and suggest that concreteness can be considered an independent attribute of words, since their frequency of occurrence, as well as participants' age and undergraduate program year did not influence the ratings attributed. The words classified in the high concreteness category were subsequently used for the elaboration of a speech recognition test.
Nama Peneliti/Mahasiswa: RIYAN ADI LESMANA (G64086037I) Judul: Pengembangan Word Net Bahasa Indonesia Berbasis Web Judul (English): Building Web Based Indonesian WordNet Pembimbing/Supervisor: Yani Nurhadriyani Abstrak/Abstract: RIYAN ADI LESMANA. Building Web Based Indonesian WordNet. Under supervision of AHMAD RIDHA and SRI NURDIATI. WordNet is an electronic lexical database that classifies categories of words such as nouns, adjectives, verbs, adverbs, and others in a set called synonim set (synset). WordNet for English has been developed by Princeton University. The purpose of this study is to build an web-based WordNet for Bahasa Indonesia. We used evolutionary method that started by building the parts that were well-understood as the first implementation and continuously developed based on comments from users until an adequate application had been developed. The data sources of the application are Kamus Besar Bahasa Indonesia (KBBI) and Indonesian thesaurus. Our application can display the meaning/definition of a word and display a word’s relations with other words. This application can also be used to add new words, definitions, examples, and word relationships; this function can only be done by users who have registered and been approved for registration by administrator. In this study we also developed an XML document format and XML Schema Definition to assist others in obtaining data from the system. Keywords: Indonesian Bahasa, Lexical Database, Wordnet.
Objectives: The present study aimed at exploring selected language development dimensions in kindergarten children with a migration background taking into consideration the duration of kindergarten attendance and home language. Participants: 60 children with normal intelligence, 3 (n = 13), 4 (n = 24), 5 (n = 23) years of age, mean: 56.1 (SD 8.7) months. Methods: Retrospective analysis of a data set. Instruments: The subtests “Phonological Non Word Repetition” and “Sentence Comprehension” from the German SETK 3-5 (Grimm, 2008); “Auditorial Sequential Memory for Digits”, “Doll Play”, “Word Explanation” from the German WET (Kastner-Koller & Deimann, 2002); and the German questionnaire SISMIK (Ulich & Mayr, 2003) for the assessment of language behaviour via the kindergarten educator. Results: The mean performance on phonological memory for nonword repetition (T-score 48.4; SD 12.0) and for digits (Centile 4.9; SD 1.9) was found normal. Lexical knowledge and usage tended towards the lower norm range (mean Centile 3.5; SD 2.1); the individual result was in 20% of the children ≪ 2 SD below the mean age norm. Lexical performance was the only one which was significantly higher in children with a kindergarten attendance > 1 year. The understanding of sentence comprehension was on average in the lower norm range (T-score 42.2; SD 11.2; Doll Play: Centile 2.8; SD 2.3); the individual result fell in 10% resp. 33% of the children below the criterion ≪ 2 SD below mean age norm. Language behaviour in contact with children (T-score 56.6; SD 10.4), in contact with kindergarten educators (T-score 56.4; SD 11.5), and linguistic competence (T-score 56.9; SD 10.9) were rated age-appropriate (SISMIK). No gender differences were observed. Children communicating bilingually within the family showed higher language competencies. Conclusion: Despite performance in the norm range, on average, still 33% of immigrant children display deficiencies both in syntactical and semantic capacities in the German language.
The paper deals with the issues of knowledge processing and transmission via metaphorical representation. It should be noted that within the cognitive paradigm in modern linguistics metaphor is considered to be a universal mental mechanism that uses previously acquired knowledge. Through preliminary theoretical literature review one can observe a trend to regards metaphors as ubiquitous components of any discourse. Due to their omnipresence metaphors are assumed to be indeliberate, subjective and therefore ruleless. The development of the conceptual metaphor theory and implementation of metaphor modelling (Lakoff, Johnson 1980) contributed greatly to the reevoluation of their (metaphors) state-of-the-art. We assume that it is of potentially high significance to reveal regularities in metaphorization using metaphor modelling in different discourse types. There is good reason to believe that the most effective method of metaphorization sudy is the composition of a metaphors thesaurus. The language material in this thesaurus is organized on the basis of two sense complexes (denotative descriptor and significative descriptor, or metaphor model), which build the lexical meaning. The metaphor model is determined as a conceptual domain (a source domain), which contains elements connected by different semantic relations (those of function, cause, example etc.). The name of the basic concept connecting all the elements of the taxons becomes the title of the whole metaphor model. Since the metaphor model results from the non-professional categorization but not from the scientific one, the distinction of metaphor models was made using the definitions of dictionaries. The basic category proves to be metaphor schemata of the discourse, that is the set of basic metaphor models including all metaphors revealed in the particular discourse. The dominating metaphor models comprise the central part of metaphor schemata of the discourse, the less representative metaphor models - its peripheral part. Metaphor schemata of the discourse is to give uniform treatment to the many different metaphor models so that different types of discourse can be contrasted and compared. Discourse is defined as verbally mediated human action represented by purpose-built text corpus (Mishlanova S. 2004), that results in creating a concept, that is specific knowledge of differen levels of abstraction, from naive, empirical to scientific, theoretical level. The concept has a vast array of verbal representation that permit a great number of differentiations. According to Ju. Karaulov they fall into three forms: "language-system", "language-text", "languagecompetence" (Karaulov Ju. 1999). The aim of this study is metaphor modelling of two ways of verbal representation of the concept (texts and associative field) in Russian and English medical discourses. Material. The volume of the thesaurus consists of 2,500 examples of metaphors chosen by the method of entire selection from scientific medical texts in the Russian and English language; 645 examples of metaphors from medical blogs in the Russian and English language; 328 responses from 164 participants (72 Russian and 92 American participants).Results.The scientific medical discourse showed the shift of the metaphor schemata nucleus to the right (Table 1).Table 1. Metaphor schemata in the scientific medical discourse in the Russian, English and German language compared (%) Scientific medical discourse Human Being Animate Nature Inanimate Nature Social Subject Russian 15* 10 15* 60** English 7 10 15* 68** ** - dominating metaphor model * - second most representative metaphor model In the scientific medical discourse Social Subject (60% in Russian, 68% in English) is the dominating metaphor model; Inanimate Nature (15% in Russian and English) is the second most representative metaphor model; Human Being and Animate Nature are the less representative metaphor models (Mishlanova S. 2004). The medical blog analysis showed parallelism of the metaphor schemata both in the Russian and English language (Table 2).Table 2. Metaphor schemata in the Russian and English medical blogs compared (%) Popular medical discourse Human Being Animate Nature Inanimate Nature Social Subject Russian 18* 1 6 75** English 15* 2 5 78** ** - dominating metaphor model * - second most representative metaphor model The dominating metaphor model Social Subject both in the Russian and English language (75% in Russian, 78% in English) is shifted to the right while the second most representative metaphor model Human Being is shifted to the left (18% in Russian, 15% in English). To study the representation of the basic medical concepts health and disease 164 participants were asked to complete the sentences: Health is similar to..., because...; Disease is similar to..., because.... The associative field was structured in the form of the metaphor schemata (Table3).Table 3. Metaphor schemata of associative field in Russian and American participants compared (%) Popular medical discourse Human Being Animate Nature Inanimate Nature Social Subject Russian participants 34** 28* 20 20 American participants 56** 5 11 28* ** - dominating metaphor model * - second most representative metaphor model In this case the popular medical discourse showed the shift of the metaphor schemata nucleus of the Russian participants to the left. The dominating metaphor model Human Being comprises 34% in the Russian participants and 56% in American participants whereas the second most representative metaphor model differs in the Russian participants (Animate Nature - 28%) and American participants (Social Subject - 28%), the latter taking up the very right position (Mishlanova S., Polyakova S. 2009). Conclusion. Different types of the medical discourse were studied. The comparative study was undertaken to analyse various ways of verbal representation of the concept - texts and associative field ("language-competence") in the Russian and English language. The research showed common characteristics of metaphorization in texts (in our study - scientific discourse and medical blogs) whereas metaphor modelling of the associative field revealed cultural differences.
In this paper, we propose a scheme for anaphora annotation in Hindi Dependency Treebank. The goal is to identify and handle the challenges that arise in the annotation of reference relations in Hindi. We identify some of the issues related to anaphora annotation specific to Hindi such as distribution of markable span, sequential annotation, representation format, annotation of multiple referents etc. The scheme hence incorporates some characteristics specific to these issues in order to achieve a consistent annotation. Most significant among these characteristics is the head-modifier separation in referent selection. The modifier-modified dependency relations inside a markable is utilized for this headmodifier distinction. A part of the Hindi Dependency Treebank, of around 2500 sentences has been annotated with anaphoric relations and an inter-annotator study was carried out which shows a significant agreement over selection of the head referent using the proposed scheme as compared to MUC annotation format. The current annotation is done for a limited set of pronominal categories.
The central problems that this paper addresses are (i) the lack of large and rich formalised lexicons for multi-word expressions for use in Natural Language Processing (NLP); (ii) the lack of proper methods and tools to extend the lexicon of an NLP-system for multi-word expressions given a text corpus in a maximally automated manner. The paper describes innovative methods and tools for the automatic identification and lexical representation of multi-word expressions. In addition, it describes a 5.000 entry corpus-based multi-word expression lexical database for Dutch developed using these methods. The database has been externally validated, and its usability has been evaluated in NLP-systems for Dutch. The MWE database developed fills a gap in existing lexical resources for Dutch. The generic methods and tools for MWE identification and lexical representation focus on Dutch, but they are largely language-independent and can also be used for other languages, new domains, and beyond this project. The research results and data described in this paper contribute directly to strengthening the digital infrastructure for Dutch.
A morphological analyser only recognizes words that it already knows in the lexical database. It needs, however, a way of sensing significant changes in the language in the form of newly borrowed or coined words with high frequency. We develop a finite-state morphological guesser in a pipelined methodology for extracting unknown words, lemmatizing them, and giving them a priority weight for inclusion in a lexicon. The processing is performed on a large contemporary corpus of 1,089,111,204 words and passed through a machine-learning-based annotation tool. Our method is tested on a manually-annotated gold standard of 1,310 forms and yields good results despite the complexity of the task. Our work shows the usability of a highly non-deterministic finite state guesser in a practical and complex application. 1
Syntactic analysis is a valuable addition to corpus annotation, since it enables the retrieval of structural information which is otherwise difficult to access. The Norwegian Infrastructure for the Exploration of Syntax and Semantics is adding syntactic information to the Norwegian Newspaper Corpus and is effectively producing a treebank as a parsed corpus.
In this paper we describe semi-automatical extending of the Czech WordNet lexical database (48,000 literals in 28,000 synsets) by translation of English literals from existing synsets in Princeton WordNet. We make use of a machine-readable bilingual dictionary to extract English-Czech translation pairs, search the English literals in Princeton WordNet and in case of a high-confidence match we transfer the literal into Czech WordNet. Along with literals, new synsets parallel to the English ones and identified by ILI are introduced into CzechWordNet, including information on their ILR (Internal Language Relations) such as hypernymy/hyponymy. The paper describes the parsing of the dictionary data, extraction of translation pairs and the criteria used for estimating the confidence level of a match. Results of the work are 36,228 added literals and 12,403 created synsets. An overview of previous similar attempts for other languages is also included.
The annotation of large corpora is usually restricted to syntactic structure and word class. Pure lexical information and information on the structure of words are stored in specialized dictionaries (Baayen et al., 1995). Both data structures ‐ dictionary and text corpus ‐ can be matched to get e.g. a distribution of certain (restricted) lexical information from a text. This procedure works fine for synchronic corpora. What is missing, however, is either a special mark-up in texts linking each of the items to a certain time or a diachronic lexical database that allows for the matching of the items over time. In what follows, we take the latter approach and present a tool set (MoreXtractor, Morphilizer, MorQuery), a database (Morphilo-DB) and the architecture of a platform (Morphorm) for a sustainable use of diachronic linguistic data for Middle English, Early Modern English and Modern English.
Annotation of discourse relations is a project related to the Prague Dependency Treebank 2.5. It represents a new manually annotated layer of language description, above the existing layers of the PDT, and it portrays linguistic phenomena from the perspective of discourse structure and coherence.
Texts The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part. Data The English part contains the entire Penn Treebank - Wall Street Journal Section (LDC99T42). The Czech part consists of Czech translations of all of the Penn Treebank-WSJ texts. The corpus is 1:1 sentence-aligned. An additional automatic alignment on the node level (different for each annotation layer) is part of this release, too. The original Penn Treebank-like file structure (25 sections, each containing up to one hundred files) has been preserved. Only those PTB documents which have both POS and structural annotation (total of 2312 documents) have been translated to Czech and made part of this release. Each language part is enhanced with a comprehensive manual linguistic annotation in the PDT 2.0 style (LDC2006T01, Prague Dependency Treebank 2.0). The main features of this annotation style are: dependency structure of the content words and coordinating and similar structures (function words are attached as their attribute values) semantic labeling of content words and types of coordinating structures argument structure, including an argument structure ("valency") lexicon for both languages ellipsis and anaphora resolution. This annotation style is called tectogrammatical annotation and it constitutes the tectogrammatical layer in the corpus. For more details see below and documentation. Annotation of the Czech part Sentences of the Czech translation were automatically morphologically annotated and parsed into surface-syntax dependency trees in the PDT 2.0 annotation style. This annotation style is sometimes called analytical annotation; it constitutes the analytical layer of the corpus. The manual tectogrammatical (deep-syntax) annotation was built as a separate layer above the automatic analytical (surface-syntax) parse. A sample of 2,000 sentences was manually annotated on the analytical layer. Annotation of the English part The resulting manual tectogrammatical annotation was built above an automatic transformation of the original phrase-structure annotation of the Penn Treebank into surface dependency (analytical) representations, using the following additional linguistic information from other sources: PropBank (LDC2004T14) VerbNet NomBank (LDC2008T23) flat noun phrase structures (by courtesy of D. Vadas and J.R. Curran) For each sentence, the original Penn Treebank phrase structure trees are preserved in this corpus together with their links to the analytical and tectogrammatical annotation.
The program part of the lexical database was developed as the Mozilla Firefox application (2005-2006); the processing software is designed as a modern lexicographic workstation. Since 2007, we have gradually complemented the database with the lexicographicallly relevant data (research project Creation of a Lexical Database of the Czech Language of the Beginning of the 21st Century, 2005-2011). We have focused predominantly on detailed treatment of the example part, which demonstrates the breadth of the collocability of the lemmas.
This article focuses on the relationship between a lexical database and a derivational map, a hierarchical representation of the lexicon that specifies lexical and morphological inheritance by means of graph theory. In order to take steps towards construing a three-dimensional lexicon, this article also puts forward the concept of semantic pole. A semantic pole is a pivot of lexical organization defined as the area of lexical space comprised of the intersection of the lexical areas of one or more derivational paradigms and the major exponents of a semantic prime. This proposal is applied to the semantic pole sōð-trēowe in Old English and two main conclusions are reached. Firstly, a semantic pole constitutes a panchronic representation of lexical relations that contributes to the development of the third-generation Internet, which aims, among other things, at compiling databases and representing contents in 3D. Secondly, the concept of semantic pole constitutes an explanatory principle of derivational morphology and lexical semantics because it explains the degree of convergence between morphological and lexical inheritance, accounts for the clustering of lexical items around certain semantic poles and predicts the rise of polysemy.
International audience
Korean is a morphologically rich language in which grammatical functions are marked by inflections and affixes, and they can indicate grammatical relations such as subject, object, predicate, etc. A Korean sentence could be thought as a sequence of eojeols. An eo- jeol is a word or its variant word form ag- glutinated with grammatical affixes, and eo- jeols are separated by white space as in En- glish written texts. Korean treebanks (Choi et al., 1994; Han et al., 2002; Korean Lan- guage Institute, 2012) use eojeol as their fun- damental unit of analysis, thus representing an eojeol as a prepreterminal phrase inside the constituent tree. This eojeol-based an- notating schema introduces various complex- ity to train the parser, for example an en- tity represented by a sequence of nouns will be annotated as two or more different noun phrases, depending on the number of spaces used. In this paper, we propose methods to transform eojeol-based Korean treebanks into entity-based Korean treebanks. The methods are applied to Sejong treebank, which is the largest constituent treebank in Korean, and the transformed treebank is used to train and test various probabilistic CFG parsers. The experi- mental result shows that the proposed transfor- mation methods reduce ambiguity in the train- ing corpus, increasing the overall F1 score up to about 9 %.
This paper describes how electronic grammars can be further enhanced by adding machine-readable grammars and treebanks. We explore the potential benefits of im- plemented grammars and treebanks for descriptive linguistics, following the discursive methodology of Bird & Simons (2003) and the values and maxims identified by Nordhoff(2008). We describe the resources which we believe make implemented grammars and treebanks feasible additions to electronic descriptive grammars, with a particular focus on the Grammar Matrix grammar customization system (Bender et al. 2010) and the Fangorn treebank search application (Ghodke & Bird 2010). By presenting an ex- ample of an implemented grammar based on a descriptive prose grammar, we show one productive method of collaboration between grammar engineer and field linguist, and propose that a tighter integration could be beneficial to both, creating a virtuous cycle that could lead to more effective and informative resources.
In this paper, we describe an ongoing research to develop an HPSG-based treebank for Persian. To this aim, we use a bootstrapping approach for the data annotation. In the first step, a set of seed rules are defined as regular expressions in the CLaRK system. Then, the data is shallow processed with this set of rules. In the next step, a human annotator completes the annotation of sentences manually. To increase automatic annotation, we extract the manual applied rules and iteratively augment the seed rules with the rules applied frequently in the manual annotation. Our experiment in building the Persian treebank which currently contains 1000 sentences shows that the proposed method reduces human intervention from 74.05% in first iterations to 39.01% in last iterations.
International audience
International audience
In spite of their superior performance, neural probabilistic language models (NPLMs) remain far less widely used than n-gram models due to their notoriously long training times, which are measured in weeks even for moderately-sized datasets. Training NPLMs is computationally expensive because they are explicitly normalized, which leads to having to consider all words in the vocabulary when computing the log-likelihood gradients. We propose a fast and simple algorithm for training NPLMs based on noise-contrastive estimation, a newly introduced procedure for estimating unnormalized continuous distributions. We investigate the behaviour of the algorithm on the Penn Treebank corpus and show that it reduces the training times by more than an order of magnitude without affecting the quality of the resulting models. The algorithm is also more efficient and much more stable than importance sampling because it requires far fewer noise samples to perform well. We demonstrate the scalability of the proposed approach by training several neural language models on a 47M-word corpus with a 80K-word vocabulary, obtaining state-of-the-art results on the Microsoft Research Sentence Completion Challenge dataset.
It has been established that incorporating word cluster features derived from large unlabeled corpora can significantly improve prediction of linguistic structure. While previous work has focused primarily on English, we extend these results to other languages along two dimensions. First, we show that these results hold true for a number of languages across families. Second, and more interestingly, we provide an algorithm for inducing cross-lingual clusters and we show that features derived from these clusters significantly improve the accuracy of cross-lingual structure prediction. Specifically, we show that by augmenting direct-transfer systems with cross-lingual cluster features, the relative error of delexicalized dependency parsers, trained on English treebanks and transferred to foreign languages, can be reduced by up to 13%. When applying the same method to direct transfer of named-entity recognizers, we observe relative improvements of up to 26%. 1
We present two dependency parsers for Persian, MaltParser and MSTParser, trained on theUppsala PErsian Dependency Treebank. The treebank consists of 1,000 sentences today. Itsannotation scheme is based on Stanford Typed Dependencies (STD) extended for Persianwith regard to object marking and light verb contructions. The parsers and the treebank aredeveloped simultanously in a bootstrapping scenario. We evaluate the parsers by experimentingwith different feature settings. Parser accuracy is also evaluated on automatically generated andgold standard morphological features. Best parser performance is obtained when MaltParseris trained and optimized on 18,000 tokens, achieving 68.68% labeled and 74.81% unlabeledattachment scores, compared to 63.60% and 71.08% for labeled and unlabeled attachmentscore respectively by optimizing MSTParser.
Chinese Treebank as a kind of valuable resource in Chinese Information Processing includes rich information on the sentence structure and the constitunent combination.A study on the combination of POS strings is the basic work for the effective use of treebank information.This paper investigates the ambiguous combination in Chinese Treebank,revealing that it largely requires semantic feature to resolve the ambiguous combination and structure in Chinese,and can not be solved simply by grammatical features(such as POS information) of words.
We address the issue of consuming heterogeneous annotation data for Chinese word segmentation and part-of-speech tagging. We empirically analyze the diversity between two representative corpora, i.e. Penn Chinese Treebank (CTB) and PKU’s People’s Daily (PPD), on manually mapped data, and show that their linguistic annotations are systematically different and highly compatible. The analysis is further exploited to improve processing accuracy by (1) integrating systems that are respectively trained on heterogeneous annotations to reduce the approximation error, and (2) re-training models with high quality automatically converted data to reduce the estimation error. Evaluation on the CTB and PPD data shows that our novel model achieves a relative error reduction of 11 % over the best reported result in the literature. 1
Parallel treebanks have received increasing attention in the past few years,\nprimarily due to their potential use in statistical machine translation. Creating\nparallel treebanks manually is a time-consuming and expensive task\nand for this reason there is considerable interest in creating treebanks automatically.\nThis task can be solved using standard tools such as parsers and\naligners. However, because parallel treebanks are based on parallel corpora,\nwe are in a special situation where the same meaning is represented\nin two different ways. This thesis is about how we can exploit this information\nto create better parallel treebanks than we can by using standard\ntools....
Training data selection is a common method for domain adaptation, the goal of which is to choose a subset of training data that works well for a given test set. It has been shown to be effective for tasks such as machine translation and parsing. In this paper, we propose several entropy-based measures for training data selection and test their effectiveness on two tasks: Chinese word segmentation and part-of-speech tagging. The experimental results on the Chinese Penn Treebank indicate that some of the measures provide a statistically significant improvement over random selection for both tasks.
This paper presents a new method to evaluate machine translation (MT) systems against a parallel treebank. This approach examines specific linguistic phenomena rather than the overall performance of the system. We show that the evaluation accuracy can be increased by using word alignments extracted from a parallel treebank. We compare the performance of our statistical MT system with two other competitive systems with respect to a set of problematic linguistic structures for translation between German and French.