Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
We evaluate the accuracy of an unlexicalized statistical parser, trained on 4K treebanked sentences from balanced data and tested on the PARC DepBank. We demonstrate that a parser which is competitive in accuracy (without sacrificing processing speed) can be quickly tuned without reliance on large in-domain manually-constructed treebanks. This makes it more practical to use statistical parsers in applications that need access to aspects of predicate-argument structure. The comparison of systems using DepBank is not straightforward, so we extend and validate DepBank and highlight a number of representation and scoring issues for relational evaluation schemes.
The Czech Academic Corpus was created during the 1970s and 1980s at the Czech Lan- guage Institute under the supervision of Marie Těsitelova. The main motivation to build it (a total of 540 thousand word tokens) was to obtain the quantitative characteristics of contemporary Czech. The corpus is structurally annotated on two levels - the morphological level and the syntactical-ana- lytical level. The original stochastic experiments in morphological tagging of Czech were performed using the corpus at the beginning of the 1990s. Given this, the corpus-based processing of Czech was launched. At the end of 1990s, work on the Prague Dependency Treebank had started (independently from the corpus) and its first edition was published in 2001. In considering future released versions of the treebank, we have decided to convert the corpus into the treebank-like format. This article focuses on the twenty-year history of the Czech Academic Corpus. Special attention is devoted to thus far un- published facts about the corpus annotation. The conversion steps resulting in the first version of the Czech Academic Corpus are described in detail.
The years 1940–2000 have been a period of uneven growth and development for Irish-language prose and drama. The most significant development of all during these decades was the emergence of a great diversity of literary forms as many writers gradually distanced themselves from the more restricting aspects of Revivalist ideology and began to experiment more with language, structure and subject matter. Critical perspectives also evolved and it was recognised that the cultural context in which the Irish language had survived now demanded expression in a manner which often defied the strictures imposed by genre or particular aesthetic paradigms. Linguistic norms are challenged in many works, both by Gaeltacht and non-Gaeltacht writers, and questions of cultural change and cultural hybridity become stylistic concerns as much as themes in much of the prose literature of the period. During this time Máirtín Ó Cadhain was to gain prominence as a major prose writer, and a host of other new and distinct voices emerged such as those of novelists and short story writers Diarmaid Ó Súilleabháin, Alan Titley, Micheál Ó Conghaile and Pádraig Ó Siadhail. These years also witnessed a continuation of a strong Gaeltacht tradition of autobiography and auto-ethnography, side by side with an unprecedented amount of experimental prose writing. The publishing house Sáirséal agus Dill, which was established in 1946, was responsible for the publication of much of the most original and innovative fiction of the period, while in the 1980s the newly established Coiscéim and Cló Iar-Chonnachta were to play a prominent role in the encouragement of many new and younger writers.
We identify problems with the Penn Treebank that render it imperfect for syntaxbased machine translation and propose methods of relabeling the syntax trees to improve translation quality. We develop a system incorporating a handful of relabeling strategies that yields a statistically significant improvement of 2.3 BLEU points over a baseline syntax-based system.
The last decade has seen a large increase in the number of available corpus query systems. Some of these are optimized for a particular kind of linguistic annotation (e.g., time-aligned, treebank, word-oriented, etc.). In this paper, we report on our own corpus query system, called Emdros. Emdros is very generic, and can be applied to almost any kind of linguistic annotation using almost any linguistic theory. We describe Emdros and its query language, showing some of the benefits that linguists can derive from using Emdros for their corpora. We then describe the underlying database model of Emdros, and show how two corpora can be imported into the system. One of the two is a parallel corpus of Hungarian and English (the Hunglish corpus), while the other is a treebank of German (the TIGER Corpus). In order to evaluate the performance of Emdros, we then run some performance tests. It is shown that Emdros has extremely good performance on “small ” corpora (less than 1 million words), and that it scales well to corpora of many millions of words. 1.
Objective:To investigate the difference of affective reactions to International Affective Picture System (IAPS) between Chinese and American adults with different culture background. Methods:By using Self Assessment Manikin (SAM), every subject had to complete the Valence and Arousal rating for all the color pictures in IAPS, and the scores obtained were compared between two groups with different cultural background. Results:Among 816 pictures, Valence scores between two groups of male showed significant difference in 70.22%(573/816) pictures (P0.05), and Arousal scores between the two groups of male showed significant difference in 69.61%(567/813) pictures(P0.05). Among 816 pictures, Valence scores between two groups of female showed significant difference in 75.00%(612/816) pictures (P 0.05), and Arousal scores between two groups of female showed significant difference in 63.24%(516/816) pictures (P 0.05). From 9 categories of IAPS, the ratio of positive and negative affective reactions due to facial pictures in two groups of male was greatly different(χ2 =4.857,P0.05), and the ratio of positive and negative affective reactions due to erotic pictures in two groups of female was greatly different (c2=25.93, P0.001). The ratios of low and high arousal due to categories of facial and object pictures were significantly different, and the ratios of low and high arousal due to categories of facial, violence, objects, sports and pollution were significantly different (P0.05). Conclusion:Under different cultural background, the affective reactions of Chinese and American adults towards IAPS are almost similar, but significantly different in the erotic and facial pictures.
This paper presents a general method to automatically build large knowledge bases from online lexical resources. While our experiments were limited to generate a knowledge base from WordNet, an online lexical database, the method is applicable to any type of dictionary organized around the elementary structure lexical entry - denition(s). The advantages of using WordNet, or richer online resources such as thesauri, as the source of a knowledge base, are outlined.
Face mask is now a common feature in our social environment. Although face covering reduces our ability to recognize other's face identity and facial expressions, little is known about its impact on the formation of first impressions from faces. In two online experiments, we presented unfamiliar faces displaying neutral expressions with and without face masks, and participants rated the perceived approachableness, trustworthiness, attractiveness, and dominance from each face on a 9-point scale. Their anxiety levels were measured by the State-Trait Anxiety Inventory and Social Interaction Anxiety Scale. In comparison with mask-off condition, wearing face masks (mask-on) significantly increased the perceived approachableness and trustworthiness ratings, but showed little impact on increasing attractiveness or decreasing dominance ratings. Furthermore, both trait and state anxiety scores were negatively correlated with approachableness and trustworthiness ratings in both mask-off and mask-on conditions. Social anxiety scores, on the other hand, were negatively correlated with approachableness but not with trustworthiness ratings. It seems that the presence of a face mask can alter our first impressions of strangers. Although the ratings for approachableness, trustworthiness, attractiveness, and dominance were positively correlated, they appeared to be distinct constructs that were differentially influenced by face coverings and participants' anxiety types and levels.
The paper presents a spoken document summarization scheme using a topic-related corpus and semantic dependency grammar. The summarization score considers speech recognition confidence, word significance, word trigram, semantic dependency grammar (SDG) and probabilistic context free grammar (PCFG). In addition, a topic-related corpus consisting of keywords as well as articles is used to estimate the word significance score using latent semantic indexing (LSI). Semantic relations between words are determined by SDG using HowNet and Sinica Treebank. A dynamic programming algorithm is applied to decide the summarization ratio and look for the best summarization result according to summarization scores. Experimental results indicate that the proposed approach effectively extracts important words with semantic dependency and gives a promising speech summary.
Abstract. The main goal of this study was to estimate the correlation between various psychophysiological variables and self-reported disgust during a picture perception paradigm. We further studied disgust sensitivity (DS) as a possible moderator variable for this relationship. Forty-seven subjects (23 females) were presented with a total of 36 pictures with different disgust intensities. Each picture was shown for 8s during which different physiological parameters were registered: heart rate (HR), skin conductance response (SCR), and electromyographic activity (EMG) of the musculus levator labii. Affective ratings and viewing times were assessed after the physiological registrations. The data were analyzed using hierarchical linear models. The degree of the disgust experience reported by the subjects showed a significantly negative correlation with HR and a significantly positive correlation with SCR. Disgust-inducing pictures resulted in higher EMG responses in comparison to neutral pictures, but there was no significant correlation between self-reported disgust and EMG activity on an individual level. Elevated DS, measured by the questionnaire by Haidt, McCauley, and Rozin, (1994), led to more intense subjective responses toward disgust-inducing pictures, but this was not true for the behavioral and physiological responses.
This paper presents the results of automatically inducing a Combinatory Categorial Grammar (CCG) lexicon from a Turkish dependency treebank. The fact that Turkish is an agglutinating free wordorder language presents a challenge for language theories. We explored possible ways to obtain a compact lexicon, consistent with CCG principles, from a treebank which is an order of magnitude smaller than Penn WSJ.
Parallel treebanks, i.e. syntactically annotated corpora of translated texts, are invaluable resources for cross-linguistic research. Whereas some parallel treebanks have been created, little work has been devoted to parallel treebank query systems. In this paper, we compare experiences from monolingual corpus query tools and project these insights for the development of parallel treebank search tools. We distinguish between two different query types, namely single constraint queries and combined constraint queries, and show how the certainty of the alignment information can be included in the search result as well. Suggestions for graphical output representation are also made. We show that a large amount of the work which has been done on monolingual treebanking can be used with parallel treebanks as well, although additional requirements need to be observed and fulfilled. Parallele Baumbanken, d.h. syntaktisch annotierte Korpora aus übersetzten Texten, sind wertvolle Ressourcen für die sprachübergreifende Forschung. Obwohl bereits einige parallele Baumbanken erstellt wurden, findet man noch fast keine Arbeiten zum Thema der dazugehörigen Suchwerkzeuge. In diesem Artikel vergleichen wir Erfahrungen, die mit einsprachigen Korpusabfragesystemen gemacht wurden, und übertragen diese Erkenntnisse auf die Entwicklung eines Abfragesystems für parallele Baumbanken. Wir unterscheiden zwei verschiedene Abfragearten, nämlich Abfragen mit einfachen Baumbedingungen und Abfragen mit kombinierten Baumbedingungen. Wir veranschaulichen, wie Angaben zur Sicherheit der Alinierung in die Suchresultate miteinbezogen werden können und machen einen Vorschlag zur graphischen Abbildung der Suchresultate. Wir zeigen auf, dass ein grosser Teil der Erfahrungen mit einsprachigen Baumbanken für parallele Baumbanken weiterverwendet werden kann, betonen aber, dass zusätzliche Anforderungen beachtet und realisiert werden müssen.
An automatic method for annotating the Penn-II Treebank (Marcus et al., 1994) with high-level Lexical Functional Grammar (Kaplan and Bresnan, 1982; Bresnan, 2001; Dalrymple, 2001) f-structure representations is presented by Burke et al. (2004b). The annotation algorithm is the basis for the automatic acquisition of wide-coverage and robust probabilistic approximations of LFG grammars (Cahill et al., 2004) and for the induction of subcategorisation frames (O’Donovan et al., 2004; O’Donovan et al., 2005). Annotation quality is, therefore, extremely important and to date has been measured against the DCU 105 and the PARC 700 Dependency Bank (King et al., 2003). The annotation algorithm achieves f-scores of 96.73% for complete f-structures and 94.28% for preds-only f-structures against the DCU 105 and 87.07% against the PARC 700 using the feature set of Kaplan et al. (2004). Burke et al. (2004a) provides detailed analysis of these results. \nThis paper presents an evaluation of the annotation algorithm against PropBank (Kingsbury and Palmer, \n2002). PropBank identifies the semantic arguments of each predicate in the Penn-II treebank and annotates their semantic roles. As PropBank was developed independently of any grammar formalism it provides a platform for making more meaningful comparisons between parsing technologies than was previously possible. PropBank also allows a much larger scale evaluation than the smaller DCU 105 and PARC 700 gold standards. In order to perform the evaluation, first, we automatically converted the PropBank annotations \ninto a dependency format. Second, we developed conversion software to produce PropBank-style semantic annotations in dependency format from the f-structures automatically acquired by the annotation algorithm from Penn-II. The evaluation was performed using the evaluation software of Crouch et al. (2002) and Riezler et al. (2002). Using the Penn-II Wall Street Journal Section 24 as the development set, currently we achieve an f-score of 76.58% against PropBank for the Section 23 test set.
Toempower thegeneral massthrough access toinformation andknowledge, organized efforts arebeing madetodevelop relevant content inlocallanguages andprovide local language capabilities toutility software. Wehavedeveloped a Question Answering (QA)System forHindidocuments that wouldberelevant formassesusingHindiasprimary language ofeducation. Theusershould beabletoaccess information fromE-learning documents ina userfriendly way,that isbyquestioning thesystem intheir native language Hindi andthesystem will return theintended answer (also in Hindi) bysearching incontext fromtherepository ofHindi documents. Thelanguage constructs, querystructure, commonwords, etc.arecompletely different inHindias compared toEnglish. A novelstrategy, inaddition to conventional search andNLP techniques, wasusedto construct theHindi QAsystem. Thefocus isoncontext based retrieval ofinformation. Forthis purpose weimplemented a Hindi search engine that works onlocality-based similarity heuristics toretrieve relevant passages fromthecollection. It alsoincorporates language analysis modules like stemmer andmorphological analyzer aswellasself constructed lexical database ofsynonyms. Theexperimental results over corpus oftwoimportant domains ofagriculture andscience showeffectiveness ofourapproach.
We introduce a method for transferring annotation from a syntactically annotated corpus in a source language to a target language. Our approach assumes only that an (unannotated) text corpus exists for the target language, and does not require that the parameters of the mapping between the two languages are known. We outline a general probabilistic approach based on Data Augmentation, discuss the algorithmic challenges, and present a novel algorithm for sampling from a posterior distribution over trees.
@conference{ai-giguet-2005-1, author = {Giguet, Emmanuel and Luquet, Pierre-Sylvain}, title = {Multilingual Lexical Database Generation from parallel texts with endogenous resources}, booktitle = {PAPILLON-2005 Workshop on Multilingual Lexical Databases}, year = {2005}, month = {December 12-14}, address = {Chiang Rai, Thaïland} }
This paper presents the design and construction of the PolyU Treebank, a manually annotated Chinese shallow treebank. The PolyU Treebank is based on shallow annotation where only partial syntactical structures within sentences are annotated. Guided by the Phrase-Standard Grammar proposed by Peking University, the PolyU Treebank has been designed and constructed to provide a large amount of annotated data containing shallow syntactical information and limited semantic information for use in natural language processing (NLP) research. This paper describes the relevant design principles, annotation guidelines, and implementation issues, including the achievement of high quality annotation through the use of well-designed annotation workflow and effective post-annotation checking tools. Currently, the PolyU Treebank consists of a one-million-word annotated corpus and has been used in a number of NLP research projects with promising results.
Abstract. The present paper focuses on representation of morphological meanings on the underlying syntactic level. The concept of semantic counterparts of morphological meanings, the so-called grammatemes, was introduced in Functional Generative Description in the 1960’s. We suggest an elaborated system of these grammatemes, which have become a part of the tectogrammatical level of the Prague Dependency Treebank.
Current trends in language technology require treebanks that do not stop at the level of constituent structure, but include deeper and richer levels of analysis, including appropriate meaning structures. Capturing sufficient detail at different levels of linguistic description is too complex a task to be practically achievable by manual annotation or shallow parsing; rather it requires sophisticated tools that help secure the consistency of parallel but different structures. We are constructing a multilevel treebanking tool that incorporates a deep parser and grammar for Norwegian. Thus, we are tightly linking our treebank to grammar development so as to achieve a sound embedding in grammatical theory and yield more useful results for applications.
Reviewed by: Word sense disambiguation: The case for combinations of knowledge sources by Mark Stevenson Cornelia Tschichold Word sense disambiguation: The case for combinations of knowledge sources. By Mark Stevenson. (CSLI studies in computational linguistics.) Stanford: CSLI Publications, 2003. Pp. 175. ISBN 1575863901. $25. Disambiguating words is easy for human beings, but difficult for computers. Computational linguistics has developed methods to reliably find the correct part of speech for the large majority of words in running text, but the disambiguation of polysemous words and homonyms (bat as animal, sports tool, or blink of the eye) is a more complex process. This difference is due mainly to the lack of sufficiently complete and formalized data about word senses. Stevenson shows how progress can be achieved by reusing existing lexical databases and combining them in an optimal way. Ch. 1 introduces the problem of polysemy and points out the potential areas of application for word sense disambiguation (WSD). Ch. 2 gives some historical background on the area, intended for readers unfamiliar with the field. In Ch. 3, lexicographic problems associated with polysemous words and attempts at arriving at suitable databases (such as Word-Net) are discussed. As it does not seem likely that machines can take over any significant part of the lexicographic work involved in the production of semantic databases, the re-use of machine-readable dictionaries appears to be the only viable solution for the immediate future. S refutes a number of criticisms that have been made against the use of a machine-readable dictionary for WSD, mainly due to their lack of alternatives, and proposes methods for at least partially remedying the known shortcomings. Ch. 4 describes the knowledge sources that can be used for WSD, that is, syntactic, semantic, and pragmatic information, and the conditions needed to combine them. WordNet and the Longman dictionary of contemporary English (LDOCE) are identified as two potentially useful on-line lexicographic databases. In Ch. 5, the computational similarities and differences of part-of-speech tagging and WSD are explained. In Ch. 6, S explains how his system combining the various knowledge sources was implemented: the preprocessing stage filters out proper names, tokenizes the input text, and identifies the part of speech for each word (using a Brill-type tagger). This is followed by a shallow syntactic analysis and finally the lexical look-up stage. At the disambiguation stage, the part-of-speech tags are used to filter out any (syntactically) incompatible senses, before a number of partial (semantic) taggers are brought into play. The first of these uses LDOCE senses, with any subsenses grouped where possible; the second uses categories of synonyms, and the last selectional restrictions. Known collocations are also taken into account. The implementation involved a memory-based machine learning system that was first trained on annotated data and then used to combine all the knowledge sources for WSD. Chs. 7 and 8 deal with evaluation of the author’s and other known systems for WSD. S illustrates the unsatisfactory state of evaluation tools and procedures in the area of WSD, before demonstrating that his system achieves better results thanks to the combination of a number of available lexical resources. [End Page 1022] The book is a readable introduction and description of the problems WSD poses for computational linguistics, making a clear case for a hybrid approach that uses knowledge-based and corpus-based sources of information to identify the sense of ambiguous words. Cornelia Tschichold University of Wales Swansea, Great Britain Copyright © 2005 Linguistic Society of America
In the year 2001, the French government made the Creole languages of Guadeloupe, Guyane, Martinique and Reunion, one single “Regional Language of France”. The main reason for this policy is educational and should lead to results in creating one single teacher’s assessment exam for that subject. Disregarding local differences, underestimating the complexity of the settling of regional linguistic norms, and blindly following the all Creole activist discourse, the French authorities launched a project that has had no significant positive results. This article leads to the conclusion that there is a need for a differentiating branch of linguistics and recommends the creation of normative commissions working in the field with a goal to implement efficient policies.
Traditional Chinese text chunking approach is to identify phrases using only one model and same features. It is shown that one model couldn't comprise each phrase's characteristics, and same features are not suitable to all phrases, data sparseness also appears. Multi-agent strategy uses several model and sensitive features of each phrase to identify different phrases. This paper describes the multi-agent strategy applied in the identification of Chinese phrases whose main features are: 1) easy and quick communication between phrases; 2) avoidance of data sparseness. Through testing on Chinese Penn Treebank, F score of Chinese text chunking using multi-agent strategy achieves to 95.82%, which is higher than the best result that has been reported.
In this article, I explore the ways in which ethnic identity is expressed by following the formulaic socio-linguistic norm, the very method of which defies the authenticity of identity itself, thereby asserting the identity's multi-facetedness as sustained in performative linguistic practice. I look at multi-sited socio-linguistic interactions among Koreans in Japan, who claim their primary identity to be that of North Korea's overseas citizens even though none of them have North Korean passport or nationality. Their identity, in other words, is based on ideological commitment, which is in reality supported by their ongoing linguistic practice. A close look at their socio-linguistic life reveals their ethnicity's dual or multiple ontology, which challenges among other things the currently dominant assertion of Japanese self in the western academic discourse.
In this paper we discuss the application of semi-supervised machine learning method-co-training on Chinese Text Chunking. Firstly, we give the definition of Chinese chunk,then the formalized definition of co-training algorithm.We proposed a example selection method based on the consistence, using two classifiers: Transductive HMM and fnTBL to combine a classification system to perform the Chinese text chunking task with the small-scale labled Chinese treebank and large-scale unlabled Chinese corpus. The result were compared with the self-training result and the result of the non co-training experiment in which we only used the small-scale Chinese treebank as training data and use one classifier(Transductive HMM or fnTBL) to recognize the Chinese chunk. The improvement is significant, the F value of the two classifiers reached 83.41%,85.34%, get a improvement of 2.13 points and 7.21 points respectively.
Considering the popularity of the Internet, an automatic interactive feedback system for e-learning websites is becoming increasingly desirable. However, there are still some problems for computers to understand the natural language, especially the Chinese language. First, because the Chinese language has no space to segment the lexical entry, its segmentation method is more difficult than that of English. Second, the lack of complete grammar for Chinese language makes parsing more difficult and complicated. Building an automatic Chinese feedback system for special application domains could solve these problems. In this paper, an interactive mechanism is proposed to parse, understand and response to a Chinese sentence. This mechanism utilizes a specific lexical database according to the particular application. In this way, a Chinese interactive feedback e-learning website is built for a special application domain that will choose the proper response in a user friendly, accurate and timely manner.
The aim of this paper is to present a lexical database of English collocations used in scientific language, which is being built in three Spanish universities (Barcelona, Illes Balears and Leon) and is mainly intended for the Spanish-speaking scientific community. The shortage of specialized dictionaries providing contextual information on the grammatical and collocational patterns in specific registers prompted the onset of this project. Our database is based on the analysis of a corpus of written texts in the areas of biology, biochemistry, and biomedicine, and provides the grammatical, semantic, and collocational information necessary for the correct and precise use of each term in scientific discourse. The paper describes the steps followed in the creation of the data base and it includes the case study of one of its entries.
Modern day lexical databases are not constructive, but differential (Miller et al 1990).We outline the logical structure of the task of a conceptual analyst who wishes to construct a constructive lexicon based on a Universal Theory Model of Concepts.Elementary notions of descriptive and explanatory adequacy are developed within this model.The diverse streams of evidence available for the conceptual analyst to engage in empirical inquiry are reviewed in the domains of light and perception. Theories for Constructive LexiconsMajor projects have been conducted for centuries to record, in one form or another, representations of word meanings.This has been in the form of the lexicographer's dictionary or the more modern computational linguist's electronic databases (WordNet, FrameNet, Verb-Net) e.g. last year's symposium (Miller, Fillmore, Palmer, Lenat and Hayes 2004).The degree to which computational linguists and cognitive scientists draw upon these resources as accurate descriptions of lexical knowledge is astonishing.Researchers speak of "putting meaning in your trees", providing "deep semantics", automatically labeling "semantic roles", or using Word-Net as the authority for word sense disambiguation, to name some key projects.(Palmer et al 2002, Fillmore et al 2001, Gildea and Jurafsky 2002) Given the increasing reliance to these databases, it may appear that the terms meaning and semantic are not being used glibly and that the theoretical foundations of these databases were sound.Even the most cursory analysis shows that this is not so.Consider the distinction made by the originators of WordNet between a constructive vs. differential lexicon: In a differential theory of the lexicon, meanings can be represented by any symbols that enable a theorist to distinguish among them; In a constructive theory of the lexicon, the representation should "contain sufficient information to support an accurate construction of the concept (by either a person or a machine)" (Miller et al 1990).Today's dictionaries and all of today's lexical databases are differential: the intension of synsets of WordNet are just sufficient so that someone who already knows English can distinguish among synsets, while the thematic roles used in VerbNet and FrameNet are notoriously difficult to define: primitive terms such as Agent,
We have constructed a large scale and detailed database of lexical types in Japanese from a treebank that includes detailed linguistic information. The database helps treebank annotators and grammar developers to share precise knowledge about the grammatical status of words that constitute the treebank, allowing for consistent large scale treebanking and grammar development. In this paper, we report on the motivation and methodology of the database construction. 1
In order to realize the full potential of dependency-based syntactic parsing, it is desirable to allow non-projective dependency structures. We show how a data-driven deterministic dependency parser, in itself restricted to projective structures, can be combined with graph transformation techniques to produce non-projective structures. Experiments using data from the Prague Dependency Treebank show that the combined system can handle non-projective constructions with a precision sufficient to yield a significant improvement in overall parsing accuracy. This leads to the best reported performance for robust non-projective parsing of Czech.
The paper deals with the preliminary findings from the morphologically annotated corpus of Lithuanian language (1 million running words). It was compiled and processed at the Center of Computational Linguistics, Vytautas Magnus University. Each annotation for an inflected word form of the corpus contains a lemma and a set of morphological features. The paper presents the strategy for automatic and manual annotation. Automatic annotation was carried out with the help of analyser-lemmatiser. Disambiguation of the homoforms was performed manually. Tag sets and the most prominent features of Lithuanian morphology are discussed in detail. The annotated corpus allowed us to measure the usage of parts of speech and their morphological features in contemporary Lithuanian language. The annotated corpus is of great importance for future development of parsing tools, treebanks and other NLP tools and resources for Lithuanian language.
Given a sequence of samples from an unknown probability distribution, a statistical estimator aims at providing an approximate guess of the distribution by utilizing statistics from the samples. One crucial property of a `good' estimator is that its guess approaches the unknown distribution as the sample sequence grows large. This property is called consistency. This paper concerns estimators for natural language parsing under the Data- Oriented Parsing (DOP) model. The DOP model specifies how a probabilistic grammar is acquired from statistics over a given training treebank, a corpus of sentence-parse pairs. Recently, Johnson [15] showed that the BOP estimator (called DOPl) is biased and inconsistent. A second relevant problem with DOP1 is that it suffers from an overwhelming computational inefficiency. This paper presents the first (nontrivial) consistent estimator for the DOP model. The new estimator is based on a combination of held-out estimation and a bias toward parsing with shorter derivations. To justify the need for a biased estimator in the case of DOP we prove that every non-overfitting DOP estimator is statistically biased. Our choice for the bias toward shorter derivations is justified by empirical experience, mathematical convenience and efficiency considerations. In support of our theoretical results of consistency and computational efficiency, we also report experimental results with the new estimator.
In this paper, we present a deterministic dependency structure analyzer for Chinese. This analyzer implements two algorithms – Yamada and Nivre models – and two sorts of classifiers – Support Vector Machines and Maximum Entropy methods. We compare the performance of these 2x2 combinations. We evaluate the method on a dependency tagged corpus derived from the CKIP Treebank corpus. Then, we analyzed the errors in the experiments and found that some errors were caused by mistakes of nominal compounds analysis. Therefore we adopt an NP-chunker to solve this problem.
We present a method for automatic RMRS semantics construction from dependency structures, following the semantic algebra of Copestake et al. (2001). We have applied this method to a subset of the TIGER Dependency Bank for German (Forst et al., 2004) to obtain a semantic treebank for (HPSG) parser evaluation. We describe the semantics construction mechanism and give evaluation figures from manual validation of the treebank. These indicate high precision of the automatic RMRS construction process.