Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Data-driven parsing techniques have a number of advantages over rule-based parsing techniques, such as fast development time, broad-coverage and robustness. Treebanks, collections of syntactically annotated sentences, are important resources for data-driven parsers. When developing a parser for Swedish one needs a treebank containing Swedish sentences, but currently there is a lack of Swedish treebanks of substantial size. This holds for the other Nordic languages too, with Danish as an exception. The absence of Swedish treebanks is remarkable considering that two corpora of Swedish text augmented with syntactic annotation have been created, one as early as 1974 named Talbanken (Einarsson 1976), and another in the 80's named Syntag (Järborg 1980). Unfortunately, the annotation formats of these resources make them cumbersome to use for modern treebank tools and parsers. In a way, Sweden can be regarded as a pioneer in this area, but thereafter the work with creating new treebanks has decreased considerably.
Great progress has been made in parsing the Wall Street Journal portion of the Penn Treebank. Now parsing languages other than English is an intensive research area. Head-driven model is one of the best English parsing models. It has been successfully applied to Czech but failed to outperform a base-line model in parsing German. This paper attempts to parse Chinese with head-driven model. Promising experimental results demonstrate that head-driven model works well for Chinese. We propose a hybrid parsing strategy, which combines head-driven model with a Chinese base phrases parsing model. The combined model not only improves the performance but also makes the parser space and time efficient. We evaluate our method in PARSEVAL measures, and the combined model performances are at 79.88% precision, 81.97% recall.
Linguistically annotated corpus based on texts in biomedical domain has been constructed to tune natural language processing (NLP) tools for biotextmining. As the focus of information extraction is shifting from "nominal" information such as named entity to "verbal " information such as function and interaction of substances, application of parsers has become one of the key technologies and thus the corpus annotated for syntactic structure of sentences is in demand. A subset of the GENIA corpus consisting of 500 MEDLINE abstracts has been annotated for syntactic structure in an XMLbased format based on Penn Treebank II (PTB) scheme. Inter-annotator agreement test indicated that the writing style rather than the contents of the research abstracts is the source of the difficulty in tree annotation, and that annotation can be stably done by linguists without much knowledge of biology with appropriate guidelines regarding to linguistic phenomena particular to scientific texts. 1
In academic courses in which one task for the students is to understand empirical methodology and the nature of scientific inquiry, the ability of students to create and implement their own experiments allows them to take intellectual ownership of, and greatly facilitates, the learning process. The Psychology Experiment Authoring Kit (PEAK) is a novel spreadsheet-based interface allowing students and researchers with rudimentary spreadsheet skills to create cognitive and cognitive neuroscience experiments in minutes. Students fill in a spreadsheet listing of independent variables and stimuli, insert columns that represent experimental objects such as slides (presenting text, pictures, and sounds) and feedback displays to create complete experiments, all within a single spreadsheet. The application then executes experiments with centisecond precision. Formal usability testing was done in two stages: (1) detailed coding of 10 individual subjects in one-on-one experimenter/subject videotaped sessions and (2) classroom testing of 64 undergraduates. In both individual and classroom testing, the students learned to effectively use PEAK within 2 h, and were able to create a lexical decision experiment in under 10 min. Findings from the individual testing in Stage 1 resulted in significant changes to documentation and training materials and identification of bugs to be corrected. Stage 2 testing identified additional bugs to be corrected and new features to be considered to facilitate student understanding of the experiment model. Such testing will improve the approach with each semester. The students were typically able to create their own projects in 2 h.
Arabic is usually written without short vowels and additional diacritics, which are nevertheless important for several applications. We present a novel algorithm for restoring these symbols, using a cascade of probabilistic finite-state transducers trained on the Arabic treebank, integrating a word-based language model, a letter-based language model, and an extremely simple morphological model. This combination of probabilistic methods and simple linguistic information yields high levels of accuracy.
There is growing evidence that Internet-mediated psychological tests can have satisfactory psychometric properties and can measure the same constructs as traditional versions. However, equivalence cannot be taken for granted. The prospective memory questionnaire (PMQ; Hannon, Adams, Harrington, Fries-Dias, & Gibson, 1995) was used in an on-line study exploring links between drug use and memory (Rodgers et al., 2003). The PMQ has four factor-analytically derived subscales. In a large (N763) sample tested via the Internet, only two factors could be recovered; the other two subscales were essentially meaningless. This demonstration of nonequivalence underlines the importance of on-line test validation. Without examination of its psychometric properties, one cannot be sure that a test administered via the Internet actually measures the intended construct.
This paper addresses the current needs for so-called emotion in speech, but points out that the issue is better described as the expression of relationships and attitudes rather than the currently held raw (or big-six) emotional states. From an analysis of more than three years of daily conversational speech, we find the direct expression of emotion to be extremely rare, and contend that when speech technologists say that what we need now is more ‘emotion’ in speech, what they really mean is that the current technologies are too text-based, and that more expression of speaker attitude, affect, and discourse relationships is required.
This paper presents the first probabilistic parsing results for French, using the recently released French Treebank. We start with an unlexicalized PCFG as a baseline model, which is enriched to the level of Collins' Model 2 by adding lexicalization and subcategorization. The lexicalized sister-head model and a bigram model are also tested, to deal with the flatness of the French Treebank. The bigram model achieves the best performance: 81% constituency F-score and 84% dependency accuracy. All lexicalized models outperform the unlexicalized baseline, consistent with probabilistic parsing results for English, but contrary to results for German, where lexicalization has only a limited effect on parsing performance.
The NITE XML Toolkit (NXT) is open source software for working with language corpora, with particular strengths for multimodal and heavily cross-annotated data sets. In NXT, annotations are described by types and attribute value pairs, and can relate to signal via start and end times, to representations of the external environment, and to each other via either an arbitrary graph structure or a multi-rooted tree structure characterized by both temporal and structural orderings. Simple queries in NXT express variable bindings for n-tuples of objects, optionally constrained by type, and give a set of conditions on the n-tuples combined with boolean operators. The defined operators for the condition tests allow full access to the timing and structural properties of the data model. A complex query facility passes variable bindings from one query to another for filtering, returning a tree structure. In addition to describing NXTȁ9s core data handling and search capabilities, we explain the stand-off XML data storage format that it employs and illustrate its use with examples from an early adopter of the technology.
In this paper we argue for the importance of doing inference over the information expressed by the annotations of temporally annotated corpora. We describe the process of inferential closure which can be applied to determine the full temporal content that follows from an annotation. We illustrate the importance of temporal inference and temporal closure in relation to three tasks, which are: (a) the comparison of different temporal annotations, (b) facilitating the manual annotation process needed to create temporally annotated corpora and (c) empirical investigations done over temporally annotated data.
A growing number of studies have supported the use of unidimensional psychometric test instruments administered via the Internet; however, support for the use of multidimensional scales is weak. The present study compares paper and Internet administrations of the Multidimensional Health Locus of Control (MHLC) Scale (Wallston & Wallston, 1981). In terms of reliabilities and factor structures, the Internet data were found to be at least as good as the paper data. MHLC scores were comparable for paper and Internet administrations, although the Internet sample scored significantly lower on the Powerful Others subscale. Overall, the results show that administration of the MHLC Scale via the Internet can produce data comparable to that obtained by pen-and-paper methods. However, it is concluded that generalization of these findings beyond the psychometric test instrument and sampling procedures used here is not warranted.
A mathematical model previously developed for use in computer vision applications is presented as an empirical model for face space. The term appearance space is used to distinguish this from previous models. Appearance space is a linear vector space that is dimensionally optimal, enables us to model and describe any human facial appearance, and possesses characteristics that are plausible for the representation of psychological face space. Randomly sampling from a multivariate distribution for a location in appearance space produces entirely plausible faces, and manipulation of a small set of defining parameters enables the automatic generation of photo-realistic caricatures. The appearance space model leads us to the new concept of nonlinear caricatures, and we show that the accepted linear method for caricature is only a special case of a more general paradigm. Nonlinear methods are also viable, and we present examples of photographic quality caricatures, using a number of different transformation functions. Results of a simple experiment are presented that suggest that nonlinear transformations can accurately capture key aspects of the caricature effect. Finally, we discuss the relationship between appearance space, caricature, and facial distinctiveness. On the basis of our new theoretical framework, we suggest an experimental approach that can yield new evidence for the plausibility of face space and its ability to explain processes of recognition.
The annotations of the Penn Discourse Treebank (PDTB) include (1) discourse connectives and their arguments, and (2) attribution of each argument of each connective and of the relation it denotes. Because the PDTB covers the same text as the Penn TreeBank WSJ corpus, syntactic and discourse annotation can be compared. This has revealed significant differences between syntactic structure and discourse structure, in terms of the arguments of connectives, due in large part to attribution. We describe these differences, an algorithm for detecting them, and finally some experimental results. These results have implications for automating discourse annotation based on syntactic annotation.
These two letters and two inventories preserved in the rich heritage of Anton Hodinka in the manuscript depository of the Hungarian Academy of Sciences Library present an exciting picture of the everyday life of the 18th century. Nevertheless, I find these documents valuable not because of this fact but due to their vocabulary which reflects the Rusyn language adequately. These original sources are the splendid illustrations of the Rusyn language wordstock used in everyday life of that period. Therefore I have not spared myself to copy, study and publish the manuscripts in question because I should like to contribute to enriching the Rusyn language history. As a matter of fact the Rusyn language of the 18th century reflects the synthesis of three elements: the Church Slavonic liturgy language, the Old Ukrainian language and the living folk language. The formation and unification of the literary language norm, which was not regulated by grammars and dictionaries, was greatly influenced by the bishop's office documents due to the great authority and prestige of the church in the region. The three above-mentioned elements of the Rusyn literary language of the 18th century can be revealed in all language layers (phonetical, morphological, syntactical, lexical, semantical). I shall give several examples on the elements of the Rusyn folk language.
Many recent annotation efforts for English have focused on pieces of the larger problem of semantic annotation, rather than initially producing a single unified representation. This paper discusses the issues involved in merging four of these efforts into a unified linguistic structure: PropBank, NomBank, the Discourse Treebank and Coreference Annotation undertaken at the University of Essex. We discuss resolving overlapping and conflicting annotation as well as how the various annotation schemes can reinforce each other to produce a representation that is greater than the sum of its parts.
TextTrees, introduced in (Newman, 2005), are skeletal representations formed by systematically converting parser output trees into unlabeled indented strings with minimal bracketing. Files of TextTrees can be read rapidly to evaluate the results of parsing long documents, and are easily edited to allow limited-cost treebank development. This paper reviews the TextTree concept, and then describes the implementation of the almost parser- and grammar-independent TextTree generator, as well as auxiliary methods for producing parser review files and inputs to bracket scoring tools. The results of some limited experiments in TextTree usage are also provided.
It is widely believed that the difference between regular and irregular verbs is restricted to form. This study questions that belief. We report a series of lexical statistics showing that irregular verbs cluster in denser regions in semantic space. Compared to regular verbs, irregular verbs tend to have more semantic neighbors that in turn have relatively many other semantic neighbors that are morphologically irregular. We show that this greater semantic density for irregulars is reflected in association norms, familiarity ratings, visual lexical-decision latencies, and word-naming latencies. Meta-analyses of the materials of two neuroimaging studies show that in these studies, regularity is confounded with differences in semantic density. Our results challenge the hypothesis of the supposed formal encapsulation of rules of inflection and support lines of research in which sensitivity to probability is recognized as intrinsic to human language.
This article considers approaches which rerank the output of an existing probabilistic parser. The base parser produces a set of candidate parses for each input sentence, with associated probabilities that define an initial ranking of these parses. A second model then attempts to improve upon this initial ranking, using additional features of the tree as evidence. The strength of our approach is that it allows a tree to be represented as an arbitrary set of features, without concerns about how these features interact or overlap and without the need to define a derivation or a generative model which takes these features into account. We introduce a new method for the reranking task, based on the boosting approach to ranking problems described in Freund et al. (1998). We apply the boosting method to parsing the Wall Street Journal treebank. The method combined the log-likelihood under a baseline model (that of Collins [1999]) with evidence from an additional 500,000 features over parse trees that were not included in the original model. The new model achieved 89.75% F-measure, a 13% relative decrease in F-measure error over the baseline model's score of 88.2%. The article also introduces a new algorithm for the boosting approach which takes advantage of the sparsity of the feature space in the parsing data. Experiments show significant efficiency gains for the new algorithm over the obvious implementation of the boosting approach. We argue that the method is an appealing alternative-in terms of both simplicity and efficiency-to work on feature selection methods within log-linear (maximum-entropy) models. Although the experiments in this article are on natural language parsing (NLP), the approach should be applicable to many other NLP problems which are naturally framed as ranking tasks, for example, speech recognition, machine translation, or natural language generation.
Recent years have seen a revived interst in semantic parsing by applying statistical and machinelearning methods to semantically annotated corpora such as the FrameNet and the Proposition Bank. So far much of the research has been focused on English due to the lack of semantically annotated resources in other languages. In this paper, we report first results on semantic role labeling using a pre-release version of the Chinese Proposition Bank. Since the Chinese Proposition Bank is superimposed on top of the Chinese Treebank, i.e., the semantic role labels are assigned to constituents in a treebank parse tree, we start by reporting results on experiments using the handcrafted parses in the treebank. This will give us a measure of the extent to which the semantic role labels can be bootstrapped from the syntactic annotation in the treebank. We will then report experiments using a fully automatic Chinese parser that integrates word segmentation, POS-tagging and parsing. This will gauge how successful semantic role labeling can be done for Chinese in realistic situations. We show that our results using hand-crafted parses are slightly higher than the results reported for the state-of-the-art semantic role labeling systems for English using the Penn English Proposition Bank data, even though the Chinese Proposition Bank is smaller in size. When
We formalize weighted dependency parsing as searching for maximum spanning trees (MSTs) in directed graphs. Using this representation, the parsing algorithm of Eisner (1996) is sufficient for searching over all projective trees in O(n3) time. More surprisingly, the representation is extended naturally to non-projective parsing using Chu-Liu-Edmonds (Chu and Liu, 1965; Edmonds, 1967) MST algorithm, yielding an O(n2) parsing algorithm. We evaluate these methods on the Prague Dependency Treebank using online large-margin learning techniques (Crammer et al., 2003; McDonald et al., 2005) and show that MST parsing increases efficiency and accuracy for languages with non-projective dependencies.
The acquisition of grammar from a corpus is a challenging task in the preparation of a knowledge bank. In this paper, we discuss the extraction of Chinese grammar oriented to a restricted corpus. First, probabilistic context-free grammars (PCFG) are extracted automatically from the Penn Chinese Treebank and are regarded as the baseline rules. Then a corpusoriented grammar is developed by adding specific information including head information from the restricted corpus. Then, we describe the peculiarities and ambiguities, particularly between the phrases “PP” and “VP” in the extracted grammar. Finally, the parsing results of the utterances are used to evaluate the extracted grammar.
There has been a contemporary surge of interest in the application of stochastic models of parsing. The use of tree-adjoining grammar (TAG) in this domain has been relatively limited due in part to the unavailability, until recently, of large-scale corpora hand-annotated with TAG structures. Our goals are to develop inexpensive means of generating such corpora and to demonstrate their applicability to stochastic modeling. We present a method for automatically extracting a linguistically plausible TAG from the Penn Treebank. Furthermore, we also introduce labor-inexpensive methods for inducing higher-level organization of TAGs. Empirically, we perform an evaluation of various automatically extracted TAGs and also demonstrate how our induced higher-level organization of TAGs can be used for smoothing stochastic TAG models.
Whether or not the ‘-(eu)ro’ Kasus phrase as a sentence constituent is indispensable is fundamentally related with the argument structure of verb. That is to say, it means whether the ‘-(eu)ro’ Kasus phrase is syntactically required or not. But in spite of such cognition, it is really hard that we judge whether it is true or not. Therefore, I approach the problem through picking out the constituent which has semantically adjunct function. Then, the word order has to be considered. With the result, the ‘-(eu)ro’ Kasus phrases of the functions such as via, cause or reason, sense or value, instrumental, material, time or space background, norm, spacial background, and the ‘-(eu)ro’ Kasus phrases or the ‘-eseo’ Kasus phrases which express semantically equal meaning with NP of ‘NP+ha-’ construction or with Subject NP of verb ‘byunha-’ construction, and the idiomatic ‘-(eu)ro’ Kasus phrases or ‘-(eu)rosseo’ Kasus phrases, etc. are all estimated as the adjuncts. But the ‘-(eu)ro’ Kasus phrases of the functions such as goal, result or product, status property, and those related to time in the verb ‘(jeong)ha-’ construction, etc. are all estimated as the arguments. If such ‘-(eu)ro’ Kasus phrases are adjuncts, then they have a modifier scope individually such as adverb. It is deeply related with the unmarked positions of those. According to such description some adjuncts modify the lexical category, some adjuncts modify the sentence category, some adjuncts modify the mediate category which is larger than the lexical category but is smaller than the sentence category. However, idiomatic ‘-(eu)ro’ Kasus phrases or ‘-(eu)rosseo’ Kasus phrases are always the sentence category adjuncts, but ‘-(eu)roseo’ Kasus phrases in front of the sentence are not only those but also the ‘situation presentive word’ related to topic.
Abstract. We propose to apply classical development methodologies to the design and implementation of Lexical Databases(LDB), which embody conceptual and linguistic knowledge. We represent the conceptual knowledge as an ontology, and the linguistic knowledge, which depends on each language, in lexicons. Our approach is based on a single language-independent ontology. Besides, we study some conceptual and linguistic requirements; in particular, meaning classifications in the ontology, focusing on taxonomies. We have followed a classical software development methodology for implementing lexical information systems in order to reach robust, maintainable, and integrateable relational databases (RDB) for storing the conceptual and linguistic knowledge. 1
In order to realize the full potential of dependency-based syntactic parsing, it is desirable to allow non-projective dependency structures. We show how a data-driven deterministic dependency parser, in itself restricted to projective structures, can be combined with graph transformation techniques to produce non-projective structures. Experiments using data from the Prague Dependency Treebank show that the combined system can handle non-projective constructions with a precision sufficient to yield a significant improvement in overall parsing accuracy. This leads to the best reported performance for robust non-projective parsing of Czech.
This article is devoted to the problem of quantifying noun groups in German. After a thorough description of the phenom ena, the results of corpus-based investigations are described. Moreover, some examples are given that underline the necessity of integrating some kind of information other than grammar sensu stricto into the treebank. We argue that a more sophisticated and fine-grained annotation in the treebank would have very positve effects on stochastic parsers trained on the treebank and on grammars induced from the treebank, and it would make the treebank more valuable as a source of data for theoretical linguistic investigations. The information gained from corpus research and the analyses that are proposed are realized in the framework of SILVA, a parsing and extraction tool for German text corpora.
In this paper, we extend an existing parser to produce richer output annotated with function labels. We obtain state-of-the-art results both in function labelling and in parsing, by automatically relabelling the Penn Treebank trees. In particular, we obtain the best published results on semantic function labels. This suggests that current statistical parsing methods are sufficiently general to produce accurate shallow semantic annotation.
This paper investigates the automatic identification of aspects of Information Structure (IS) in texts. The experiments use the Prague Dependency Treebank which is annotated with IS following the Praguian approach of Topic Focus Articulation. We automatically detect t(opic) and f(ocus), using node attributes from the treebank as basic features and derived features inspired by the annotation guidelines. We show the performance of C4.5, Bagging, and Ripper classifiers on several classes of instances such as nouns and pronouns, only nouns, only pronouns. A baseline system assigning always f(ocus) has an F-score of 42.5%. Our best system obtains 82.04%.
Abstract This study explores the relevance of suffix allomorphy for processing complex words. The question is whether structural invariance of the morphological category (i.e., lack of allomorphy) would affect the processing of Finnish derived words. A series of four visual lexical decision experiments in which alternatively surface and base frequency was manipulated showed that the two invariant suffixes, namely denominal –stO and deadjectival –hkO, showed reliable effects of base frequency, whereas for the two categories with suffix allomorphy, deverbal –Us and deadjectival –(U)Us, only surface frequency played a role. A further experiment showed that even with the most frequent variant of –(U)Us, namely –Ude-, response latencies were a function of surface frequency only. It is shown that neither the results from the experiments here nor previous findings from processing Finnish words can be accounted for by suffix frequency, the frequency ratio between the derived word and its base, or morphological productivity in any straightforward manner. We conclude that the lack of allomorphy, that is, structural invariance, significantly adds to affixal salience and therefore enhances morphological decomposition. The implications of this finding for models of lexical processing are discussed. We would like to thank Matti Laine and an anonymous reviewer for their helpful comments. Notes 1To be precise, typologically Finnish morphology is fusional-agglutinative in distinction to that of Turkish, which gives us a schoolbook example of agglutinative morphology. 2The capital vowels in –jA and the suffixes investigated here, namely –stO, –hkO, – (U)Us, and –Us, refer to the archiphoneme marking phonological adjustment following from Finnish vowel harmony. That is, for instance in the case of –jA, the suffix is realised as either /ja/ or /jæ/ depending on whether the stem has a back or front vowel respectively. 3Although perceptually homonymous with the deverbal –jA, the particular variant of partitive plural inflection is structurally bimorphemic with a combination of a partitive allomorph –A and a phonologically motivated change of the plural marker –i to –j in /V_V/ contexts. 4The frequency and productivity counts are based on the lexical database compiled from seven consecutive annual volumes of a Finnish newspaper Karjalainen (1991–1997). The material consists of 34.5 million word tokens and covers a reasonably long stretch of time to provide a representative view of the use and productivity of derivational affixation. The corpus (Karjalaisen korpus, 34.5 million-word token computer-based newspaper corpus of Finnish based on Karjalainen (Joensuu), compiled by J. Niemi and his associates at the Linguistics Department, University of Joensuu, SGML form created at the Department of General Linguistics, University of Helsinki) is available from Kielipankki at http://www.csc.fi/kielipankki/. 5This is also reflected in the fact that in the Karjalainen database –(U)Us is attested in noun bases with about 120 types and 12 hapaxes, resulting in a comparably low index of productivity, p =.0001 (cf. Table 1). 6The results from another run of the surface frequency experiment with different items (16 per condition) and 25 participants showed a significant difference between the High Surface (674 ms) and Low Surface (722 ms) conditions, t 1(24) = 4.05, p <.001, t 2(30) = 2.46, p <.05, confirming that surface frequency is indeed a significant factor in the processing of words in –hkO. 7In line with the tradition in terminology, we have also here used the term Surface Frequency to reflect the cumulative frequency of the word and all its inflected forms. It should be noted though that the surface frequency proper (here the genitive form) highly correlates with the surface frequency as we defined it. That is, for the base frequency experiment, the genitive frequencies are matched (High: 0.2; Low: 0.3; t<1), for the surface frequency experiment they are manipulated (High: 4.3; Low: 0.3; p<.001). 8Although they do note that the connectionist three-layer network model of Davis, van Casteren, and Marslen-Wilson (Citation2003) can also accommodate their pattern of findings. 9More precisely, –Us and –(U)Us are homonymous only in nominative singular and in all plural cases except the nominative. Furthermore, words in deadjectival –(U)Us quite rarely occur in any of the plurals, where the overlap of the two suffixes would be greatest. For example, for a frequent noun in –(U)Us, 'kauneus' (beauty), that is found 722 times in the Karjalainen database, there is not a single instance of plural inflection. 10Also, Finnish stem allomorphy is abundant and often not amenable to straightforward rule-based generalisations. It might therefore come as no surprise, that, in contrast to the wealth of experimental results supporting the phonological underspecification account for the representation of English stem allomorphy, several studies have found evidence for the position that in Finnish stem allomorphs have separate representations in the mental lexicon. This evidence is obtained from a variety of experimental paradigms, such as lexical decision (e.g., Laine, Vainio, & Hyönä, Citation1999), priming (Järvikivi & Niemi, Citation2002a), masked priming (Järvikivi & Niemi, Citation2002b), slips of the tongue (Niemi & Laine, Citation1992), and aphasia studies (Laine et al, Citation1995; Laine & Niemi, Citation1997). Additional informationNotes on contributorsJuhani Järvikivi Correspondence should be addressed to Juhani Järvikivi, Department of Psychology, Assistentinkatu 7, FIN-20014 University of Turku, Finland. juhani.jarvikivi@utu.fi
We present a methodology for extracting subcategorization frames based on an automatic lexical-functional grammar (LFG) f-structure annotation algorithm for the Penn-II and Penn-III Treebanks. We extract syntactic-function-based subcategorization frames (LFG semantic forms) and traditional CFG category-based subcategorization frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. In contrast to many other approaches, ours does not predefine the subcategorization frame types extracted, learning them instead from the source data. Including particles and prepositions, we extract 21,005 lemma frame types for 4,362 verb lemmas, with a total of 577 frame types and an average of 4.8 frame types per verb. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource. To our knowledge, this is the largest and most complete evaluation of subcategorization frames acquired automatically for English.
We introduce a method for transferring annotation from a syntactically annotated corpus in a source language to a target language. Our approach assumes only that an (unannotated) text corpus exists for the target language, and does not require that the parameters of the mapping between the two languages are known. We outline a general probabilistic approach based on Data Augmentation, discuss the algorithmic challenges, and present a novel algorithm for sampling from a posterior distribution over trees.
The Korean Treebank Annotations Version 2.0 is a second volume of The Korean Treebank Annotations (Palmer et al., 2002; Han et al., 2002). It contains new texts that are from the news domain: the original corpus for the Korean Treebank 2.0 was extracted from The Korean Newswire corpus published by LDC, catalog number LDC2000T45. The Korean Treebank Annotations Version 2.0 consists of 647 news articles in 112 files which contain 132,040 words and 5,010 sentences. There are 40,252 unique words and 13,844 unique morphemes (12,681 unique morphemes excluding foreign characters and arabic numbers). The annotated text measures about 8.5MB in size.\nWhile annotating the new texts, many new linguistic constructions and phenomena were encountered which called for setting additional guidelines. Furthermore, a few guidelines used for the first volume of the Korean Treebank were re-examined and modified in the second volume. This document outlines the guidelines that were newly introduced for the second volume of the Penn Korean Treebank, as well as the ones that have been revised since the publication of volume 1.0. Therefore, this is not a self-contained document, but is rather an addendum to the two previously published guidelines for the Penn Korean Treebank (Han and Han, 2001; Han et al., 2001).
OBJECTIVE: In multiple sclerosis (MS), magnetic resonance imaging (MRI) predictors of cognitive impairment are based on sophisticated computer-generated analyses that are difficult to apply in clinical settings. This study investigated the clinical usefulness of a new visual rating scale, the Cholinergic Pathways Hyperintensities Scale (CHIPS), in detecting cognitive dysfunction. METHODS: Forty clinically definite MS patients underwent a brain MRI. Based on the CHIPS, cholinergic pathway hyperintensities were rated in 10 regions on four axial slices. Computerized hyperintense lesion volumes were also obtained. For cognitive testing, The Neuropsychological Screening Battery for Multiple Sclerosis was used. "Low" and "High" lesion score groups were computed based on the mean of the total CHIPS score. Optimal sensitivity and specificity of the total CHIPS score in detecting cognitive impairment were determined using a receiver operator characteristic curve. RESULTS: Despite a similar demographic profile, subjects with a "High" lesion score performed significantly worse than the "Low" lesion score group on verbal (P =.007) and visuospatial (P =.02) memory, and on a global index of cognitive functioning (P =.001). Optimal sensitivity (82%) and specificity (83%) were reached with a threshold total CHIPS score of 18 points. Total CHIPS score and total hyperintense lesion load were correlated (sigma = 0.82, P <.0001). CONCLUSION: CHIPS is helpful in clinically predicting cognitive impairment in MS.
We describe a parallel annotation approach for PubMed abstracts. It includes both entity/relation annotation and a treebank containing syntactic structure, with a goal of mapping entities to constituents in the treebank. Crucial to this approach is a modification of the Penn Treebank guidelines and the characterization of entities as relation components, which allows the integration of the entity annotation with the syntactic structure while retaining the capacity to annotate and extract more complex events.
Linguistic research and language technology development employ large repositories of ordered trees. XML, a standard ordered tree model, and XPath, its associated language, are natural choices for storing and querying linguistic data. However, several important expressive features required for linguistic queries are missing in XPath. In this paper, we motivate and illustrate these features with a variety of linguistic queries. Then we define extensions to XPath which support linguistic tree queries. We provide a relational representation for trees, and define an SQL translation for queries. Experiments demonstrate that the query system is significantly faster than other linguistic tree query systems for a wide range of queries. 1
We present the case for an extensive scientific effort to build up large treebanks for the Nordic and Baltic languages, as a step towards developing advanced multilingual communication technologies for these languages in the future.
Abstract The essential first step in the development of any quantitative method is identifying something to measure. If we want to use a quantitative technique to establish whether languages are likely to be related, or whether they fall into the same subgroup, or how similar two languages are relative to a third language, we need to decide what we are going to count; and to ensure that we are comparing like with like, the something we are counting has to be the same across all the languages we are comparing. Indeed, to make the method as flexible as possible we need that something to be the same in all the languages we might ever want to compare. In comparative linguistics this is a tall order. Our first thought might be to turn to sociolinguistics, where quantitative methods have enjoyed such success, and adopt the strategies developed there. However, sociolinguistic studies tend to operate within a single speech community (Patrick 2002), where speakers share the same norms of behaviour and attitude, and the same variable elements of linguistic structure. It is possible, therefore, to isolate a set of variables and to study the circumstances, both linguistic and non-linguistic, under which the different variants emerge. For example, we might consider a phonological variable (t), with variants [t], used mainly by women, and glottal stop [?], mainly used by men. We might find a syntactic variable (negation), with multiple negation (I didn’t do nothing ) favoured by lower-class speakers, and single negative markers (I didn’t do anything) the dominant variant for middle-class speakers; or a lexical variable, where the meaning (be sick) might be expressed by vomit for older speakers and throw up for younger ones.