Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Abstract This study attempted to demonstrate an elevated disgust sensitivity in bulimia nervosa. Eleven bulimic patients and 12 control subjects underwent a functional magnetic resonance imaging (fMRI) study in which they were presented with alternating blocks of 40 disgust‐inducing, 40 fear‐inducing and 40 affectively neutral scenes. Each scene was shown for 1.5 s. After completion of all blocks, affective ratings were then determined. The viewing of the disgusting pictures, which had been rated as highly repulsive by the bulimic females, was associated with an activation of the left amygdala and the occipito‐temporal visual cortex. The subjective and brain‐physiological responses did not differ from those of the healthy control subjects. This held true for the fear‐inducing scenes as well. Thus, bulimic patients are not characterized by an increased global disgust sensitivity and they do not show any indication of an altered central processing of generally disgust and fear‐inducing visual stimuli. Copyright © 2004 John Wiley & Sons, Ltd and Eating Disorders Association.
The preceding articles by Piek Vossen (PV), Willy Martin (WM), and Marc van Campenoudt (MC) give a much more detailed account on their respective multilingual lexical database designs than the article by myself in this same journal (MJ). At the same time, they indicate some points of concern regarding the set-up of the SIM<it>u</it>LLDA system. Rather than responding directly to the points raised, this response elaborates on the two aspects of the SIM<it>u</it>LLDA system that seem to form the main sources of these issues: the status of the definitional attributes, and the practical usability of the system. The issues raised in the preceding articles will be explicitly addresses in the course of this elaboration. For even more details on these topics, see Janssen (2002).
This paper presents a system for automatically generating discourse structures from written text. The system is divided into two levels: sentence-level and text-level. The sentence-level discourse parser uses syntactic information and cue phrases to segment sentences into elementary discourse units and to generate discourse structures of sentences. At the text-level, constraints about textual adjacency and textual organization are integrated in a beam search in order to generate best discourse structures. The experiments were done with documents from the RST Discourse Treebank. It shows promising results in a reasonable search space compared to the discourse trees generated by human analysts.
Function tags are a context-sensitive annotation applied to words and phrases of natural language text, marking their syntactic or semantic role within a larger utterance. As researchers improve results on various other problems in “pure” natural language processing (e.g part-of-speech tagging, parsing), those who work in the more “applied” NLP fields (e.g. question-answering, temporal analysis) are seeking more powerful sorts of linguistic annotation as input for their own systems. Hence, function tags. In the first part of the thesis, I present the problem of function tagging: why it is an interesting problem, who has worked on similar thing, and what exactly I intend to do. I briefly review the function tags of the Penn treebank, and explain the specific metrics by which I will evaluate my work. In the second part of the thesis, I introduce the many features that I will use to train a function tagging system, and then I present some systems that make use of them: one using feature trees, one using decision trees (briefly), and one using perceptron models. For each system, I give a brief historical perspective, an overview of where it has been used before and why I think it will be useful in this task. I will then try a number of feature combinations with interesting properties; and finally, present the best-performing tweaked-out version of that system. Finally, in the third part of the thesis, I bring them all together and discuss the advantages and disadvantages of each system in various situations. More interestingly, I will present an analysis of what features prove to be the most helpful for the different function tagging subtasks. Lastly, I will present a comparison to other systems performing related tasks, and speculate on some interesting future work.
機械翻訳システムによる翻訳を人間による翻訳に近づけるために取り組むべき課題を明らかにしようという試みの一環として, 本稿では, ニュース記事から無作為抽出した英文を英日機械翻訳システムで翻訳した結果と, これらの英文を人間が翻訳した結果を照らし合わせ, 両者の間で使用されている動詞の馴染み度の分布に違いがあるかどうかを計量的に分析した. 動詞の馴染み度を測る尺度としては, NTTの単語親密度データベースを利用した. 分析の結果, 機械翻訳システムによる翻訳と人間による翻訳の間で単語親密度の分布に統計的有意差は認められず, 使用されている動詞の馴染み度に関しては両者の間で違いがないということが示唆された. 従って, 格要素などとの共起関係を考えず動詞だけに着目した場合, 調査対象とした機械翻訳システムでは動詞の翻訳品質は一定のレベルに達していると判断できる.
Predicting the location of phrase breaks within an utterance is an important task in text-to-speech synthesis, and can be done with reasonable accuracy using part-of-speech (POS) tags as features. However, it seems unlikely that the 40 or more different tags used by most taggers all contribute to this task, and in fact many may contribute noise. In this paper, we present an algorithm for reducing the standard Penn Treebank POS tag set for use in predicting phrase breaks. Using the best first search approach, the algorithm considers possible groupings of tags, searching the groupings that yield the highest overall performance. The reduced tag sets were evaluated by an n-gram model trained on POS sequences along with their associated juncture (break/non-break), the reduced tag set raised the model's performance on junctures correct from 90.38% to 92.43%, and reduced insertions from 2.89% to 1.83%.
The aim of this study is to highlight what kind of information distinguishes abstract and concrete conceptual knowledge in different aged children. A familiarity-rating task has shown that 8-year-olds judged concrete concepts as very familiar while abstract concepts were judged as much less familiar with ratings increasing substantially from age 10 to age 12, according to literature showing that abstract terms are not mastered until adolescence (Schwanenflugel, 1991). The types of relation elicited by abstract and concrete concepts during development were investigated in an association production task. At all considered age levels, concrete concepts mainly activated attributive and thematic relations as well as, to a much lesser extent, taxonomic relations and stereotypes. Abstract concepts, instead, elicited mainly thematic relations and, to a much lesser extent, examples and taxonomic relations. The patterns of relations elicited were already differentiated by age 8, becoming more specific in abstract concepts with age.
This paper argues for the development of parallel treebanks. It summarizes the work done in this area and reports on experiments for building a Swedish-German treebank. And it describes our approach for reusing resources from one language while annotating another language.
Corpus linguistics prompts a lexicocentric approach to linguistic theory. The theory of norms and exploitations (TNE; Hanks, forthcoming) is such a theory, applying the insights of prototype theory and Sinclairian text analysis to the empirical evidence of large corpora. By studying words in context, we can identify the normal patterns of usage that are associated with each word. A meaning, or meaning potential, can then be associated with each pattern. Thereby, lexical entropy is reduced. A central question in this approach to language analysis concerns metaphors and idioms. In the present paper, conventional metaphors and idioms are classified as ‘norms’ (i.e. conventional uses), while dynamic, ad-hoc metaphors are classified as ‘exploitations’ of norms. However, conventional metaphors can still be distinguished from literal meanings. At least in some cases, conventional metaphors differ from literal senses by their particular syntagmatic patterns. The paper also discusses the importance of text type and domain in achieving a satisfactory interpretation of idiomatic expressions.
This paper describes an incremental parsing approach where parameters are estimated using a variant of the perceptron algorithm. A beam-search algorithm is used during both training and decoding phases of the method. The perceptron approach was implemented with the same feature set as that of an existing generative model (Roark, 2001a), and experimental results show that it gives competitive performance to the generative model on parsing the Penn treebank. We demonstrate that training a perceptron model to combine with the generative model during search provides a 2.1 percent F-measure improvement over the generative model alone, to 88.8 percent.
In natural language processing a huge amount of structured data is constantly used for the extraction and presentation of grammatical structures in sentences. For example the Chinese Treebank corpus developed at the Institute of Information Science Academia Sinica Taiwan is a semantically annotated corpus that has been used to help parse and study Chinese sentences. In this setting users usually use structured tree patterns instead of keywords to query the corpus.
The ideas presented in this paper have arisen during the work on a database system designed for the description and comparison of human languages in a theoretically well-founded and systematic way. The Cross-linguistic Reference Grammar (CRG) is a database application for unified descriptions of natural languages of any kind, including sign languages, that has been developed over the last decade at LMU Munich. The core of the application is an XML-based client-server-DBMS called Systematics (Nickles 2001) that implements a class system of linguistic phenomena as a kind of Be-Have-tree (Peterson 2002). Examples are uniformly coded in one of the interlinear representation formats provided by CRG for three ontological categories and modalities of linguistic signs: spoken, signed, and written (Zaefferer 2003). The present paper is about one of the criteria of adequacy that has informed CRG system design from the beginning. I call it pro-competitivess and by a pro-competive database I understand a database that is designed to include dissenting entries or competing descriptions, i.e., entries that share a common basic datum to be described, but disgree on the way it is described or categorized. This is to be contrasted with apodictic databases, where no basic datum can possibly be assigned more than one category or description.
Style is embodied with different meanings in the eyes of different stylists. This paper concentrates on one of the views on style,deviation.Style is deviation of the norm which helps achieve the effect of foregrounding.Some specific examples are to show that foregrounding is prominence that is motivated.
Scientific Reasoning in Day-to-Day Research Janet Bond-Robinson (jrobinso@ku.edu) Amy Preece Stucky (apreece@ku.edu) University of Kansas, 1251 Wescoe Hall Drive, 2010 Malott Hall, Lawrence, KS 66045 USA Introduction proposals; each project proceeds from a different foundational molecule, however uses similar techniques, equipment, and instruments to perform chemical reactions. Long series of reactions and what makes them work (a mechanical system) lead to a molecule engineered to possess specific and valuable properties. Problems punctuate researchers’ progress. We define a problem as a difficulty when the issue shows a basic lack of understanding of the process or inability to get the mechanical system working whereas an anomaly is an unexpected and therefore, problematic, piece of evidence. Experience with COP problems inspires integration of explicit declarative knowledge of chemical properties and mechanisms with functional procedural knowledge, whose product is often tacit expertise. Scientific reasoning is instantiated as “street smarts” developed in a specific research COP where reasoning: (a) Is guided by expectations of the organic synthesis COP’s norms and standards (constraints) while researchers do valued COP work. (b) Leads to and develops further apprentices’ learning in what to notice, understand, and take advantage of in terms of physical, human, and disciplinary COP resources (affordances). (c) Determines causal interactions of relevant variables in a mechanical system causing difficulties. (d) Is learning how to interpret the COP’s typical kinds of evidentiary formats in feedback because evidence is often evident only to COP members. (e) Recognizes anomalies in feedback. (f) Deciphers and explains anomalies. Klahr and Simon (1999) identified four approaches to scientific studies of science emerging in recent decades: (a) Historical accounts of scientific advances, (b) psychological experiments of non-scientists on structured and ill-structured problems, (c) observations of researchers’ daily work in science, and (d) computational modeling of scientific discovery processes. Our study fits as (c) observations of daily work in organic synthesis laboratories as others have done in biomechanical engineering (Nersessian, et al, 2002) and molecular biology (Dunbar, 1995). We expect to develop a grounded theory (Glaser & Strauss, 1967) of scientific reasoning within a community of practice (COP). Theoretical Framework & Methodology Cognitive apprenticeship is situated learning within a proficient COP through each participant’s immersion with frequent opportunities for practice, reflection and discussion while pursuing goals (Lave & Wenger, 1991). When Dunbar studied four different laboratories, all four COPs practicing molecular biology reasoned very similarly, i.e., similar experimental heuristics, mental representations, and problem solving heuristics, and differed only in their own combinations of these features. He noted that researchers interacted with the COP’s domain knowledge and fellow researchers to reduce reasoning errors. Logic in scientific reasoning requires substantial leaps from the data to infer conclusions (Toulmin, 1977). Toulmin explains that each field (COP) has different things to reason about, different consequences to gauge, and thus, different criteria for justifying inferred conclusions. Thus, apprentices must learn COP-specific standards of justifiable reasoning. Video data collected included 80 hours of researchers working in the lab, gathering and interpreting data, interacting with mentors, and attending group meetings. Semi-structured interviews, field notes, and laboratory notebook pages supplemented video data. All COP data were analyzed for norms, practices and reasoning. Acknowledgements National Science Foundation Grant REC-0093319 References Dunbar, K. (1995). How scientists really reason: In R. J. Sternberg & J. E. Davidson (Eds.), The Nature of Insight Cambridge, MA: MIT Press. Glaser, B. G., & Strauss, A. L. (1967). The Discovery of Grounded Theory: Chicago: Aldine. Klahr, D., & Simon, H. (1999). Studies of scientific discovery: Complementary approaches and convergent findings. Psychological Bulletin, 125(5), 524-543. Lave, J., & Wenger, E. (1991). Situated learning: Legitimate peripheral participation. Cambridge: University Press. Nersessian, N. J., Newstetter, W. C., Kurz-Milcke, E., & Davies, J. (2002). A Mixed-method Approach to Studying Distributed Cognition in Evolving Environments. Proceedings of International Conference on Learning Sciences, Seattle. Toulmin, S. (l977) Uses of argument (updated ed.). Cambridge, UK. Cambridge University Press. Results & Conclusions We asked how scientific reasoning, is instantiated when apprentice researchers pursue their daily work towards Ph.D. “certification” as scientists. This organic COP synthesizes compounds for potential in treatment of diseases, e.g., HIV. The research director determines norms of distributed work from success in funding
Abstract. This paper describes experiments on using inductive machine learning to guide a deterinistic dependency parser for unrestricted natural language text. Using data from a small treebank of Swedish, an eager probabilistic learning algorithm is used to induce context-sensitive parse tables. Evaluation shows a significant improvement over the baseline, which uses a table without contextual information. 1
The primary data of many experimental studies of animal learning and performance consist of the times at which stimuli and reinforcers were delivered, and the times at which responses occurred. The articles based on most of these studies report selected data, either from some sessions or some animals, or summary measures of the animals’ behavior. The primary data are sufficient to produce any of the selected and summary measures, but the selected and summarized data cannot produce many of the measures used in other experimental reports. It is now feasible to archive the primary data from animal behavior experiments so that they are accessible for others to perform secondary analysis. The value of such secondary analysis of archived data is described with a case study in which rats were trained on three fixed-interval schedules of reinforcement. The full data set may be downloaded fromwww.psychonomic.org/archive/.
I shall explore the implications for lexical resources of my work on ATT-Meta, a reasoning system designed to work out the signicance of a broad class of metaphorical utterances. This class includes imap-transcendingi utterances, resting on familiar, general conceptual metaphors but go beyond them by including source- domain elements that are not handled by the mappings in those metaphors. The system relies heavily on doing reasoning within the terms of the source domain rather than trying to construct new mapping relationships to handle the unmapped source-domain elements. The approach would therefore favour the use of WordNet-like resources that facilitate rich within-domain reasoning and the retrieval of known cross-domain mappings without being constrained to facilitate the creation of new mappings. The approach also seeks to get by with a small number of very general mappings per conceptual metaphor. The research has also led me to a radical language-user-relative view of metaphor. The question of whether an utterance is metaphorical, what conceptual metaphors it involves, what mappings those metaphors involve, what word-senses are recorded in a lexicon, etc. are all relative to specic language users and shouldn’t be regarded as something we have to make objective decisions about. This favours a practical approach where natural language applications can differ widely on how they handle the same potentially metaphorical utterance because of differences in lexical resources used. The user-relativity is also friendly to a view where the presence of a word-sense in a lexicon has little to do with whether that sense is gurati ve or not. This stance is related to, Patrick Hanks’ view that we should focus on norms and exploitations rather than on gurati vity. The research has furthermore led me to a deep scepticism about the ability to rely in denitions of metaphor on qualitative differences between domains. Scepticism about domains then causes additional difculty in distinguishing between metaphor and metonymy. At the panel I will outline a particular view of the distinction.
This paper deals with the relationship between rules and conventionality, as reflected in a case study of the two structures N 1 de N 2 and N 1 du N 2. Most occurrences of these can be explained via two rules which evoke ease of referent identification and are a subset of the principles governing the ± definiteness opposition in French. Via dictionaries and Google searches, however, we detect variation that our rules cannot predict, typically when N 2 is non-countable and abstract. Local semantic contrasts, i.e. dependent on the lexical content of N 1 or N 2, further complicate the picture. We conclude that Coseriu's Norm, alias conventionality, must be recognized as a factor coexisting with the rules.
The aim of this paper is to present the theoretical principles underlying the making of a lexical database of English collocations of non-specialized words used in scientific language. This project was prompted by the shortage of reference tools providing information about the use and combinatorial properties of general words in specific registers. A case study will illustrate that in scientific texts, words, especially polysemous verbs, have a distinct semantic and combinatorial behaviour. Following the assumption that the meaning and the grammatical and collocational patterns of words are interrelated, we suggest that context-specific information should be included in specialized reference tools to facilitate the written production of scientific texts by nonnative speakers ofEnglish.
The development of wide-coverage grammars is at the core of robust NLP systems.This paper addresses the problem of grammar extraction from treebanks with respect to the issue of broad coverage along three dimensions: the grammar formalism (contextfree grammar, dependency grammar, lexicalized tree adjoining grammar), the domain of the annotated corpus (press reports, civil law) and the language of the corpus (English, Korean, Chinese, Italian).We have extracted three grammars from an annotated corpus of Italian and we have comparatively analyzed the coverage of a test set; then, working on two different domain subcorpora we have compared the cross-domain coverage of the extracted grammars; finally, we have compared the grammars for four different languages.The results are that there are relevant differences in coverage among formalisms and domains; a more limited difference appears in the crosslinguistic comparison.
Abstract. Syntactic disambiguators for natural language often use ”Treebank Grammars”: probabilistic grammars which are directly projected from an annotated corpus. In this paper we show that for describing these systems in the framework of Estimation Theory, we must generalize this theory so that it allows for an infinite number of parameters. Embracing this generalization will also bring the justification of statistical smoothing techniques within the scope of Estimation Theory. 1
Starting from the analysis of the verbal entries in the letter A of An Anglo-Saxon Dictionary (Bosworth and Toller 1973), this paper establishes the descriptive bases necessary for the elaboration of a lexical database of Old English. The basic terms (morphological process, derivational chain, inflexion, compounding, derivation, paradigm and root) are defined, the units (affixal and non-affixal predicates, primitive predicates and derived predicates) are classified and a methodology for the identification of the affixal predicates based on distribution and decomposition is put forward. Conclusions are drawn about the nature of derivational chains, the direction of derivation, and the distinction between inflexion and derivation, on the one hand, and compounding and derivation on the other.
The American poetess Emily Dickinson in the 19th century is well-known for her original imagery and unbound language style. Her discarding of the linguistic norms makes her verse darting and whimsical as well as sophisticated to understand. This paper is to interpret the mysterious poetess by means of exploring the linguistic deviations in her poetry. The two deviations that will be focused on are grammatical deviation and graphological deviation.
Statistics from about 17,000 occurrences of the structures “N1 is N1” and “N1 is N7” have proved that (a) there is a functional difference between the two predicative cases and (b) there are strong norms for selecting one of the two cases in communication. The nominative case is a strong norm if the communicative function of the sentence predicate is to (a) identify a sort of denotation in a demonstrative act, (b) identify a sort of denotation in a nominative act, (c) define using qualification or (d) qualify in an expressive manner. The nominative is preferred when stating a person’s profession, in most cases (except for professions such as minister, director, manager etc., especially in a specific sentence structure and discourse function). The statistics show that (a) for 630 different predicate nouns (PN), only the nominative is used and for 270 PN, only the instrumental is used; (b) for 900 PN, one of the two predicative cases is used either exclusively or with a strong preference; (c) for only 89 PN, the use of the two cases is balanced. The corpus statistics for the two predicative cases show that the selection of one of these cases is semantically determined and to a great degree lexically bound.
Abstract. An important aspect of discourse understanding and generation involves the recognition and processing of discourse relations. These are conveyed by discourse connectives, i.e., lexical items like because and as a result or implicit connectives expressing an inferred discourse relation. The Penn Discourse TreeBank (PDTB) provides annotations of the argument structure, attribution and semantics of discourse connectives. In this paper, we provide the rationale of the tagset, detailed descriptions of the senses with corpus examples, simple semantic definitions of each type of sense tags as well as informal descriptions of the inferences allowed at each level. 1
In the design of a Multilingual Lexical Database, one of the biggest problems is constituted by conceptual mismatches between languages, and the resulting matter of lexical gaps. Lexical gaps concern words for which there is no direct translation in a target language, but which nonetheless need to receive a translation within the system. In this article, it will be shown that the various possible ways of dealing with these lexical gaps can be classified in four basic groups. Using the SIM<it>u</it>LLDA system as an example (Janssen 2002), the advantages of the structured interlingua approach over the other possibilities will be explained. With the SIM<it>u</it>LLDA set-up, it is possible to derive correct lexical definitions for lexical gaps from the lexical database. How this process of “lexical gap filling” works will be shown using a concrete example of a lexical gap: the treatment of the English words <it>river</it> and <it>stream </it>in contrast with the French words <it>fleuve</it> and <it>rivière</it>.
The purpose of this paper is to describe the TuBa-D/Z treebank of written German and to compare it to the independently developed TIGER treebank (Brants et al., 2002). Both treebanks, TIGER and TuBa-D/Z, use an annotation framework that is based on phrase structure grammar and that is enhanced by a level of predicate-argument structure. The comparison between the annotation schemes of the two treebanks focuses on the different treatments of free word order and discontinuous constituents in German as well as on differences in phrase-internal annotation.
This paper is a contribution to the issue -- which has, in the course of the last decade, become critical -- of the basic requirements and validation criteria for lexical language resources in Standard Arabic. The work is based on a critical analysis of the architecture of the DIINAR.1 lexical database, the entries of which are associated with grammar-lexis relations operating at word-form level (i.e. in morphological analysis). Investigation shows a crucial difference, in the concept of 'lexical database', between source program and generated lexica. The source program underlying DIINAR.1 is analysed, and some figures and ratios are presented. The original categorisations are, in the course of scrutiny, partly revisited. Results and ratios given here for basic entries on the one hand, and for generated lexica of inflected word-forms on the other. They aim at giving a first answer to the question of the ratios between the number of lemma-entries and inflected word-forms that can be expected to be included in, or generated by, a Standard Arabic lexical dB. These ratios can be considered as one overall language-specific criterion for the analysis, evaluation and validation of lexical dB-s in Arabic.
We present the architectural design rationale of a Sanskrit computational linguistics platform, where the lexical database has a central role. We explain the structuring requirements issued from the interlinking of grammatical tools through its hypertext rendition.
In the present paper we discuss some issues connected with the condition of projectivity in a dependency based description of language (see Sgall, Hajičová, and Panevová (1986), Hajičová, Partee, and Sgall (1998)), with a special regard to the annotation scheme of the Prague Dependency Treebank (PDT, see Hajič (1998)). After a short Introduction (Section 1), the condition of projectivity is discussed in more detail in Section 2, presenting its formal definition and formulating an algorithm for testing this condition on a subtree (Section 2.1); the introduction of the condition of projectivity in a formal description of language is briefly substantiated in Section 2.2. and some problematic cases are discussed in Section 2.3. In Section 3, a preliminary classification into three main groups and several subgroups of Czech non-projective constructions on the analytical level is presented (Section 3.1),
Word-to-word dependency structures are useful for consistent representation and comparable evaluation of parsing results. However, most large-scale treebanks contain various variants of phrase structure trees, since automatic parsers usually produce constituent struc-tures. We present a freely available extensible tool for converting phrase structure to dependencies automatically, and discuss its appli-cation to the NEGRA treebank of German. 1.
This paper describes a method for conducting evaluations of Treebank and non-Treebank parsers alike against the English language U. Penn Treebank (Marcus et al., 1993) using a metric that focuses on the accuracy of relatively non-controversial aspects of parse structure. Our conjecture is that if we focus on maximal projections of heads (MPH), we are likely to find much broader agreement than if we try to evaluate based on order of attachment. We hope that this method may find wider acceptance and be useful in establishing a generally applicable framework for evaluation in natural language parsing. We employ this method in an evaluation of NLPWin (Heidorn, 2000), a parser developed at Microsoft Research without reference to the Penn Treebank, and, for comparison, the well-known statistical Treebank parser of Charniak (2000). 1.
Nous décrivons l’extraction d’une grammaire d’arbres adjoints (TAG) à partir d’une banque d’arbres de l’arabe écrit. Nous montrons queleques exemples d’arbres élémentaires ainsi obtenus, et les structures de dérivation (donc, de dépendence syntaxique) qui y correspondent. We describe the extraction of a Tree Adjoining Grammar from the Penn Arabic Treebank.1 We show some examples of extracted trees for different constructions, and the corresponding derivation structures (which represent syntactic dependency). 1
Among the variety of proposals currently making the dependency perspective on grammar more concrete, there are several treebanks whose annotation exploits some form of Relational Structure that we can consider a generalization of the fundamental idea of dependency at various degrees and with reference to different types of linguistic knowledge. The paper describes the Relational Structure as the common underlying representation of treebanks which is motivated by both theoretical and task-dependent considerations. Then it presents a system for the annotation of the Relational Structure in treebanks, called Augmented Relational Structure, which allows for a systematic annotation of various components of linguistic knowledge crucial in several tasks. Finally, it shows a dependency-based annotation for an Italian treebank, i.e. the Turin University Treebank, that implements the Augmented Relational Structure. 1
Scaling wide-coverage, constraint-based grammars such as Lexical-Functional Grammars (LFG) (Kaplan and Bresnan, 1982; Bresnan, 2001) or Head-Driven Phrase Structure Grammars (HPSG) (Pollard and Sag, 1994) from fragments to naturally occurring unrestricted text is knowledge-intensive, time-consuming and (often prohibitively) expensive. A number of researchers have recently presented methods to automatically acquire wide-coverage, probabilistic constraint-based grammatical resources from treebanks (Cahill et al., 2002, Cahill et al., 2003; Cahill et al., 2004; Miyao et al., 2003; Miyao et al., 2004; Hockenmaier and Steedman, 2002; Hockenmaier, 2003), addressing the knowledge acquisition bottleneck in constraint-based grammar development. Research to date has concentrated on English and German. In this paper we report on an experiment to induce wide-coverage, probabilistic LFG grammatical and lexical resources for Chinese from the Penn Chinese Treebank (CTB) (Xue et al., 2002) based on an automatic f-structure annotation algorithm. Currently 96.751% of the CTB trees receive a single, covering and connected f-structure, 0.112% do not receive an f-structure due to feature clashes, while 3.137% are associated with multiple f-structure fragments. From the f-structure-annotated CTB we extract a total of 12975 lexical entries with 20 distinct subcategorisation frame types. Of these 3436 are verbal entries with a total of 11 different frame types. We extract a number of PCFG-based LFG approximations. Currently our best automatically induced grammars achieve an f-score of 81.57% against the trees in unseen articles 301-325; 86.06% f-score (all grammatical functions) and 73.98% (preds-only) against the dependencies derived from the f-structures automatically generated for the original trees in 301-325 and 82.79% (all grammatical functions) and 67.74% (preds-only) against the dependencies derived from the manually annotated gold-standard f-structures for 50 trees randomly selected from articles 301-325.
In this paper we present a methodology for extracting subcategorisation frames based on an automatic LFG f-structure annotation algorithm for the Penn-II Treebank. We extract abstract syntactic function-based subcategorisation frames (LFG semantic forms), traditional CFG category-based subcategorisation frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach does not predefine frames, associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. We extract 3586 verb lemmas, 14348 semantic form types (an average of 4 per lemma) with 577 frame types. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource.