Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The current research explored the processes that predominate during the anticipation of an emotionally salient event. Experiment 1 (N536), employed three different conditional stimuli followed by pictorial pleasant, unpleasant or neutral unconditioned stimuli. Half the participants were trained with visual CSs, the other half with tactile CSs. In the group trained with visual CSs, startle eyeblinks were larger and faster during CSs that were paired with unpleasant pictures than CSs paired with neutral or pleasant pictures respectively, indicating an affect startle pattern. This linear trend was not found in the group trained with tactile CSs. Experiment 2 (N564) aimed to investigate whether the affective pattern found in the startle data in Experiment 1 could also be found using a behavioural measure of emotion. This time participants’ reaction time during a post-experimental affective priming taskwas used as dependantmeasure to assess the presence of emotional learning. Instead of a simple differential conditioning task, an occasion setting paradigm was employed and participants were trained using either a feature positive or feature negative design with pleasant or unpleasant picture USs. For participants trained with unpleasant USs, valence ratings collected before and after conditioning training suggested the presence of emotional learning, whereas no such pattern was found for participants trained with pleasant USs. These findings were not confirmed in the priming data.
We investigated subjective and hemodynamic responses towards disgust-inducing, fear-inducing, and neutral pictures in a functional magnetic resonance imaging study. Within an interval of 1 week, 24 male subjects underwent the same block design twice in order to analyze possible response changes to the repeated picture presentation. The results showed that disgust-inducing and fear-inducing scenes provoked a similar activation pattern in comparison to neutral scenes. This included the thalamus, primary and secondary visual fields, the amygdala, the hippocampus, and various regions of the prefrontal cortex. During the retest, the affective ratings hardly changed. In contrast, most of the previously observed brain activations disappeared, with the exception of the temporo-occipital activation. An additional analysis, which compared the emotion-related activation patterns during the two presentations, showed that the responses to the fear-inducing pictures were more stable than the responses to the disgust-inducing ones.
A statistical estimator attempts to guess an unknown probability distribution by analyzing a sample from this distribution. One desirable property of an estimator is that its guess is increasingly likely to get arbitrarily close to the actual distribution as the sample size increases. This property is called consistency. Data Oriented Parsing (DOP) employs all fragments of the trees in a training treebank, including the full parse-trees themselves, as the rewrite rules of a probabilistic tree-substitution grammar. Since the most popular DOP-estimator (DOP1) was shown to be inconsistent, there is an outstanding theoretical question concerning the possibility of DOP-estimators with reasonable statistical properties. This question constitutes the topic of the current paper. First, we show that, contrary to common wisdom, any unbiased estimator for DOP is futile because it will not generalize over the training treebank. Subsequently, we show that a consistent estimator that generalizes over the treebank should involve a local smoothing technique. This exposes the relation between DOP and existing memory-based models that work with full memory and an analogical function such as k-nearest neighbor, which is known to implement backoff smoothing. Finally, we present a new consistent backoff-based estimator for DOP and discuss how it combines the memory-based preference for the longest match with the probabilistic preference for the most frequent match.
We describe the automatic conversion of English Penn Treebank (PTB) annotations into Language Neutral Syntax (LNS) (Campbell and Suzuki, 2002a,b). In this paper, we describe LNS and why it is useful, describe the conversion algorithm, present an evaluation of the conversion, and discuss some uses of the converted annotations and the potential for extending the coverage to other languages. The work described here is in the spirit of other automatic re-annotations of PTB trees (e.g. Frank, 2000 and Meyers, 2001), but differs in the nature of the output.
This paper discusses an annotation scheme for Korean null pronouns, which were used in annotating three kinds of Korean text corpora including Penn Korean Treebank. In annotating the corpora, null pronouns and their antecedents were marked up for their type and reference, with coreference relation tracked by numeric identifiers. Based on the annotation scheme, an outline of a potential pronoun resolution strategy is also proposed. The resulting dataset of annotated text is rather small at 11,834 words; we hope the null pronoun classification and annotation scheme proposed in this study will serve as a basis in developing a large-scale annotated corpus in the future.
Abstract This study attempted to demonstrate an elevated disgust sensitivity in bulimia nervosa. Eleven bulimic patients and 12 control subjects underwent a functional magnetic resonance imaging (fMRI) study in which they were presented with alternating blocks of 40 disgust‐inducing, 40 fear‐inducing and 40 affectively neutral scenes. Each scene was shown for 1.5 s. After completion of all blocks, affective ratings were then determined. The viewing of the disgusting pictures, which had been rated as highly repulsive by the bulimic females, was associated with an activation of the left amygdala and the occipito‐temporal visual cortex. The subjective and brain‐physiological responses did not differ from those of the healthy control subjects. This held true for the fear‐inducing scenes as well. Thus, bulimic patients are not characterized by an increased global disgust sensitivity and they do not show any indication of an altered central processing of generally disgust and fear‐inducing visual stimuli. Copyright © 2004 John Wiley & Sons, Ltd and Eating Disorders Association.
Abstract. An important aspect of discourse understanding and generation involves the recognition and processing of discourse relations. These are conveyed by discourse connectives, i.e., lexical items like because and as a result or implicit connectives expressing an inferred discourse relation. The Penn Discourse TreeBank (PDTB) provides annotations of the argument structure, attribution and semantics of discourse connectives. In this paper, we provide the rationale of the tagset, detailed descriptions of the senses with corpus examples, simple semantic definitions of each type of sense tags as well as informal descriptions of the inferences allowed at each level. 1
In this paper, we consider the problem of enriching a Thai lexical database by extending the semantic information with selectional preferences. We propose a novel approach for acquiring selectional preferences of verbs, which is motivated by the tree cut model. We apply a model selection technique called the Bayesian Information Criterion (BIC). Given a semantic hierarchy, our goal is to generalize initial noun classes to the most plausible levels on that hierarchy. We present an iterative algorithm for generalization. The algorithm performs agglomerative merging on the semantic hierarchy in a bottomup manner. The BIC is used to measure the improvement of the model both locally and globally. In our experiments, we consider the Web as large corpora. We also propose approaches for extracting examples from the Web. Preliminarily experimental results are given to show the feasibility and effectiveness of our approach.
The present work falls in the line of activities promoted by the European Languguage Resource Association (ELRA) Production Committee (PCom) and raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources. A two-fold purpose lies behind this experiment. The first aim is to investigate the feasibility, define methods and procedures for combining two Italian lexical resources that have incompatible formats and complementary information into a Unified Lexicon (UL). The adopted strategy and the procedures appointed are described together with the driving criterion of the merging task, where a balance between human and computational efforts is pursued. The coverage of the UL has been maximized, by making use of simple and fast matching procedures. The second aim is to exploit this newly obtained resource for implementing the phonological and morphological layers of the CLIPS lexical database. Implementing these new layers and linking them with the already exisitng syntactic and semantic layers is not a trivial task. The constraints imposed by the model, the impact at the architectural level and the solution adopted in order to make the whole database ‘speak ’ efficiently are presented. Advantages vs. disadvantages are discussed. 1. Background and Motivations The work described here raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources (LRs) and has been
We present discourse-level annotation of newspaper texts in German and English, as part of an ongoing project aimed at investigating information structure from a cross-linguistic perspective. Rather than annotating some specific notion of information structure, we propose a theory-neutral annotation of basic features at the levels of syntax, prosody and discourse, using treebank data as a starting point. Our discourse-level annotation scheme covers properties of discourse referents (e.g., semantic sort, delimitation, quantification, familiarity status) and anaphoric links (coreference and bridging). We illustrate what investigations this data serves and discuss some integration issues involved in combining different levels of stand-off annotations, created by using different tools.
Rating agencies' track record is good in developed countries but poor in emerging economies. Why? Given the almost-monopolistic structure of the industry, we conjecture that agencies might underinvest in information gathering. We propose an indicator quantifying the agencies' effort to gather information and assess whether greater effort affects rating levels. We detect: (i) absolute underinvestment for non-OECD sovereigns (less effort in spite of greater opaqueness); (ii) relative underinvestment for non-OECD firms compared with OECD ones (though the former receive a larger effort, more intense effort boosts firm ratings in non-OECD countries while depressing them in OECD countries).
The aim of this paper is to present the theoretical principles underlying the making of a lexical database of English collocations of non-specialized words used in scientific language. This project was prompted by the shortage of reference tools providing information about the use and combinatorial properties of general words in specific registers. A case study will illustrate that in scientific texts, words, especially polysemous verbs, have a distinct semantic and combinatorial behaviour. Following the assumption that the meaning and the grammatical and collocational patterns of words are interrelated, we suggest that context-specific information should be included in specialized reference tools to facilitate the written production of scientific texts by nonnative speakers ofEnglish.
We present a relational learning framework for grammar induction that is able to learn meaning as well as syntax. We introduce a type of constraint-based grammar, lexicalized well-founded grammar (lwfg), and we prove that it can always be learned from a small set of semantically annotated examples, given a set of assumptions. The semantic representation chosen allows us to learn the constraints together with the grammar rules, as well as an ontology-based semantic interpretation. We performed a set of experiments showing that several fragments of natural language can be covered by a lwfg,and that it is possible to choose the representative examples heuristically, based on linguistic knowledge.
In this paper, a hybrid language model is defined as a combination of a word-based <i>n</i>-gram, which is used to capture the local relations between words, and a category-based stochastic context-free grammar (SCFG) with a word distribution into categories, which is defined to represent the long-term relations between these categories. The problem of unsupervised learning of a SCFG in General Format and in Chomsky Normal Form by means of estimation algorithms is studied. Moreover, a bracketed version of the classical estimation algorithm based on the Earley algorithm is proposed. This paper also explores the use of SCFGs obtained from a treebank corpus as initial models for the estimation algorithms. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment.
This paper describes a method for conducting evaluations of Treebank and non-Treebank parsers alike against the English language U. Penn Treebank (Marcus et al., 1993) using a metric that focuses on the accuracy of relatively non-controversial aspects of parse structure. Our conjecture is that if we focus on maximal projections of heads (MPH), we are likely to find much broader agreement than if we try to evaluate based on order of attachment. We hope that this method may find wider acceptance and be useful in establishing a generally applicable framework for evaluation in natural language parsing. We employ this method in an evaluation of NLPWin (Heidorn, 2000), a parser developed at Microsoft Research without reference to the Penn Treebank, and, for comparison, the well-known statistical Treebank parser of Charniak (2000). 1.
This paper presents a prosodic phrasing model for Korean to be used in a text-to-speech synthesis (TTS) system. Read text corpora were morpho-syntactically parsed and prosodically labeled following the Penn Korean Treebank (Han, Chunghye, Ko, Eon-Suk, Yi, Heejong, Palmer, M., 2002. Penn Korean Treebank: development and evaluation. In: Proceedings of the 16th Pacific Asian Conference on Language and Computation. Korean Society for Language and Information.) and K-ToBI prosodic labeling conventions (Sun-Ah, J., 2000. K-ToBI (Korean ToBI) labelling conventions. Version 3.1. Available from: URL.), respectively. Decision trees were trained with morpho-syntactic and textual distance features to predict locations of accentual and intonational phrase breaks. Our phrasing model cross-validated on a 300-sentence corpus (6936 words or 21,436 syllables, with an average of 72 syllables or 23 words per sentence) predicted non-breaks with F=92.4% and breaks with F=88.0% (F=72.8% for accentual phrase breaks and F=71.3% for intonational phrase breaks).
In this paper we will present work carried out on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples – Federico II. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. We will also present the linguistic annotation tools needed to allow the parser to produce syntactic structures automatically.\nIn particular, in order to produce appropriate linguistic annotations, all transcribed materials need to be transliterated from the audio-transcription to a more standard orthographic format. In that way, the parser receives as an input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts “overlap”, i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts or backchannels the current speaker. This phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically both to a following turn and to the local turn where it may produce conversational moves to repair what has been previously said by the current speaker.
This paper explores the interaction between conceptual structure and morpho-syntax. In particular, we show that ontology-based conceptual classification can be used to predict internal relations in compounds. We propose an ontology-based approach to predict the semantic relation between the two component words in Mandarin VV compounds. A Mandarin VV compound is classified according to the eventive relation between the two simplex verbs. These relations specify how the eventive meanings of the two simplex verbs combine to form the meaning of the compound. The three types of eventive relations that we deal with in this paper are: coordinate, modificational, and resultative. Since the way in which two events combine with each other depends upon their event types, we hypothesize that the eventive relations can be predicted by the conceptual classified event types of the two simplex verbs. An approach of ontology-based prediction is proposed based on this hypothesis. The assignment of ontology classification for each simplex verb is based on SUMO and Sinica BOW. The correlation between the ontology class of each verb position and each eventive type is trained and scored based on a manually tagged lexical database. We encode the ontology information of each VV compound in a 3-tuple based on these correlation scores. This 3-tuple is represented as a three-dimensional vector and used to predict the eventive type of new VV compounds. Our classification experiment on unknown VV compounds yields good recall and precision. 1.
This report presents an approach to enriching flat and robust predicate argument structures with more fine-grained semantic information, extracted from underspecified semantic representations and encoded in Minimal Recursion Semantics (MRS). Such representations are provided by a hand-built HPSG grammar with a wide linguistic coverage. A specific semantic representation, called linked predicate argument structure (LPAS), has been worked out, which describes the explicit embedding relationships among predicate argument structures. LPAS can be used as a generic interface language for integrating semantic representations with different granularities. Some initial experiments have been conducted to convert MRS expressions into LPASs. A simple constraint solver is developed to resolve the underspecified dominance relations between the predicates and their arguments in MRS expressions. LPASs are useful for high-precision information extraction and question answering tasks because of their fine-grained semantic structures. In addition, I have attempted to extend the lexicon of the HPSG English Resource Grammar (ERG) exploiting WordNet and to disambiguate the readings of HPSG parsing with the help of a probabilistic parser, in order to process texts from application domains. Following the presented approach, the HPSG ERG grammar can be used for annotating some standard treebank, e.g., the Penn Treebank, with its fine-grained semantics. In this vein, I point out opportunities for a fruitful cooperation of the HPSG annotated Redwood Treebank and the Penn PropBank. In my current work, I exploit HPSG as an additional knowledge resource for the automatic learning of LPASs from dependency structures.
Function tags are a context-sensitive annotation applied to words and phrases of natural language text, marking their syntactic or semantic role within a larger utterance. As researchers improve results on various other problems in “pure” natural language processing (e.g part-of-speech tagging, parsing), those who work in the more “applied” NLP fields (e.g. question-answering, temporal analysis) are seeking more powerful sorts of linguistic annotation as input for their own systems. Hence, function tags. In the first part of the thesis, I present the problem of function tagging: why it is an interesting problem, who has worked on similar thing, and what exactly I intend to do. I briefly review the function tags of the Penn treebank, and explain the specific metrics by which I will evaluate my work. In the second part of the thesis, I introduce the many features that I will use to train a function tagging system, and then I present some systems that make use of them: one using feature trees, one using decision trees (briefly), and one using perceptron models. For each system, I give a brief historical perspective, an overview of where it has been used before and why I think it will be useful in this task. I will then try a number of feature combinations with interesting properties; and finally, present the best-performing tweaked-out version of that system. Finally, in the third part of the thesis, I bring them all together and discuss the advantages and disadvantages of each system in various situations. More interestingly, I will present an analysis of what features prove to be the most helpful for the different function tagging subtasks. Lastly, I will present a comparison to other systems performing related tasks, and speculate on some interesting future work.
We present a linguistically-motivated algorithm for reconstructing nonlocal dependency in broad-coverage context-free parse trees derived from treebanks. We use an algorithm based on loglinear classifiers to augment and reshape context-free trees so as to reintroduce underlying nonlocal dependencies lost in the context-free approximation. We find that our algorithm compares favorably with prior work on English using an existing evaluation metric, and also introduce and argue for a new dependency-based evaluation metric. By this new evaluation metric our algorithm achieves 60% error reduction on gold-standard input trees and 5% error reduction on state-of-the-art machine-parsed input trees, when compared with the best previous work. We also present the first results on non-local dependency reconstruction for a language other than English, comparing performance on English and German. Our new evaluation metric quantitatively corroborates the intuition that in a language with freer word order, the surface dependencies in context-free parse trees are a poorer approximation to underlying dependency structure.
Word-to-word dependency structures are useful for consistent representation and comparable evaluation of parsing results. However, most large-scale treebanks contain various variants of phrase structure trees, since automatic parsers usually produce constituent struc-tures. We present a freely available extensible tool for converting phrase structure to dependencies automatically, and discuss its appli-cation to the NEGRA treebank of German. 1.
In this paper we describe on-going work aimed at creating a dependency-based annotated treebank for the BioMedical domain. Our starting point is the GENIA corpus, which is a corpus of 2000 MEDLINE abstracts, which has been manually annotated for various biological entities, according to the GENIA Ontology. There is an exponential growth of published research in this sector, which makes it difficult even for the experts to follow the recent developments. This creates the need for tools that can automatically process the research literature and extract only relevant information, such as interactions between genes and proteins. In order for these tools to be developed, annotated resources, such as corpora and Treebanks are of fundamental importance. Such resources will support the development of practical domain-specific information extraction tools.
We describe the development of a Dutch memory-based shallow parser. The availability of large treebanks for Dutch, such as the one provided by the Spoken Dutch Corpus, allows memory-based learners to be trained on examples of shallow parsing taken from the treebank, and act as a shallow parser after training. An overview is given of a modular memory-based learning approach to shallow parsing, composed of a part-of-speech tagger-chunker and two grammatical relation finders, which has originally been developed for English. This approach is applied to the syntactically annotated part of the Spoken Dutch Corpus to construct a Dutch shallow parser. From the generalisation scores of the parser we conclude that existing memory-based parsing approaches can be applied to spoken Dutch successfully, but that there is room for improvement in the tagger-chunker
Schema matching is prerequisite to an automated transformation of XML documents. Because previous works about schema matching compute all semantically-possible matchings, they produce many-to-many matching relationships. Such imprecise matchings are inappropriate for an automated transformation of XML documents. This paper presents an efficient schema matching algorithm that computes precise one-to-one matchings between two schemas. The proposed algorithm consists of two steps: preliminary matching relationships between leaf nodes in the two schemas are computed and one-to-one matchings are finally extracted based on a proposed path similarity. Specifically, for a sophisticated schema matching, the proposed algorithm is based on a domain ontology as well as a lexical database that includes abbreviations and synonyms. Experimental results with real schemas from an e-commerce field show that the proposed method is superior to previous works, resulting in an accuracy of 97% in average.
Published in: Proceedings of the Third Workshop on Treebanks and Linguistic Theories (TLT 2004), Tübingen, December 10–11, 2004.
The majority of electronic data today is in textual form.Financial data such as articles in the Wall Street Journal are written as texts.These electronic documents contain a wealth of information but require human interpretation.For financial analysis, rapid up-to-date information is critical.Most software tools currently require data which are better structured than text (such as data in relational databases).Thus, our research goal is to build a system, "FIRST" (Flexible Information extRaction SysTem), that will extract data from financial articles and store the output in an explicit format.FIRST uses natural language processing techniques and resources such as the lexical database WordNet and collocation information to extract information.We hope to be able to extract data such as an organization's name, its profit/loss status, and sales status, from financial articles to input into a database.The data will come from international corporate reports which appear in the Wall Street Journal.
An automatic method for annotating the Penn-II Treebank (Marcus et al., 1994) with high-level Lexical Functional Grammar (Kaplan and Bresnan, 1982; Bresnan, 2001; Dalrymple, 2001) f-structure representations is described in (Cahill et al., 2002; Cahill et al., 2004a; Cahill et al., 2004b; O’Donovan et al., 2004). The annotation algorithm and the automatically-generated f-structures are the basis for the automatic acquisition of wide-coverage and robust probabilistic approximations of LFG grammars (Cahill et al., 2002; Cahill et al., 2004a) and for the induction of LFG semantic forms (O’Donovan et al., 2004). The quality of the annotation algorithm and the f-structures it generates is, therefore, extremely important. To date, annotation quality has been measured in terms of precision and recall against the DCU 105. The annotation algorithm currently achieves an f-score of 96.57% for complete f-structures and 94.3% for preds-only \nf-structures. There are a number of problems with evaluating against a gold standard of this size, most \nnotably that of overfitting. There is a risk of assuming that the gold standard is a complete and balanced \nrepresentation of the linguistic phenomena in a language and basing design decisions on this. It is, therefore, \npreferable to evaluate against a more extensive, external standard. Although the DCU 105 is publicly available, \n1 a larger well-established external standard can provide a more widely-recognised benchmark against which the quality of the f-structure annotation algorithm can be evaluated. For these reasons, we present an evaluation of the f-structure annotation algorithm of (Cahill et al., 2002; Cahill et al., 2004a; Cahill et al., 2004b; O’Donovan et al., 2004) against the PARC 700 Dependency Bank (King et al., 2003). Evaluation against an external gold standard is a non-trivial task as linguistic analyses may differ systematically between the gold standard and the output to be evaluated as regards feature geometry and nomenclature. We present conversion software to automatically account for many (but not all) of the systematic differences. Currently, we achieve an f-score of 87.31% for the f-structures generated from the original Penn-II trees and \nan f-score of 81.79% for f-structures from parse trees produced by Charniak’s (2000) parser in our pipeline \nparsing architecture against the PARC 700.
OBJECTIVE: To validate a culturally relevant body image instrument among urban African Americans through three distinct studies. RESEARCH METHODS AND PROCEDURES: In Study 1, 38 medical practitioners performed content validity tests on the instrument. In Study 2, three research staff rated the body image of 283 African-American public housing residents (75% women, mean age = 44 years), with the residents completing body image, BMI, and percentage body fat measures. In Study 3, 35 African Americans (57% men, mean age = 42) completed body image measures and evaluated their cultural relevance. RESULTS: In Study 1, 97% to 100% of practitioners sorted the jumbled figures into the correct ascending order. The correlation between the body image figures and the practitioners' weight classifications of the figures was high (r = 0.91). In Study 2, observers arrived at similar ratings of body size with excellent consistency (alpha = 0.95). Ratings of body image were strongly correlated with participant BMI (r = 0.89 to 0.93 across observers and 0.81 for all participants) and percentage of body fat (r = 0.77 to 0.89 across observers and 0.76 for all participants). In Study 3, body image ratings with the new scale were positively correlated with other validated figural scales. The majority of participants reported that figures in the new body image scale looked most like themselves and other African Americans and were easiest to identify themselves with. DISCUSSION: The instrument displayed strong psychometric performance and cultural relevance, suggesting that the scale is a promising tool for examining body image and obesity among African Americans.
This paper investigates the usefulness of sentence-internal prosodic cues in syntactic parsing of transcribed speech. Intuitively, prosodic cues would seem to provide much the same information in speech as punctuation does in text, so we tried to incorporate them into our parser in much the same way as punctuation is. We compared the accuracy of a statistical parser on the LDC Switchboard treebank corpus of transcribed sentence-segmented speech using various combinations of punctuation and sentence-internal prosodic information (duration, pausing, and f0 cues).
In the design of a Multilingual Lexical Database, one of the biggest problems is constituted by conceptual mismatches between languages, and the resulting matter of lexical gaps. Lexical gaps concern words for which there is no direct translation in a target language, but which nonetheless need to receive a translation within the system. In this article, it will be shown that the various possible ways of dealing with these lexical gaps can be classified in four basic groups. Using the SIM<it>u</it>LLDA system as an example (Janssen 2002), the advantages of the structured interlingua approach over the other possibilities will be explained. With the SIM<it>u</it>LLDA set-up, it is possible to derive correct lexical definitions for lexical gaps from the lexical database. How this process of “lexical gap filling” works will be shown using a concrete example of a lexical gap: the treatment of the English words <it>river</it> and <it>stream </it>in contrast with the French words <it>fleuve</it> and <it>rivière</it>.
The requirements of the depth and precision of annotation vary for different intended uses of the corpus but it has been commonly accepted nowadays that the standard annotations of surface structure are only the first steps in a more ambitious research program, aiming at a creation of advanced resources for most different systems of natural language processing and for testing and further enrichment of linguistic and computational theories. Among the several possible directions in which we believe the standard annotation systems should go (and in some cases already attempt to go) beyond the POS tagging or shallow syntactic annotations, the following four are characterized in the present contribution: (i) predicateargument representation of the underlying syntactic relations as basically corresponding to a rooted tree that can be univocally linearized, (ii) the inclusion of the information structure using very simple means (the left-to-right order of the nodes and three attribute values), (iii) relating this underlying structure (rendering the ”linguistic meaning,” i.e. the semantically relevant counterparts of the grammatical means of expression) to certain central aspects of referential semantics (reference assignment and coreferential relations), and (iv) handling of word sense disambiguation. The first three issues are documented in the present paper on the basis of our experience with the development of the structure and scenario of the Prague Dependency Treebank which provides for syntactico-semantic annotation of large text segments from the Czech National Corpus and which is based on a solid theoretical framework.
In natural language processing a huge amount of structured data is constantly used for the extraction and presentation of grammatical structures in sentences. For example the Chinese Treebank corpus developed at the Institute of Information Science Academia Sinica Taiwan is a semantically annotated corpus that has been used to help parse and study Chinese sentences. In this setting users usually use structured tree patterns instead of keywords to query the corpus.
Introduction The CorpusEye project (http://corp.hum.sdu.dk ) at the University of Denmark aims at designing and programming an internet based corpus search interface that (1) offers standardised search tools and a unified descriptive formalism across different corpus types and different languages, and (2) allows users to exploit grammatical information in annotated corpora in a user-friendly and menubased way. All corpora in CorpusEye have been annotated with VISL's Constraint Grammar based parsers, in the case of treebanks using an additional PSG module or equivalent (Bick 2003). At the time of writing, the material covers 8 languages and ca. 600 million words.
In this paper we address the following questions from our experience of the last two and a half years in developing a large-scale corpus of Arabic text annotated for morphological information, part-of-speech, English gloss, and syntactic structure: (a) How did we 'leapfrog' through the stumbling blocks of both methodology and training in setting up the Penn Arabic Treebank (ATB) annotation? (b) How did we reconcile the Penn Treebank annotation principles and practices with the Modern Standard Arabic (MSA) traditional and more recent grammatical concepts? (c) What are the current issues and nagging problems? (d) What has been achieved and what are our future expectations?
The classification algorithm based on SVM (support vector machine) attracts more attention from researchers due to its perfect theoretical properties and good empirical results. Compared with other classification algorithms, structural risk minimizations based SVM achieve high generalization performance with small number of samples. The text chunking, as a preprocessing step for parsing, is to divide text into syntactically related non-overlapping groups of words (chunks), reducing the complexity of the full parsing. In this paper, we treat Chinese text chunking as a classification problem, and apply SVM to solve it. The chunking experiments were carried out on the HIT Chinese Treebank corpus. Experimental results show that it is an effective approach, achieving an F score of 88.67%, especially for a small number of Chinese labeled samples.
All in-text\treferences\tunderlined\tin\tblue\tare\tlinked\tto\tpublications\ton\tResearchGate, letting you\taccess\tand\tread\tthem\timmediately.
Biomimetic design uses ideas from biological phenomena as inspiration in design. To support biomimetic design, biological analogies are identified by finding instances of functional keywords that describe the engineering problem in biological knowledge in natural-language format. Challenges in using this approach include the identification of keywords, and the quantity and quality of results found. WordNet, a lexical database, is used as a language framework to systematically generate alternative keywords to find matches and analyze the results of searches. Troponyms from WordNet were found to provide better and more plentiful keywords than did synonyms. Due to the potentially large number of matches to keywords, matches are analyzed to facilitate extraction of dominant biological phenomena associated with keywords. This analysis found that words that frequently collocated with keywords tend to be objects of the keyword verb or agents that carry out the actions of the keyword. Furthermore, nouns that are inanimate, e.g., substances, tend to be objects, and nouns that are animate e.g., animals, organs, tend to be agents. Distinguishing frequently collocated words and their relationships to keywords can be used to facilitate identification of biological analogies in natural-language format to support design.Copyright © 2004 by ASME