Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Word sense disambiguation automatically determines the appropriate senses of a word in context. We have previously shown that self-organized document maps have properties similar to a large-scale semantic structure that is useful for word sense disambiguation. This work evaluates the impact of different linguistic features on self-organized document maps for word sense disambiguation. The features evaluated are various qualitative features, e.g. part-of-speech and syntactic labels, and quantitative features, e.g. cut-off levels for word frequency. It is shown that linguistic features help make contextual information explicit. If the training corpus is large even contextually weak features, such as base forms, will act in concert to produce sense distinctions in a statistically significant way. However, the most important features are syntactic dependency relations and base forms annotated with part of speech or syntactic labels. We achieve 62.9% ± 0.73% correct results on the fine grained lexical task of the English SENSEVAL-2 data. On the 96.7% of the test cases which need no back-off to the most frequent sense we achieve 65.7% correct results.
We investigated subjective and hemodynamic responses towards disgust-inducing, fear-inducing, and neutral pictures in a functional magnetic resonance imaging study. Within an interval of 1 week, 24 male subjects underwent the same block design twice in order to analyze possible response changes to the repeated picture presentation. The results showed that disgust-inducing and fear-inducing scenes provoked a similar activation pattern in comparison to neutral scenes. This included the thalamus, primary and secondary visual fields, the amygdala, the hippocampus, and various regions of the prefrontal cortex. During the retest, the affective ratings hardly changed. In contrast, most of the previously observed brain activations disappeared, with the exception of the temporo-occipital activation. An additional analysis, which compared the emotion-related activation patterns during the two presentations, showed that the responses to the fear-inducing pictures were more stable than the responses to the disgust-inducing ones.
This paper presents the preliminary analysis of Kannada WordNet and the set of relevant computational tools. Although the design has been inspired by the famous English WordNet, and to certain extent, by the Hindi WordNet, the unique features of Kannada WordNet are graded antonyms and meronymy relationships, nominal as well as verbal compoundings, complex verb constructions and efficient underlying database design (designed to handle storage and display of Kannada unicode characters). Kannada WordNet would not only add to the sparse collection of machine-readable Kannada dictionaries, but also will give new insights into the Kannada vocabulary. It provides sufficient interface for applications involved in Kannada machine translation, spell checker and semantic analyser.
Most statistical parsers have used the grammar induction approach, in which a stochastic grammar is induced from a treebank. An alternative approach is to induce a controller for a given parsing automaton. Such controllers may be stochastic; here, we focus on greedy controllers, which result in deterministic parsers. We use decision trees to learn the controllers. The resulting parsers are surprisingly accurate and robust, considering their speed and simplicity. They are almost as fast as current part-ofspeech taggers, and considerably more accurate than a basic unlexicalized PCFG parser. We also describe Markov parsing models, a general framework for parser modeling and control, of which the parsers reported here are a special case.
In the present paper detailed schemes are proposed for an annotated corpus and a lexical database of Modern Hebrew. They are meant to be primary linguistic sources for more empirical studies of the grammatical and lexical structure of Modern Hebrew to shed light on aspects and phenomena hitherto unknown. XML (Extensible Markup Language) is chosen as their storage format because of its machine- and humanreadability, crossplatform-compatibility, crosslinguistic-compatibility, self-descriptiveness and capability of nesting structure. The corpus will be annotated in four levels, i.e., syntactically, morphosyntactically, lexically and morphologically. The lexical database will include modules of morphosyntax, inflection, word-formation and syntacticosemantics. The data will be recorded in Unicode, whether in Hebrew characters or in Latin transcription. Although countless numbers of revisions have been made since the idea of building these two sources in XML was first born a few years ago, what is proposed here is essentially by one individual. It might, therefore, need minor (or even major) revisions and/or expansions.
NLP applications in all domains require more than a formal grammar to process the input in a practical way, because natural language contains phenomena that a formal grammar is usually not able to describe. Such phenomena are typically disfluencies and extra-grammaticality. Some robust technique is needed to deal with them. An important issue in the development of robust parsing techniques is the choice of flexibility. What precise phenomena outside the systems grammar shall the parser be able to handle? Another question is how to select the correct analysis among the great number of solutions, which are produced as a consequence of the flexibility. This report presents experiments done with two different techniques. One is based on the combination of partial parses, the other on controlled relaxation of grammar rules. In both techniques the selection of the?best? analysis is done with a statistically based ranking procedure. The grammars and test sentences are extracted from two treebanks, ATIS and Susanne. Experimental results show that the first technique has the advantage of full coverage, while the other has a better accuracy. The best performance is achieved by parsing in three passes, first with the initial grammar, then with the rule-relaxation approach and finally, if still no analysis was found, with combination of partial analyses.
This paper presents a Chinese treebank based decision tree approach to identify Chinese BNP. A self-learning mechanism is integrated into our model which includes the following steps: auto-extraction of POS string sequences (BNP rules) and their context information from the corpora and ID3 algorithm based tree training. Experimental results show good performances of our method.
Tree-based approaches to alignment model translation as a sequence of probabilistic operations transforming the syntactic parse tree of a sentence in one language into that of the other. The trees may be learned directly from parallel corpora (Wu, 1997), or provided by a parser trained on hand-annotated treebanks (Yamada and Knight, 2001). In this paper, we compare these approaches on Chinese-English and French-English datasets, and find that automatically derived trees result in better agreement with human-annotated word-level alignments for unseen test data.
An essential component of Language Engineering (LE) tools are verb class descriptors that provide information about the relations of the predicates to their arguments. The production of computationally tractable language resources necessitates the assignment of types of predicate-argument relations to a great variety of verb-centered structures: it is necessary to define not only the initial, canonical valency frame of a great number of verb lexemes, but also the diathesis alternations, which reflect the real-life usage of verbs. This paper describes the implementation of descriptors of the valency properties of Bulgarian verbs used in the production of a syntactic treebank of Bulgarian. The descriptors are based on available LE resources for Bulgarian: a verb subcategorization model implemented in the lexical data base that is used; a chunk grammar that recognizes verb form patterns. Predictive models are built and applied in a grammar that annotates grammatical relations inferred from the combination of morphosyntactic and shallow syntactic processing cues. The real significance of this particular processing is the resolution, in relation to the valency properties of many verbs, of the discrepancy or the contradiction between the verb lexicon specifications and the verb syntagmatic realization. 1.
Reviewed by: A rainbow of corpora: Corpus linguistics and the languages of the world ed. by Andrew Wilson, Paul Rayson, and Tony McEnery Heiko Narrog A rainbow of corpora: Corpus linguistics and the languages of the world. Ed. by Andrew Wilson, Paul Rayson, and Tony McEnery. (Linguistics edition 40.) Munich: LINCOM Europa, 2003. Pp. 165. ISBN 3895868728. $73.20 (Hb). This volume is a collection of papers that were originally presented at Corpus Linguistics 2001, a conference held 30 March–2 April 2001 at Lancaster University (UK). The sister volume to the ‘English-oriented’ Corpus linguistics by the Lune: A festschrift for Geoffrey Leech (Frankfurt: Peter Lang, 2003), it includes contributions that are concerned with non-English languages. The languages dealt with in the present volume range from the better-known Indo-European languages to Biblical Hebrew, Korean, and Arabic. Content-wise, the individual articles can be roughly divided into three categories. First, there are three diachronically oriented papers from the workshop ‘Corpus linguistics, ancient languages, and older language periods’: Beatrix Färber on a corpus of Medieval Irish (19–26), Wolf-Dieter Syring on the design and usage of a text database of Biblical Hebrew (141–52), and Matthew Brook O’Donnell, Stanley E. Porter, and Jeffrey T. Reed on the database-assisted discourse analysis of texts from the New Testament (109–21). The other papers, which come from the main session, deal with modern languages. Half of them are primarily concerned with corpora as such that is, their design, mark-up, usage, and so on. Anne Abeillé, Lionel Clément, Alexandra Kinyon, and François Toussenel discuss the PARIS 7 annotated corpus for French (1–10), and Martin Beaudoin and Michel Samard present their work on a large corpus of written Canadian French (11–18). R. Rossini Favretti, F. Tamburini, and C. de Santis deal with the CORIS corpus of written Italian (27–38), and Eva Hajičová and Petr Sgall discuss the Prague Dependency Treebank (39–50). Shereen Khoja, Roger Garside, and Gerry Knowles present a tagset for the morphosyntactic tagging of Arabic (59–72),and Kiril Simov, Gergana Popova, and Petya Osenova introduce a HPSG-based syntactic treebank of Bulgarian (133–40). Apart from language-specific interests, the papers by Abeillé and colleagues and Hajičová and Sgall are particularly impressive. The former shows how a corpus tagged with higher accuracy than previous corpora can completely overturn research results based on corpora tagged with lower accuracy. The latter presents a corpus that is marked up for syntactic features, including topic-focus structure, with a depth that is probably unmatched in any language. The rest of the papers present corpus-based linguistic research. The paper by Beom-mo Kang, Hung-gyu Kim, and Myung-hoe Huh applies Douglas Biber’s multidimensional text analysis to Korean (51–57); Maarten Lemmens investigates the functions and meaning range of posture verbs in Swedish from a typological perspective (73–85); Martina Möllering analyzes the use of the German modal particle eben in a corpus of telephone conversations (87–97); P.-O. Nilsson shows how Swedish texts translated from English exhibit systematically different [End Page 904] lexical and grammatical patterns from those found in texts written originally in Swedish (99–107); Katja Ploog explores the syntax of pronominal subjects in Abidjanee French in contrast to standard French (123–32); and Adriana Vlad, Adrian Mitrea, and Mihai Mitrea investigate Romanian texts from a stochastic perspective (153–65). Many of these contributions combine quantitative corpus methods with qualitative methods of investigation and persuasively demonstrate how corpus-based research can contribute to broader linguistic issues. On the whole, not all papers in this volume are of the same theoretical interest, but each of them is at least informative. In contrast, the editing is rather disappointing. There is a table of contents and a...
A novel methodology is presented to enhance Chinese text chunking with the aid of transductive Hidden Markov Models (transductive HMMs),where the chunking is considered as a special tagging problem. An attempt is thus made to utilize it via a number of transformation functions to introduce as much relevant contextual information as possible in model training. These functions enable the models to make use of contextual information to a greater extent and keep away from costly changes of the original training and tagging process. Each of them results in an individual model with certain pros and cons. Through a number of experiments, the best two models are integrated into a significantly better one. The chunking experiments were carried out on the HIT Chinese Treebank corpus, of which the results show that it is an effective approach to the recognition of Chinese chunk, achieving an F score of 8238%.
This paper presents the construction of a manually annotated Chinese shallow Treebank, named PolyU Treebank. Different from traditional Chinese Treebank based on full parsing, the PolyU Treebank is based on shallow parsing in which only partial syntactical structures are annotated. This Treebank can be used to support shallow parser training, testing and other natural language applications. Phrase-based Grammar, proposed by Peking University, is used to guide the design and implementation of the PolyU Treebank. The design principles include good resource sharing, low structural complexity, sufficient syntactic information and large data scale. The design issues, including corpus material preparation, standard for word segmentation and POS tagging, and the guideline for phrase bracketing and annotation, are presented in this paper. Well-designed workflow and effective semiautomatic and automatic annotation checking are used to ensure annotation accuracy and consistency. Currently, the PolyU Treebank has completed the annotation of a 1-million-word corpus. The evaluation shows that the accuracy of annotation is higher than 98%. 1
The American poetess Emily Dickinson in the 19th century is well-known for her original imagery and unbound language style. Her discarding of the linguistic norms makes her verse darting and whimsical as well as sophisticated to understand. This paper is to interpret the mysterious poetess by means of exploring the linguistic deviations in her poetry. The two deviations that will be focused on are grammatical deviation and graphological deviation.
Rating agencies' track record is good in developed countries but poor in emerging economies. Why? Given the almost-monopolistic structure of the industry, we conjecture that agencies might underinvest in information gathering. We propose an indicator quantifying the agencies' effort to gather information and assess whether greater effort affects rating levels. We detect: (i) absolute underinvestment for non-OECD sovereigns (less effort in spite of greater opaqueness); (ii) relative underinvestment for non-OECD firms compared with OECD ones (though the former receive a larger effort, more intense effort boosts firm ratings in non-OECD countries while depressing them in OECD countries).
This paper discusses an annotation scheme for Korean null pronouns, which were used in annotating three kinds of Korean text corpora including Penn Korean Treebank. In annotating the corpora, null pronouns and their antecedents were marked up for their type and reference, with coreference relation tracked by numeric identifiers. Based on the annotation scheme, an outline of a potential pronoun resolution strategy is also proposed. The resulting dataset of annotated text is rather small at 11,834 words; we hope the null pronoun classification and annotation scheme proposed in this study will serve as a basis in developing a large-scale annotated corpus in the future.
The aim of this paper is to present the theoretical principles underlying the making of a lexical database of English collocations of non-specialized words used in scientific language. This project was prompted by the shortage of reference tools providing information about the use and combinatorial properties of general words in specific registers. A case study will illustrate that in scientific texts, words, especially polysemous verbs, have a distinct semantic and combinatorial behaviour. Following the assumption that the meaning and the grammatical and collocational patterns of words are interrelated, we suggest that context-specific information should be included in specialized reference tools to facilitate the written production of scientific texts by nonnative speakers ofEnglish.
OBJECTIVE: Transcutaneous electrical nerve stimulation (TENS) is a technique widely used in clinical practice to control pain, although its clinical efficacy remains controversial. Though many mechanisms have been proposed for its analgesic effects, there is a conspicuous lack of experimentally controlled research investigating whether TENS analgesia is related to its effects on the sympathetic nervous system (SNS). METHODS: Using an established psychophysiological paradigm, the present study investigated the effects of high-frequency/low-intensity TENS, low-frequency/high-intensity TENS, and sham TENS on the perception of experimental pain and SNS function in healthy volunteers. Measures of heart rate, digital pulse volume, and skin conductance were recorded during a 20-minute TENS stimulation period and in anticipation of a series of painful electric shocks prior to and following TENS stimulation. Healthy volunteers rated the intensity of the shocks using a 0-10-point verbal pain rating scale. RESULTS: The three TENS conditions failed to differentially effect SNS responses during either the 20-minute TENS treatment period or the shock anticipation periods, and TENS did not affect ratings of pain intensity to the shock stimuli. CONCLUSIONS: While these results may not generalize to acute or chronic pain patients, within the limitations of the present experimental paradigm, no support was found for TENS affecting either SNS function or acute experimental pain perception.
The motivation of the Papillon project is to encourage the development of freely accessible Multilingual Lexical Resources by way of online collaborative work on the Internet. For this, we developed a generic community website originally dedicated to the diffusion and the development of a particular acception based multilingual lexical database.
A critical issue in the study of speech motor control is the identification of the mechanisms that generate the temporal flow of serially ordered articulatory events. Two staged models of serial ordered events (Lashley, 1951; Lindblom, 1963) claim that time controls events whereas dynamic models predict a relative relation between time and space. Each of these models predicts a different relation between the acoustic measures of formant frequency and segmental duration. The most recent method described herein provides a sensitive index of speech deterioration which is both acoustically robust and phonetically systematic. Both acoustic and magnetic resonance imaging measures were used to describe the speech disturbance in two neurologically distinct groups of cerebellar ataxia: Friedreich’s ataxia and olivo-ponto cerebellar ataxia. The speaking task was designed to elicit six different prosodic conditions and four prosodic contrasts. All subjects read the same syllable embedded in a sentence, under six different prosodic conditions. Pair-wise comparisons derived from the six conditions were used to describe (1) final lengthening, (2) phrasal accent, (3) nuclear accent and (4) syllable reduction. An estimate of speech deterioration as determined by individual and normal subects’ acoustic values of syllable duration, formant and fundamental frequencies was used in correlation analyses with magnetic resonance imaging ratings.
In natural language processing a huge amount of structured data is constantly used for the extraction and presentation of grammatical structures in sentences. For example the Chinese Treebank corpus developed at the Institute of Information Science Academia Sinica Taiwan is a semantically annotated corpus that has been used to help parse and study Chinese sentences. In this setting users usually use structured tree patterns instead of keywords to query the corpus.
We describe an approach to two areas of biomedical information extraction, drug development and cancer genomics. We have developed a framework which includes corpus annotation integrated at multiple levels: a Treebank containing syntactic structure, a Propbank containing predicate-argument structure, and annotation of entities and relations among the entities. Crucial to this approach is the proper characterization of entities as relation components, which allows the integration of the entity annotation with the syntactic structure while retaining the capacity to annotate and extract more complex events. We are training statistical taggers using this annotation for such extraction as well as using them for improving the annotation process.
Abstract. WordNet is an electronic lexical database structured around psychological and linguistic principles. As such, it should play a part in any e ort to integrate cognitive factors into knowledge based systems. Yet, some of its basic assumptions have been attacked, and the suggestion made that a major restructuring would make it more cognitively transparent. We investigate these allegations from a psycholinguistic perspective and conclude that WordNet is in fact rigorous in terms of the cognitive principles it embodies. What is lacking is a methodology for translating the explicit and implicit knowledge in WordNet into a usable, formal ontologies. We show some ways in which WordNet should be extended to facilitate this process. We agree that WordNet is not in itself ready for use as a formal ontology, but we argue that it is an invaluable tool for describing the conceptualized structure of our world, and should be used as a fundamental resource. 1
Abstract. This paper describes experiments on using inductive machine learning to guide a deterinistic dependency parser for unrestricted natural language text. Using data from a small treebank of Swedish, an eager probabilistic learning algorithm is used to induce context-sensitive parse tables. Evaluation shows a significant improvement over the baseline, which uses a table without contextual information. 1
Expert observers discriminate elementary features (e.g., colour, orientation) of salient objects outside their current focus of attention, but not complex features such as T/L shape or colour arrangement. Surprisingly, natural scene category behaves like an elementary feature in this context, in that scenes with vehicles, animals, etc. are successfully classified in a dual-task situation (Li et al., 2002). To confirm and extend this finding, we used a similar dual-task paradigm and presented a natural scene in the near periphery (8 to 14 eccentricity) simultaneously with an attention-demanding task near fixation (seven rotated Ts/Ls). Visual persistence was controlled by a particularly effective form of masking. For scenes of animals or vehicles (black and white with matched luminance scale), categorization performance exhibited little or no attentional cost, i.e., dual- and single-task performance were comparable. Thus, scene categorization under dual-task conditions can be explained neither by inadequate masking nor by trivial colour cues. In a second experiment, observers categorized scenes from the International Affective Image System database (IAPS, Lang et al., 1995) as “pleasant” or “unpleasant”. Although absolute performance was now lower (∼75% of the level reached under ideal viewing conditions), there was no significant attentional cost. As low-level features do not identify affective content, this implies some comprehension of scene gist. A control experiment highlighted the contrast between natural scenes and geometric stimuli in the dual-task situation (and also confirmed the peripheral absence of attention), in that observers failed to discriminate the colour arrangement of a peripherally presented ‘pill’ (half red, half green, inclined ±45 ). Li FF et al. (2002) Rapid natural scene categorization in the near absence of attention. PNAS 99: 9596ff. Lang PJ et al. (1995) IAPS: Technical Manual and Affective Ratings. Gainsville, FL.
The subject matter of this article concerns objectivity of lexicographic description. Here objectivity is being associated with the names of “functional” places, in D. Le Pesant’s understanding of the term. The author presents principles of the object-oriented description of lexicosemantic data, according to the concept by W. Banyś. Later on, she distinguishes a couple of basic classes of appropries predicates, which are typical for the locative nouns suggested by D. Le Pesant. The elements of the object-oriented approach are confronted with their counterparts in such lexicographic theories as: I. Melczuk and A.K. Zholkovsky’s frames, J. Pustejovsky’s qualia structure system, G. Gross’s object classes and unique beginners from the WordNet lexical database. In the suggested descriptive scheme of entries the author places results of her own analyses concerning descriptions of names of functional places, in this case-buildings.
This paper presents a prosodic phrasing model for Korean to be used in a text-to-speech synthesis (TTS) system. Read text corpora were morpho-syntactically parsed and prosodically labeled following the Penn Korean Treebank (Han, Chunghye, Ko, Eon-Suk, Yi, Heejong, Palmer, M., 2002. Penn Korean Treebank: development and evaluation. In: Proceedings of the 16th Pacific Asian Conference on Language and Computation. Korean Society for Language and Information.) and K-ToBI prosodic labeling conventions (Sun-Ah, J., 2000. K-ToBI (Korean ToBI) labelling conventions. Version 3.1. Available from: URL.), respectively. Decision trees were trained with morpho-syntactic and textual distance features to predict locations of accentual and intonational phrase breaks. Our phrasing model cross-validated on a 300-sentence corpus (6936 words or 21,436 syllables, with an average of 72 syllables or 23 words per sentence) predicted non-breaks with F=92.4% and breaks with F=88.0% (F=72.8% for accentual phrase breaks and F=71.3% for intonational phrase breaks).
The present work falls in the line of activities promoted by the European Languguage Resource Association (ELRA) Production Committee (PCom) and raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources. A two-fold purpose lies behind this experiment. The first aim is to investigate the feasibility, define methods and procedures for combining two Italian lexical resources that have incompatible formats and complementary information into a Unified Lexicon (UL). The adopted strategy and the procedures appointed are described together with the driving criterion of the merging task, where a balance between human and computational efforts is pursued. The coverage of the UL has been maximized, by making use of simple and fast matching procedures. The second aim is to exploit this newly obtained resource for implementing the phonological and morphological layers of the CLIPS lexical database. Implementing these new layers and linking them with the already exisitng syntactic and semantic layers is not a trivial task. The constraints imposed by the model, the impact at the architectural level and the solution adopted in order to make the whole database ‘speak ’ efficiently are presented. Advantages vs. disadvantages are discussed. 1. Background and Motivations The work described here raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources (LRs) and has been
In this paper we will present work carried out on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples – Federico II. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. We will also present the linguistic annotation tools needed to allow the parser to produce syntactic structures automatically.\nIn particular, in order to produce appropriate linguistic annotations, all transcribed materials need to be transliterated from the audio-transcription to a more standard orthographic format. In that way, the parser receives as an input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts “overlap”, i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts or backchannels the current speaker. This phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically both to a following turn and to the local turn where it may produce conversational moves to repair what has been previously said by the current speaker.
An automatic method for annotating the Penn-II Treebank (Marcus et al., 1994) with high-level Lexical Functional Grammar (Kaplan and Bresnan, 1982; Bresnan, 2001; Dalrymple, 2001) f-structure representations is described in (Cahill et al., 2002; Cahill et al., 2004a; Cahill et al., 2004b; O’Donovan et al., 2004). The annotation algorithm and the automatically-generated f-structures are the basis for the automatic acquisition of wide-coverage and robust probabilistic approximations of LFG grammars (Cahill et al., 2002; Cahill et al., 2004a) and for the induction of LFG semantic forms (O’Donovan et al., 2004). The quality of the annotation algorithm and the f-structures it generates is, therefore, extremely important. To date, annotation quality has been measured in terms of precision and recall against the DCU 105. The annotation algorithm currently achieves an f-score of 96.57% for complete f-structures and 94.3% for preds-only \nf-structures. There are a number of problems with evaluating against a gold standard of this size, most \nnotably that of overfitting. There is a risk of assuming that the gold standard is a complete and balanced \nrepresentation of the linguistic phenomena in a language and basing design decisions on this. It is, therefore, \npreferable to evaluate against a more extensive, external standard. Although the DCU 105 is publicly available, \n1 a larger well-established external standard can provide a more widely-recognised benchmark against which the quality of the f-structure annotation algorithm can be evaluated. For these reasons, we present an evaluation of the f-structure annotation algorithm of (Cahill et al., 2002; Cahill et al., 2004a; Cahill et al., 2004b; O’Donovan et al., 2004) against the PARC 700 Dependency Bank (King et al., 2003). Evaluation against an external gold standard is a non-trivial task as linguistic analyses may differ systematically between the gold standard and the output to be evaluated as regards feature geometry and nomenclature. We present conversion software to automatically account for many (but not all) of the systematic differences. Currently, we achieve an f-score of 87.31% for the f-structures generated from the original Penn-II trees and \nan f-score of 81.79% for f-structures from parse trees produced by Charniak’s (2000) parser in our pipeline \nparsing architecture against the PARC 700.
This paper explores the interaction between conceptual structure and morpho-syntax. In particular, we show that ontology-based conceptual classification can be used to predict internal relations in compounds. We propose an ontology-based approach to predict the semantic relation between the two component words in Mandarin VV compounds. A Mandarin VV compound is classified according to the eventive relation between the two simplex verbs. These relations specify how the eventive meanings of the two simplex verbs combine to form the meaning of the compound. The three types of eventive relations that we deal with in this paper are: coordinate, modificational, and resultative. Since the way in which two events combine with each other depends upon their event types, we hypothesize that the eventive relations can be predicted by the conceptual classified event types of the two simplex verbs. An approach of ontology-based prediction is proposed based on this hypothesis. The assignment of ontology classification for each simplex verb is based on SUMO and Sinica BOW. The correlation between the ontology class of each verb position and each eventive type is trained and scored based on a manually tagged lexical database. We encode the ontology information of each VV compound in a 3-tuple based on these correlation scores. This 3-tuple is represented as a three-dimensional vector and used to predict the eventive type of new VV compounds. Our classification experiment on unknown VV compounds yields good recall and precision. 1.
The task of inducing grammar structures has received a great deal of attention. The reasons why researchers have studied are different; to use grammar induction as the first stage in building large treebanks or to make up better language models. However, grammar induction has inherent computational complexity. To overcome it, some grammar induction algorithms add new production rules incrementally. They refine the grammar while keeping their computational complexity low. In this paper, we propose a new efficient grammar induction algorithm. Although our algorithm is similar to algorithms which learn a grammar incrementally, our algorithm uses the graphical EM algorithm instead of the Inside-Outside algorithm. We report results of learning experiments in terms of learning speeds. The results show that our algorithm learns a grammar in constant time regardless of the size of the grammar. Since our algorithm decreases syntactic ambiguities in each step, our algorithm reduces required time for learning. This constant-time learning considerably affects learning time for larger grammars. We also reports results of evaluation of criteria to choose nonterminals. Our algorithm refines a grammar based on a nonterminal in each step. Since there can be several criteria to decide which nonterminal is the best, we evaluate them by learning experiments.
Nowadays language learning systems beyond hypermedia technologies increasingly adopt Computational Linguistics technologies, network technologies and Artificial Intelligence to improve the learning process. However, although such technologies are thought to be very useful to support authentic language learning environments and to simulate real communication situations a certain neglect of such systems is lamented, even within the research community. One reason is that specific knowledge out of different disciplines is necessary to provide an innovative language learning system. This development process can be made more efficient, if methods are found to join independently developed mature products to a single learning environment. In this article we will present two interdisciplinary research projects out of the domain of lexicography, which are both ready to be integrated within other e-learning resources. Word Manager (WM) is a reusable lexical database that provides information about inflection, word formation and orthography. ELDIT is a web-based educational system for learners of the German and Italian languages. Both systems can be integrated into other learning environments as well as extended by new modules by themselves. The collaboration between ELDIT and WM has shown that these advantages can also be realized in practice. Joining the two products efficiently supported the creation of an innovative language learning environment.
The OmniPaper project has implemented several information retrieval prototypes in the area of electronic news publishing. One prototype uses SOAP as communication protocol between the central system and a number of distributed news archives. The second prototype uses an RDF metadata database, enabling direct metadata queries to the central system. Finally the Topic Map prototype uses query expansion and semantic linking for smart metadata search. The Topic Map prototype enhances the search experience by implementing a knowledge layer that combines the semantic content of a lexical database, consisting of concepts and keywords, with a metadata-set of newspaper articles. After developing and testing three smaller prototypes, the OmniPaper consortium has combined these prototypes in one. In this final prototype a kind of “enhanced full-text search” engine is implemented. This means that the prototype is an interface on top of existing search engines. When a user submits a query, this query is forwarded to several distributed news archives to retrieve relevant news articles. Next to this, the system: 1) translates queries to enable multilingual search, 2) provides a query refinement mechanism, both in graphic and text-based form, allowing users to adapt their query and 3) provides uniform result ranking algorithm across the different news archives. In this prototype querying and navigation are considered as alternative methods to find relevant information. Both interact with each other and together they produce a combined user experience that can be expressed as find what you were looking for and then browse away from it. In fact, the prototype considers both querying and navigation as a kind of search action and tries to integrate both. In concrete, keywords in a query are looked up in a dictionary and shown to the user. In the background, the keywords are translated and expanded to related terms. These expanded queries are sent to the underlying full-text search engine(s) in all requested languages. In the graphical tool (“web of concepts”) users can redefine the meaning of their query words, resulting in an updated query and result set. Both disambiguation (choosing one meaning of a word out of many) and refinement (browsing to related words) are possible. Figure 1 shows the web of concepts for the query “poll Indonesia”. The word “Indonesia” is recognized in only one concept, “Dutch East Indies”, whereas the word “poll” has many different meanings. If you select the meaning “canvass” for example, this word is replacing the word “poll” in the original query. After selection the concept “canvass” can again be expanded to related concepts, be it more general or more specific in meaning. In the textual tool only refinement is possible. The user gets a list of words that are related to the words appearing in the query, grouped into more similar, more specific and more general terms. Then the user can change his/her query using these proposed words.
In the context of the Papillon project, which aims at creating a multilingual lexical database (MLDB), we have developed Jeminie, an adaptable system that helps automatically building interlingual lexical databases from existing lexical resources. In this article, we present a taxonomy of criteria for evaluating a MLDB, that motivates the need for arbitrary compositions of criteria to evaluate a whole MLDB. A quality measurement method is proposed, that is adaptable to different contexts and available lexical resources.
This paper tries to analyse translation criticism on published books, using concept and sub-categories of translation norms. Analysed in this paper translation criticism divided into shcolars’ (but not in translation studies area) and general readers’. Translation norms as an analysing tool consist of 1) adequacy vs. acceptability, 2) translation policy, 3) directness 4) matricial norm 5) text-linguistic norm. Results are as follows: scholars’ criticism tends to focus on adequacy (accordance with source text culture), whereas general readers point out insufficient acceptability(accordance with source text culture); it seems that translation policy and directness are not the main concern in translation criticism; adequacy vs. acceptability norm is closely related with matricial and text-linguistic norm. The difference of scholars’ and general readers’ criticism tendency implies that different views possibly exist between various reader groups. This point needs further examination. Close relationship between adequacy vs. acceptability and matricial vs. text-linguistic norm also needs further research. The latter may be concrete basis for the former. It it is confirmed, we can make the analysing tool more simple. As for general readers, it is observed that readers, publishing companies and tranlators interact very actively on Internet sites. Although based on limited examples, this paper shows current state of translation criticism by Korean readers.
This report presents an approach to enriching flat and robust predicate argument structures with more fine-grained semantic information, extracted from underspecified semantic representations and encoded in Minimal Recursion Semantics (MRS). Such representations are provided by a hand-built HPSG grammar with a wide linguistic coverage. A specific semantic representation, called linked predicate argument structure (LPAS), has been worked out, which describes the explicit embedding relationships among predicate argument structures. LPAS can be used as a generic interface language for integrating semantic representations with different granularities. Some initial experiments have been conducted to convert MRS expressions into LPASs. A simple constraint solver is developed to resolve the underspecified dominance relations between the predicates and their arguments in MRS expressions. LPASs are useful for high-precision information extraction and question answering tasks because of their fine-grained semantic structures. In addition, I have attempted to extend the lexicon of the HPSG English Resource Grammar (ERG) exploiting WordNet and to disambiguate the readings of HPSG parsing with the help of a probabilistic parser, in order to process texts from application domains. Following the presented approach, the HPSG ERG grammar can be used for annotating some standard treebank, e.g., the Penn Treebank, with its fine-grained semantics. In this vein, I point out opportunities for a fruitful cooperation of the HPSG annotated Redwood Treebank and the Penn PropBank. In my current work, I exploit HPSG as an additional knowledge resource for the automatic learning of LPASs from dependency structures.
To assign semantic roles in building Treebanks, there is a need for annotators having a guideline in determining semantic relations between phrasal head and its modifiers or arguments. Semantic roles are hard to have clear-cut definitions. It is not always easy to determine thematic relations between two concepts. This paper aims to introduce an integrated nominal modifier system. Basically we adopt other scholars ' incisive idea in analyzing semantic roles that modify general nouns. We use the approach of building a fine-grain taxonomy of role system. The taxonomy of fine-grain thematic roles makes the role determination easier for human annotators, since the meaning of a fine-grain semantic role is self explanatory and a higher-level semantic role is described by its hyponyms. The proposed taxonomy has been attested during construction of Sinica TreeBank and HowNet definitions of nominal concepts and proven to be more applicable than conventional flat structures. 1
Recent evaluation techniques applied to corpus-based systems have been introduced that can predict quantitatively how well surface realizers will generate unseen sentences in isolation. We introduce a similar method for determining the coverage on the Fuf/Surge symbolic surface realizer, report that its coverage and accuracy on the Penn TreeBank is higher than that of a similar statistics-based generator, describe several benefits that can be used in other areas of computational linguistics, and present an updated version of Surge for use in the NLG community.
Introduction The CorpusEye project (http://corp.hum.sdu.dk ) at the University of Denmark aims at designing and programming an internet based corpus search interface that (1) offers standardised search tools and a unified descriptive formalism across different corpus types and different languages, and (2) allows users to exploit grammatical information in annotated corpora in a user-friendly and menubased way. All corpora in CorpusEye have been annotated with VISL's Constraint Grammar based parsers, in the case of treebanks using an additional PSG module or equivalent (Bick 2003). At the time of writing, the material covers 8 languages and ca. 600 million words.
The current research explored the processes that predominate during the anticipation of an emotionally salient event. Experiment 1 (N536), employed three different conditional stimuli followed by pictorial pleasant, unpleasant or neutral unconditioned stimuli. Half the participants were trained with visual CSs, the other half with tactile CSs. In the group trained with visual CSs, startle eyeblinks were larger and faster during CSs that were paired with unpleasant pictures than CSs paired with neutral or pleasant pictures respectively, indicating an affect startle pattern. This linear trend was not found in the group trained with tactile CSs. Experiment 2 (N564) aimed to investigate whether the affective pattern found in the startle data in Experiment 1 could also be found using a behavioural measure of emotion. This time participants’ reaction time during a post-experimental affective priming taskwas used as dependantmeasure to assess the presence of emotional learning. Instead of a simple differential conditioning task, an occasion setting paradigm was employed and participants were trained using either a feature positive or feature negative design with pleasant or unpleasant picture USs. For participants trained with unpleasant USs, valence ratings collected before and after conditioning training suggested the presence of emotional learning, whereas no such pattern was found for participants trained with pleasant USs. These findings were not confirmed in the priming data.
The modern dominant approach to translation is that which aims at intelligibility of the meaning and normality of the style in the translated text. Many translation theorists maintain that a translator should attempt to produce a target text which is clear and understandable in meaning, normal in style, and natural in language. To attain this goal, one is allowed and sometimes obliged, to make some adjustments such as expansion and reduction in the process of transfer according to the linguistic norms of the receptor language. However, in translating a highly sensitive religious text like the Qur'an, the accuracy of the meaning conveyed in the translated text is of primary importance, since it is the source text which is considered as the main criterion in evaluating the translation. Therefore, no translator is allowed to sacrifice accuracy of the meaning for the sake of intelligibility and naturalness in the translation, nor is it legitimate to make the meaning unintelligible or distorted by producing a target text which is awkward andunnatural. What on should do is to produce a translation natural in language, intelligible and accurate in meaning
The purpose of this study is to identify the properties of special-word, and to show the process of extracting special-words from a large corpus. A special-word corresponds to the notion of unknown words, which is a counterpart of the lexical database in Natural Language Process(NLP). Generally unknown words cause a lot of ambiguities and thus decline the accuracy of NLP systems. The special-word in this work includes various expressions about the events of the day or the fashions, abbreviated words and naturalized word. We came up with a semi-automatic procedure of constructing a special-word dictionary mainly based on the language-dependent heuristics. We, however, also feel that other statistical considerations including frequencies, and probability distributions may be required for unknown word extractions in a higher automatic fashion.
The preceding articles by Piek Vossen (PV), Willy Martin (WM), and Marc van Campenoudt (MC) give a much more detailed account on their respective multilingual lexical database designs than the article by myself in this same journal (MJ). At the same time, they indicate some points of concern regarding the set-up of the SIM<it>u</it>LLDA system. Rather than responding directly to the points raised, this response elaborates on the two aspects of the SIM<it>u</it>LLDA system that seem to form the main sources of these issues: the status of the definitional attributes, and the practical usability of the system. The issues raised in the preceding articles will be explicitly addresses in the course of this elaboration. For even more details on these topics, see Janssen (2002).