Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
In my thesis I have attempted to develop an integrated translation approach materialized in the form of a Dynamic Translation Model (DTM). This endeavour can be justified to the extent that Translation Studies is perceived so far as a fragmentary discipline with implicitly and explicitly opposed and apparently irreconcilable points of view: linguistics-oriented approaches and culture-and-literature-oriented approaches. The main problem arising from this lack of common ground for further developing Translation Studies is that the disciplinary boundaries are not well-established and therefore the discipline itself cannot be developed coherently. Besides, Translation Studies is still to be constructed as an autonomous and an independent discipline that has a common core of theoretical and practical problems. This lack of coherent development of the discipline is due, I think, to an epistemological mistake: to believe that one single approach can account for (that is, describe and explain) all the translational reality. I propose to distinguish a two-phase epistemological move: 1. each translation approach works on its own research interests and acknowledges that its approach deals only with one part of the whole subject matter of Translation Studies; and 2. the results obtained by each translation approach are incorporated into a holistic integrative model like the Dynamic Translation Model I propose. In order to achieve this goal I have attempted to show the key tenets of modern translation approaches, both linguistics-oriented and culture-and-literature-oriented, by quoting the main theses of the representatives of these approaches. I have then presented the most important criticisms that have been raised in relation to these diverse translation approaches, together with my own criticisms (chapters 1 and 2). Also, I have introduced the theoretical basis for an integrated approach taking Holmes’ differentiation between theoretical (product-, process-, and function-oriented) and practical approaches as a point of departure. Likewise, I have discussed the problems of integrating Translation Studies, as well as Snell-Hornby’s integrated proposal and some key aspects of literary translation relevant for my integrative endeavour (chapter 3). Finally, I have developed my proposal for a Dynamic Translation Model (chapter 4). As to the conclusions of my thesis, I can say that my holistic DTM was able to integrate functionally aspects from both linguistics-oriented and culture-and-literature-oriented approaches: historico-cultural context (Leipzig School and postcolonial studies); norms, ideology and power (Descriptive Translation Studies; G. Toury and A. Lefevere); translation commisioner (Skopos theory); sender’s communicative purpose (linguistic and pragmatic approaches: W. Koller, J. House, H. Gerzymisch-Arbogast, etc); importance of source language text (linguistic and textlinguistic approaches; stylistic approaches; B. Spillner, B. Sandig); translator’s comprehension process (hermeneutic, deconstructive, and poststructural approaches), target language receiver in the target language historico-cultural context (Descriptive Translation Studies; postcolonial and gender studies). On the other hand, the three levels of the Dynamic Translation Model help to explain the flux of translational proceses and the variables that are activated or neutralized therein. They also incorporate concepts from other disciplines such as text linguistics, pragmatics, stylistics, and the communication theory. In my integrative endeavour I also proposed new concepts and, accordingly, coined new terms: Compulsory Translational Forces (CTF) (which include both Initiator’s Translational Instructions (ITI) and Target Language Valid Translational Norms (TL-VTN), Default Equivalence Position (DEP). In the pragmatic dimension of the model special attention is paid to what I call Text Illocutionary Indicators (TII) as well as the strengthening (upgraders) and weakening (downgraders) illocutionary mechanisms in relation to the Source Language Text (SLT) and the Target Language Text (TLT). Semantic/lexical fields play a crucial role in the establishment of equivalences between SLT and TLT in the text semantic dimension, as well as what I have called Fictionalizing Stylistic Shifts in the text stylistic dimension. As to the future developments of translation research within the framework of the Dynamic Translation Model I would say that some modificationbs may be called for so that interpretation can also be accounted for. This proposal can be used profitably in the field of translation criticism. As is the case with any other integrative approach, DTM should be widely discussed and criticized in order to validate its theoretical soundness and its application in Translation Studies. This thesis is an attempt to contribute in this research direction.
The paper presents a set of tools designed for the Czech syntax parser Synt. It desribes the development as well as the data used in the testing, newly created Brno Phrasal Treebank.
Semantic Network Manual Annotation and its Evaluation The present contribution is a brief extract of (Novák, 2008). The Prague Dependency Treebank (PDT) is a valuable resource of linguistic information annotated on several layers. These layers range from morphemic to deep and they should contain all the linguistic information about the text. The natural extension is to add a semantic layer suitable as a knowledge base for tasks like question answering, information extraction etc. In this paper I set up criteria for this representation, explore the possible formalisms for this task and discuss their properties. One of them, Multilayered Extended Semantic Networks (Multi-Net), is chosen for further investigation. Its properties are described and an annotation process set up. I discuss some practical modifications of MultiNet for the purpose of manual annotation. MultiNet elements are compared to the elements of the deep linguistic layer of PDT. The tools and problems of the annotation process are presented and initial annotation data evaluated.
Both impression and function are important issues for design creation. These two factors compose meanings of designed objects. Design process can be viewed as a process of development of the structure of meanings. This research presents a new design methodology by focusing on the structure of meanings. The processes of search and evaluation of meanings form the fundamental phases of this method. In order to facilitate the searching for the meanings, the WordNet lexical database and an existing visualization tool (Visuwords) are adopted. The basic tool used for evaluation process is the WordNet::Similarity software, measuring the relatedness of meanings in the database. The measures of relatedness of meanings are developed as convergence criteria for application in the processes of evaluation. In this research, the steps of the design methodology, including the search and evaluation processes involved in the development of the structure of the meanings, are elucidated by carrying out demonstrations of proposed system.
Speech monitoring encompasses detection and self-repair of errors.This paper first reviews types of errors and self-repairs,then focuses on three theoretical accounts of how the monitoring mechanism works to detect and correct errors.Product-based theory assumes that there is a monitor which is equipped with phonological,lexical and syntactical rules and pragmatic norms and whose sole function is to monitor errors at varying levels when language is produced.Perception-based theory posits that a central monitor within the conceptualizer functions to accomplish the monitoring job.Node structure theory accounts for monitoring from the node activation hypothesis,i.e.,detection and correction of errors is tied to the activation strength or level of node committed or uncommitted.
<h3>Introduction</h3><br> Penn Discourse Treebank (PDTB) Version 3.0 is the third release in the Penn Discourse Treebank project, the goal of which is to annotate the Wall Street Journal (WSJ) section of Treebank-2 (<a href="http://catalog.ldc.upenn.edu/LDC95T7" rel="nofollow">LDC95T7</a>) with discourse relations. Penn Discourse Treebank Version 2 (<a href="../../../LDC2008T05" rel="nofollow">LDC2008T05</a>) contains over 40,600 tokens of annotated relations. In Version 3, an additional 13,000 tokens were annotated, certain pairwise annotations were standardized, new senses were included and the corpus was subject to a series of consistency checks. Details concerning the development of PDTB Version 3.0 can be found in the documentation accompanying this release. <br> Largely because the PDTB project was based on the idea that discourse relations are grounded in an identifiable set of explicit words or phrases (discourse connectives) or simply in the adjacency of two sentences, the PTDB has been used by many researchers in the natural language processing community and more recently, by researchers in psycholinguistics. It has also stimulated the development of similar resources in other languages and domains. <br> <h3>Data</h3><br> Annotations are provided in the form of separate text files (<em>standoff annotation</em>) that are byte-indexed into the raw WSJ text files in Treebank-2. The raw WSJ files are also included in this release. All text files are plain text, encoded in UTF-8. <br> This corpus contains two tools: (1) The Annotator, used for annotation and adjudication, and which can also be used for viewing the corpus; and (2) The Conversion Tool for converting Version 2 annotation files into the Version 3 format. <br> The documentation directory contains a manual describing what is new in Version 3 and how Version 3 differs from Version 2; the methods and guidelines used in annotating PDTB Version 3; and a range of statistics on the tokens, including the frequency of each connective, its sense labels and its modifiers. More information about the corpus and research carried out by the developers and others using the corpus can be found on the <a href="https://www.seas.upenn.edu/~pdtb/">PDTB website</a>. <br> <h3>Samples</h3><br> One can see samples of the annotation of different types of discourse relations, along with their visualization in the Annotator tool at: <br> <ul><br> <li><a href="desc/addenda/LDC2019T05_examples.html">Explicit relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#implicit_examples">Implicit relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#altlex_examples">Altlex and AltLexC relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#entrel_norel_examples">Entity relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#entrel_norel_examples">Hypophora relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#entrel_norel_examples">NoRel</a> (annotated only between adjacent sentences within a paragraph that are not linked to each other by a discourse relation)</li><br> </ul><br> <h3>Updates</h3><br> Experiments carried out in Fall 2019 on the intra-sentential discourse relations in the PDTB-3 revealed two problems with the corpus: (1) the final versions of two gold files of "to clause" annotation had not been loaded, and (2) several tokens were inadvertently omitted on the assumption that they were duplicates, when they were not. <br> Repairing these errors, and correcting a mis-labelled token in file wsj_1026, has added another 45 implicit intra-sentential relations to the corpus. Counts in the Annotation Manual have been adjusted to take these additional tokens into account. Specific changes/additions are recorded in the file "pdtb3-revision-jan-2020.txt". Downloads after February 3, 2020 contain the updated corpus. <br> <h3>Acknowledgment</h3><br> This work has been funded by the National Science Foundation, under grant NSF IIS 1422186 to the University of Pennsylvania and grant NSF IIS 1421067 to the University of Wisconsin, Milwaukee. The content of this publication does not necessarily reflect the position or policy of the Government, and no official endorsement should be inferred. </br> Portions © 1987-1989 Dow Jones & Company, Inc., © 2008, 2012, 2019 The Penn Discourse Treebank Group, © 2008, 2012, 2019 Trustees of the University of Pennsylvania
The present study is concerned with objectively observable factors that facilitate acceptability of translated texts in Persian. It considers textual features like information load (also known as lexical density), range of vocabulary (also known as lexical variety or type-token ratio) and average sentence length to be arguably objective observable factors. The research, that has been carried out as an MA thesis in Translation Studies at Allameh Tabataba’i University, intends to determine whether a norm could be said to operate in the form of a significant relationship between these textual features and acceptance or popularity of translated texts with readers. A comparable corpus of Persian translational and non-translational fiction was built and used as the material for the study.
In this paper, we describe our work on building a parallel treebank for a less studied and typologically dissimilar language pair, namely Swedish and Turkish. The treebank is a balanced syntactically annotated corpus containing both fiction and technical documents. In total, it consists of approximately 160,000 tokens in Swedish and 145,000 in Turkish. The texts are linguistically annotated using different layers from part of speech tags and morphological features to dependency annotation. Each layer is automatically processed by using basic language resources for the involved languages. The sentences and words are aligned, and partly manually corrected. We create the treebank by reusing and adjusting existing tools for the automatic annotation, alignment, and their correction and visualization. The treebank was developed within the project Supporting research environment for minor languages aiming at to create representative language resources for language pairs dissimilar in language structure. Therefore, efforts are put on developing a general method for formatting and annotation procedure, as well as using tools that can be applied to other language pairs easily. 1.
We describe our initial efforts towards developing a large-scale corpus of Hindi texts annotated with discourse relations. Adopting the lexically grounded approach of the Penn Discourse Treebank (PDTB), we present a preliminary analysis of discourse connectives in a small corpus. We describe how discourse connectives are represented in the sentence-level dependency annotation in Hindi, and discuss how the discourse annotation can enrich this level for research and applications. The ultimate goal of our work is to build a Hindi Discourse Relation Bank along the lines of the PDTB. Our work will also contribute to the cross-linguistic understanding of discourse connectives. 1
This paper discusses a framework for development of bilingual and multilingual comprehension assistants and presents a prototype implementation of an English-Bulgarian comprehension assistant. The framework is based on the application of advanced graphical user interface techniques, WordNet and compatible lexical databases as well as a series of NLP preprocessing tasks, including POS-tagging, lemmatisation, multiword expressions recognition and word sense disambiguation. The aim of this framework is to speed up the process of dictionary look-up, to offer enhanced look-up functionalities and to perform a context-sensitive narrowing-down of the set of translation alternatives proposed to the user.
Graph-based and transition-based approaches to dependency parsing adopt very different views of the problem, each view having its own strengths and limitations. We study both approaches under the framework of beam-search. By developing a graph-based and a transition-based dependency parser, we show that a beam-search decoder is a competitive choice for both methods. More importantly, we propose a beam-search-based parser that combines both graph-based and transition-based parsing into a single system for training and decoding, showing that it outperforms both the pure graph-based and the pure transition-based parsers. Testing on the English and Chinese Penn Treebank data, the combined system gave state-of-the-art accuracies of 92.1% and 86.2%, respectively.
This paper deals with a multilingual relational lexical database of proper name, Prolexbase, a free resource available on the CNRTL website. The Prolex model is based on two main concepts: firstly, a language independent pivot and, secondly, the prolexeme (the projection of the pivot onto particular language), that is a set of lemmas (names and derivatives). These two concepts model the variations of proper name: firstly, independent of language and, secondly, language dependent by morphology or knowledge. Variation processing is very important for NLP: the same proper name can be written in different instances, maybe in different parts of speech, and it can also be replaced by another one, a lexical anaphora (that reveals semantic link). The pivot represents different referent&apos;s points of view, i.e. language independent variations of name. Pivots are linked by three semantic relations (quasi-synonymy, partitive relation and associative relation). The prolexeme is a set of variants (aliases), quasi-synonyms and morphosemantic derivatives. Prolexemes are linked to classifying contexts and reliability code.
Abstract Affective ratings of multiple religious (sub)groups (Muslims, Christians, Jews and non-believers, as well as Sunni, Alevi and Sjiit Muslims), the endorsement of Islamic minority rights and religious group identification were examined among Sunni and Alevi Turkish-Dutch participants. The findings show that both groups differ in important ways. Some Alevi participants considered themselves Muslims but others interpreted Alevi identity in a secular way. The Sunnis were quite negative towards Jews and non-believers, they more strongly endorsed Islamic minority rights and they had very high Muslim group identification. Furthermore, the Sunnis were negative towards Alevis and the Alevis were negative towards the Sunnis. Muslim group identification was positively and strongly related to feelings towards Muslims and to the endorsement of Islamic group rights.
Conventional n-best reranking techniques often suffer from the limited scope of the n-best list, which rules out many potentially good alternatives. We instead propose forest reranking, a method that reranks a packed forest of exponentially many parses. Since exact inference is intractable with non-local features, we present an approximate algorithm inspired by forest rescoring that makes discriminative training practical over the whole Treebank. Our final result, an F-score of 91.7, outperforms both 50-best and 100-best reranking baselines, and is better than any previously reported systems trained on the Treebank. 1
Objective: To carry out the native assessment of International Affective Picture System(IAPS) among Chinese older adults.Methods:Altogether 116 Chinese older adults,including 51 male and 65 female,from three communities in Dalian City,aged from 60 to 80 years,rated 60 pictures(positive:25,neutral:12,negative:23) selected from the IAPS in terms of valence,arousal and dominance with Self-Assessment Manikin(SAM).The mean affective ratings were compared to the normative ratings of USA National Institute of Mental Health(NIMH).Result: Reliability analysis indicated that the affective ratings of our sample were stable and highly internally consistent.The affective ratings of Chinese older participants were strongly correlated with the normative ratings of NIMH(r=0.92,0.54 and 0.88 respectively for valence,arousal and dominance,P0.001).But paired t test showed there were still significant differences between the two samples.Chinese aged reported relatively higher arousal and dominance than NIMH sample for all pictures [(5.33±0.93)vs.(4.83±1.25),(5.60±1.20)vs.(5.19±1.21),P0.001],but lower valence than NIMH sample[(4.99±2.28)vs.(5.28±1.85),P=0.020].Male and female Chinese older participants showed similar emotional responses to most pictures.But female Chinese older participants reported higher valence than male ones(5.05±2.33/4.93±2.24,P0.05).The 60 pictures were distributed as shape in the two-dimensional affective space(valence-arousal).The association between valence and arousal was pronounced and linear for positive pictures(r=0.71,P0.001),but unpronounced for negative pictures,(r=-0.35,P0.05).Conclusion: IAPS is highly internationally accessible just as the expectation of its designers.However,considering about great differences in many aspects such as culture,social living and age between Chinese aged and NIMH sample,they may have different affective experiences to the same emotional stimuli.Therefore it is necessary to do some revisal before the IAPS is applied to Chinese aged.
Morphological processes in Semitic languages deliver space-delimited words which introduce multiple, distinct, syntactic units into the structure of the input sentence. These words are in turn highly ambiguous, breaking the assumption underlying most parsers that the yield of a tree for a given sentence is known in advance. Here we propose a single joint model for performing both morphological segmentation and syntactic disambiguation which bypasses the associated circularity. Using a treebank grammar, a data-driven lexicon, and a linguistically motivated unknown-tokens handling technique our model outperforms previous pipelined, integrated or factorized systems for Hebrew morphological and syntactic processing, yielding an error reduction of 12% over the best published results so far. 1
Abstract Large linguistic databases, especially databases having a global coverage, such as the World Atlas of Language Structures, the Automated Similarity Judgment Program, and Ethnologue, are making it possible to systematically investigate many aspects of how languages change and compete for viability. Agent‐based computer simulations supplement such empirical data by analyzing the necessary and sufficient parameters for the current global distributions of languages or linguistic features. By combining empirical datasets with simulations and applying quantitative methods, it is now possible to address fundamental questions, such as ‘what are the relative rates of change in different parts of languages?’, ‘why are there a few large language families, many intermediate ones, and even more small ones?’, ‘do small languages change faster or slower than large ones?’, or ‘how does the borrowing of words relate to the borrowing of structural features?’
We have built a parallel treebank that includes word and phrase alignment. The alignment information was manually checked using a graphical tool that allows the annotator to view a pair of trees from parallel sentences. We found the compilation of clear alignment guidelines to be a difficult task. However, experiments with a group of students have shown that we are on the right track with up to 89% overlap between the student annotation and our own. At the same time these experiments have helped us to pin-point the weaknesses in the guidelines, many of which concerned unclear rules related to differences in grammatical forms between the languages.
To date, parsers have made limited use of semantic information, but there is evidence to suggest that semantic features can enhance parse disambiguation. This paper shows that semantic classes help to obtain significant improvement in both parsing and PP attachment tasks. We devise a gold-standard sense- and parse tree-annotated dataset based on the intersection of the Penn Treebank and SemCor, and experiment with different approaches to both semantic representation and disambiguation. For the Bikel parser, we achieved a maximal error reduction rate over the baseline parser of 6.9% and 20.5%, for parsing and PP-attachment respectively, using an unsupervised WSD strategy. This demonstrates that word sense information can indeed enhance the performance of syntactic disambiguation. © 2008 Association for Computational Linguistics.
Cornetto byl dvouletý projekt (STE05039), ve kterem byla vytvořena lexikalni semanticka databaze kombinujici Wordnet s informacemi typu FrameNet pro holandstinu. Kombinaci těchto lexikalnich zdrojů vznikla výrazně bohatsi databaze jazykových vztahů, ktera umožni kvalitnějsi výsledky technologii zpracovani přirozeneho jazyka, jako je desambiguace významu slov (WSD) a systemy generovani jazyka. Kromě propojeni Wordnetu s informacemi typu FrameNet je databaze take mapovana na formalni ontologii, ktera poskytuje přesne semanticke popisy.
This paper presents an effective dependency parsing approach of incorporating short dependency information from unlabeled data. The unlabeled data is automatically parsed by a deterministic dependency parser, which can provide relatively high performance for short dependencies between words. We then train another parser which uses the information on short dependency relations extracted from the output of the first parser. Our proposed approach achieves an unlabeled attachment score of 86.52, an absolute 1.24% improvement over the baseline system on the data set of Chinese Treebank. 1
We present the STYX system, which is designed as an electronic corpus-based exercise book of Czech morphology and syntax with sentences directly selected from the Prague Dependency Treebank, the largest annotated corpus of the Czech language. The exercise book offers complex sentence processing with respect to both morphological and syntactic phenomena, i. e. the exercises allow students of basic and secondary schools to practice classifying parts of speech and particular morphological categories of words and in the parsing of sentences and classifying the syntactic functions of words. The corpus-based exercise book presents a novel usage of annotated corpora outside their original context.
We present a description of a new resource (Prague Dependency Treebank of Spoken Language) being created for English and Czech to be used for the task of speech understanding, broad natural language analysis for dialog systems and other speech-related tasks, including speech editing. The resources we have created so far contain audio and a standard transcription of spontaneous speech, but as a novel layer, we add an edited (ldquoreconstructedrdquo) version of the spoken utterances. These edits go beyond the scope of current speech reconstruction efforts in that we allow, on top of the usual deletions of speech artifacts, fillers, etc. also for word modifications, insertions and word order changes. We have used both monologue and dialogue recordings in English and Czech to verify the feasibility of such transcription. We have also assessed the quality of the resulting annotation since the relative freedom of the editing raises an issue of what a ldquocorrectrdquo annotation is.
We present the second version of the Penn Discourse Treebank, PDTB-2.0, describing its lexically-grounded annotations of discourse relations and their two abstract object arguments over the 1 million word Wall Street Journal corpus. We describe all aspects of the annotation, including (a) the argument structure of discourse relations, (b) the sense annotation of the relations, and (c) the attribution of discourse relations and each of their arguments. We list the differences between PDTB-1.0 and PDTB-2.0. We present representative statistics for several aspects of the annotation in the corpus. 1.
With the advent of the Internet, billions of images are now freely available online and constitute a dense sampling of the visual world. Using a variety of non-parametric methods, we explore this world with the aid of a large dataset of 79,302,017 images collected from the Internet. Motivated by psychophysical results showing the remarkable tolerance of the human visual system to degradations in image resolution, the images in the dataset are stored as 32 x 32 color images. Each image is loosely labeled with one of the 75,062 non-abstract nouns in English, as listed in the Wordnet lexical database. Hence the image database gives a comprehensive coverage of all object categories and scenes. The semantic information from Wordnet can be used in conjunction with nearest-neighbor methods to perform object classification over a range of semantic levels minimizing the effects of labeling noise. For certain classes that are particularly prevalent in the dataset, such as people, we are able to demonstrate a recognition performance comparable to class-specific Viola-Jones style detectors.
Sciendo provides publishing services and solutions to academic and professional organizations and individual authors. We publish journals, books, conference proceedings and a variety of other publications.
"Treebanks allow for the creation of a valence lexicon per side effect. The TüBa-D/Z valence lexicon has been created in lockstep with the development of the TüBa- D/Z treebank as such. For each verb encountered in the treebank, the annotators created a lexical entry that records the valence frames of the verbs contained in the sentence, unless they are already contained in the valence lexicon as result of previous annotation. The TüBa-D/Z valence lexicon currently contains a total of 8013 frames for 4896 distinct verb lemmas. Since treebank annotation is still ongoing, the lexicon will continue to grow. Such a lexicon has utility in its own right as a resource for lexicalized parsing and a variety of NLP applications. At the same time, the lexicon can serve as a source for aiding consistency of annotation and automatic detection of annotation errors"
A Floresta Sintá(c)tica tem como objetivo criar e disponibilizar um corpus sintaticamente anotado. Neste artigo, são apresentados dois novos materiais do projeto: Selva (300 mil palavras e parcialmente revisto) e Amazônia (3.8 milhões de palavras, não revisto). Para lidar com um material tão grande e variado foi construída a interface Milhafre. O artigo mostra, ainda, como vem sendo enfrentado o desafio de compatibilizar, de uma lado, o usuário lingüista, que pode ter um perfil muito heterogêneo e, em geral, pouca familiaridade determinadas formalizações mais utilizadas em informática e, de outro, um único modelo de anotação sintática, freqüentemente pouco conhecido do lado “lingüístico não-computacional” e uma interface de acesso e manipulação de corpora capaz de lidar com um objeto tão complexo como a língua. Palavras-chave: árvores sintáticas, corpus anotado, corpus revisto, busca em corpora.
The progress of Chinese dependency treebank construction has fallen behind other languages, such as English, in terms of scale and quality. Building a large scale treebank needs a lot of human and material resources. Meanwhile, it is very difficult to guarantee the quality of the treebank. In this paper, we explore a new method which combines rule-based method and statistical-based method to convert a constituent treebank named Penn Chinese Treebank to a dependency treebank which follows the annatation standard of HIT Chinese Dependency Treebank (HIT-IR-CDT). We increase the size of training data by adding converted treebank into HIT-IR-CDT and re-train the dependency parser. Experiments show that small addition of converted treebank can improve the performance of dependency parser, while large addition will bring it down. Through detailed analysis, we believe that convertion of constituent-to-dependency treebank, being a method of improving performance of dependency parser by utilizing different treebanks, still needs in-depth research.
We present a robust parser which is trained on a treebank of ungrammatical sentences. The treebank is created automatically by modifying Penn treebank sentences so that they contain one or more syntactic errors. We evaluate an existing Penn-treebank-trained parser on the ungrammatical treebank to see how it reacts to noise in the form of grammatical errors. We re-train this parser on the training section of the ungrammatical treebank, leading to an significantly improved performance on the ungrammatical test sets. We show how a classifier can be used to prevent performance degradation on the original grammatical data.
We describe a parsing approach that makes use of the perceptron algorithm, in conjunction with dynamic programming methods, to recover full constituent-based parse trees. The formalism allows a rich set of parse-tree features, including PCFG-based features, bigram and trigram dependency features, and surface features. A severe challenge in applying such an approach to full syntactic parsing is the efficiency of the parsing algorithms involved. We show that efficient training is feasible, using a Tree Adjoining Grammar (TAG) based parsing formalism. A lower-order dependency parsing model is used to restrict the search space of the full model, thereby making it efficient. Experiments on the Penn WSJ treebank show that the model achieves state-of-the-art performance, for both constituent and dependency accuracy.
Graph-based and transition-based approaches to dependency parsing adopt very different views of the problem, each view having its own strengths and limitations. We study both approaches under the framework of beamsearch. By developing a graph-based and a transition-based dependency parser, we show that a beam-search decoder is a competitive choice for both methods. More importantly, we propose a beam-search-based parser that combines both graph-based and transitionbased parsing into a single system for training and decoding, showing that it outperforms both the pure graph-based and the pure transition-based parsers. Testing on the English and Chinese Penn Treebank data, the combined system gave state-of-the-art accuracies
A web-based collaborative environment including on-line authoring tools that is managed by a central database was developed in collaboration with several countries including Peru, Bolivia, and the United States. The application involved developing a linguistics database and eLearning environment for documenting, preserving, and promoting language training for Aymara, a language indigenous to Peru and Bolivia. The database, an ontology management system called Lyra, incorporates all elements of the language (dialogues, phrase patterns, phrases, words, and morphemes) as well as cultural multimedia resources (images and sound recordings). The organization of the database enables a high level of integration among language elements and cultural resources. Authoring tools are used by experts in the Aymara language to build the linguistic database. These tools are accessible on-line as part of the collaborative environment using standard web browsers incorporating the Java plug-in. The eLearning student interface is a web-based program written in Flash. The Flash program automatically interprets and formats data objects retrieved from the database in XML format. The student interface is presented in Spanish and English. A web service architecture is used to publish the database on-line so that it can be accessed and utilized by other application programs in a variety of formats
In the paper we describe the results in the development of the tools for handling diverse multilingual lexical resources such as monolingual and multilingual dictionaries, terminological dictionaries, complex lexicographic databases or WordNet semantic networks. All the presented tools are based on the Dictionary Editor and Browser (DEB) platform which uses standard XML formats. In this direction we strive to standardization of the lexical resources and also their interoperability. All the presented tools are freely available. We summarize the basic features of the DEB platform as a whole and then concentrate on four applications: DEBDict (a general dictionary browser), DEBTerm (multilingual terminological dictionary editor), PRALED (Czech Lexical Database system) and Visual Browser (graphical semantic network browser).
We outline the problem of ad hoc rules in treebanks, rules used for specific constructions in one data set and unlikely to be used again. These include ungeneralizable rules, erroneous rules, rules for ungrammatical text, and rules which are not consistent with the rest of the annotation scheme. Based on a simple notion of rule equivalence and on the idea of finding rules unlike any others, we develop two methods for detecting ad hoc rules in flat treebanks and show they are successful in detecting such rules. This is done by examining evidence across the grammar and without making any reference to context. 1
Expression of the serotonin transporter is affected by the genotype of the 5-HTTLPR (short and long forms) as well as the genotype of the SNP rs25531 within this region. Based on the combined genotypes for these polymorphisms, we designated each allele as a high or low expressing allele according to established expression levels-resulting in HiHi, HiLo, & LoLo genotype groups for analysis. We evaluated effects of gender and the promoter genotype on induction of negative affect by intravenous infusion of L: -tryptophan (TRP). The protocol consisted of a day-1 sham saline infusion and a day-2 active TRP infusion. Models assessed 5-HTTLPR composite genotype and gender as predictors of change in ratings of negative emotion during TRP infusion. During sham infusion there were no significant changes from baseline in mood ratings. During TRP infusion all negative affect ratings increased significantly from baseline (P's <.02). The genotype x gender interaction was a significant predictor of depression-dejection (P =.013), and trended towards predicting anger-hostility (P =.084). Males in the HiHi group had greater increases in negative affect during infusion, compared to all groups except LoLo females, who also showed increased negative affect.
CONTEXT: Cognitive decline, mood, behavioral and sleep disturbances, and limitations of activities of daily living commonly burden elderly patients with dementia and their caregivers. Circadian rhythm disturbances have been associated with these symptoms. OBJECTIVE: To determine whether the progression of cognitive and noncognitive symptoms may be ameliorated by individual or combined long-term application of the 2 major synchronizers of the circadian timing system: bright light and melatonin. DESIGN, SETTING, AND PARTICIPANTS: A long-term, double-blind, placebo-controlled, 2 x 2 factorial randomized trial performed from 1999 to 2004 with 189 residents of 12 group care facilities in the Netherlands; mean (SD) age, 85.8 (5.5) years; 90% were female and 87% had dementia. INTERVENTIONS: Random assignment by facility to long-term daily treatment with whole-day bright (+/- 1000 lux) or dim (+/- 300 lux) light and by participant to evening melatonin (2.5 mg) or placebo for a mean (SD) of 15 (12) months (maximum period of 3.5 years). MAIN OUTCOME MEASURES: Standardized scales for cognitive and noncognitive symptoms, limitations of activities of daily living, and adverse effects assessed every 6 months. RESULTS: Light attenuated cognitive deterioration by a mean of 0.9 points (95% confidence interval [CI], 0.04-1.71) on the Mini-Mental State Examination or a relative 5%. Light also ameliorated depressive symptoms by 1.5 points (95% CI, 0.24-2.70) on the Cornell Scale for Depression in Dementia or a relative 19%, and attenuated the increase in functional limitations over time by 1.8 points per year (95% CI, 0.61-2.92) on the nurse-informant activities of daily living scale or a relative 53% difference. Melatonin shortened sleep onset latency by 8.2 minutes (95% CI, 1.08-15.38) or 19% and increased sleep duration by 27 minutes (95% CI, 9-46) or 6%. However, melatonin adversely affected scores on the Philadelphia Geriatric Centre Affect Rating Scale, both for positive affect (-0.5 points; 95% CI, -0.10 to -1.00) and negative affect (0.8 points; 95% CI, 0.20-1.44). Melatonin also increased withdrawn behavior by 1.02 points (95% CI, 0.18-1.86) on the Multi Observational Scale for Elderly Subjects scale, although this effect was not seen if given in combination with light. Combined treatment also attenuated aggressive behavior by 3.9 points (95% CI, 0.88-6.92) on the Cohen-Mansfield Agitation Index or 9%, increased sleep efficiency by 3.5% (95% CI, 0.8%-6.1%), and improved nocturnal restlessness by 1.00 minute per hour each year (95% CI, 0.26-1.78) or 9% (treatment x time effect). CONCLUSIONS: Light has a modest benefit in improving some cognitive and noncognitive symptoms of dementia. To counteract the adverse effect of melatonin on mood, it is recommended only in combination with light. TRIAL REGISTRATION: controlled-trials.com/isrctn Identifier: ISRCTN93133646.
Robust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction).
We present a dependency-driven parser that parses both dependency structures and constituent structures. Constituency representations are automatically transformed into dependency representations with complex arc labels, which makes it possible to recover the constituent structure with both constituent labels and grammatical functions. We report a labeled attachment score close to 90% for dependency versions of the TIGER and TúBa-D/Z treebanks. Moreover, the parser is able to recover both constituent labels and grammatical functions with an F-Score over 75% for TüBa-D/Z and over 65% for TIGER.
The mere exposure effect is the commonly observed increase in pleasantness ratings of stimuli that have been given prior exposure. According to the fluency attribution account of the mere exposure effect, repeated presentations of a stimulus lead to increased ease of processing, which in turn is attributed to pleasantness. If so, processing fluency manipulated by means other than repetition should influence liking. In the present experiment, processing fluency was manipulated using a negative priming procedure, and its influence on affective judgement was examined. Previously ignored stimuli were responded to slower (negative priming) and were rated as less pleasant than controls. It was concluded that decreased processing fluency decreases liking of previously ignored stimuli.
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank. MorphoTrees is the idea of building effective and intuitive hierarchies over the information provided by computational morphological systems. MorphoTrees are implemented for Arabic as an extension to the TrEd annotation environment based on Perl. Encode Arabic libraries for Haskell and Perl serve for processing the non-trivial and multi-purpose ArabTEX notation that encodes Arabic orthographies and phonetic transcriptions in parallel.
We present the first results on parsing the SynTagRus treebank of Russian with a data-driven dependency parser, achieving a labeled attachment score of over 82% and an unlabeled attachment score of 89%. A feature analysis shows that high parsing accuracy is crucially dependent on the use of both lexical and morphological features. We conjecture that the latter result can be generalized to richly inflected languages in general, provided that sufficient amounts of training data are available.
Parser self-training is the technique of taking an existing parser, parsing extra data and then creating a second parser by treating the extra data as further training data. Here we apply this technique to parser adaptation. In particular, we self-train the standard Charniak/Johnson Penn-Treebank parser using unlabeled biomedical abstracts. This achieves an f-score of 84.3% on a standard test set of biomedical abstracts from the Genia corpus. This is a 20% error reduction over the best previous result on biomedical data (80.2% on the same test set).
Periods in the development of the lexical database in the Czech Language Institute, programmes.