Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Natural language processing modules such as part-of-speech taggers, named-entity recognizers and syntactic parsers are commonly evaluated in isolation, under the assumption that artificial evaluation metrics for individual parts are predictive of practical performance of more complex language technology systems that perform practical tasks. Although this is an important issue in the design and engineering of systems that use natural language input, it is often unclear how the accuracy of an end-user application is affected by parameters that affect individual NLP modules. We explore this issue in the context of a specific task by examining the relationship between the accuracy of a syntactic parser and the overall performance of an information extraction system for biomedical text that includes the parser as one of its components. We present an empirical investigation of the relationship between factors that affect the accuracy of syntactic analysis, and how the difference in parse accuracy affects the overall system.
We present procedures which pool lexical information estimated from unlabeled data via the Inside-Outside algorithm, with lexical information from a treebank PCFG. The procedures produce substantial improvements (up to 31.6% error reduction) on the task of determining subcategorization frames of novel verbs, relative to a smoothed Penn Treebank-trained PCFG. Even with relatively small quantities of unlabeled training data, the re-estimated models show promising improvements in labeled bracketing f-scores on Wall Street Journal parsing, and substantial benefit in acquiring the subcategorization preferences of low-frequency verbs.
Earlier work in parsing Arabic has speculated that attachment to construct state constructions decreases parsing performance. We make this speculation precise and define the problem of attachment to construct state constructions in the Arabic Treebank. We present the first statistics that quantify the problem. We provide a baseline and the results from a first attempt at a discriminative learning procedure for this task, achieving 80% accuracy.
We present a simple and effective semisupervised method for training dependency parsers. We focus on the problem of lexical representation, introducing features that incorporate word clusters derived from a large unannotated corpus. We demonstrate the effectiveness of the approach in a series of dependency parsing experiments on the Penn Treebank and Prague Dependency Treebank, and we show that the cluster-based features yield substantial gains in performance across a wide range of conditions. For example, in the case of English unlabeled second-order parsing, we improve from a baseline accuracy of 92:02% to 93:16%, and in the case of Czech unlabeled second-order parsing, we improve from a baseline accuracy of 86:13% to 87:13%. In addition, we demonstrate that our method also improves performance when small amounts of training data are available, and can roughly halve the amount of supervised data required to reach a desired level of performance.
This paper discusses the role of annotated corpora as works of reference for grammatical translation problems. Within this context, the English-German CroCo Corpus and its multi-layer alignment and annotation are introduced. It is described how the corpus is exploited as interactive resource to display translation solutions for typologically problematic constructions. Additionally, the Penn and TiGer Treebanks are used as comparable corpora for English and German. The linguistic enrichment of the treebanks, i.e. their syntactic annotation, is described and corpus query techniques relevant for translation problems are shown. Relevant structures are extracted from the treebanks and translation candidates are displayed and discussed. The advantage of this technique is that translation solutions are extracted from published translations, i.e. language in use. Consequently, they are more comprehensive and inventive than dictionary entries or descriptions in grammars are. Treebanks could thus be used as an interactive reference grammar in translation education and practice.
Automatic image annotation is very important for image retrieval. Despite continuous efforts in inventing new annotation algorithms, the annotation performance is usually unsatisfactory, and the annotation vocabulary is still limited due to the use of a small scale training set. In this paper, a novel image automatic annotation system based on the WordNet is presented, named WordNet-based image annotation. By using WordNet hierarchical structure, we collect a large image datasets. And each image is loosely labeled with one of the non-abstract nouns in English, as listed in the WordNet lexical database. Then we use PageRank method to delete the wrong images under every word, and make sure that every word covers 100 images. Hence the image database gives a comprehensive coverage of all object categories and scenes. The semantic information from WordNet can be used in conjunction with SVM classifiers to perform object classification over a range of semantic levels minimizing the effects of labeling noise. The system models a real-world situation by including pictures gathered from the Internet and is designed for exploratory large scale image retrieval system based on the internet.
We present a dependency-driven parser that parses both dependency structures and constituent structures. Constituency representations are automatically transformed into dependency representations with complex arc labels, which makes it possible to recover the constituent structure with both constituent labels and grammatical functions. We report a labeled attachment score close to 90% for dependency versions of the TIGER and TúBa-D/Z treebanks. Moreover, the parser is able to recover both constituent labels and grammatical functions with an F-Score over 75% for TüBa-D/Z and over 65% for TIGER.
Problematic types of prepositions, conjuctions, and particles and possible ways of their treatment in the lexical database.
In this paper we present the methodology for Word Sense Disambiguation based on domain information. Domain is a set of words in which there is a strong semantic relation among the words. The words in the sentence contribute to determine the domain of the sentence. The availability of WordNet domains makes the domain-oriented text analysis possible. The domain of the target word can be fixed based on the domains of the content words in the local context. This approach can be effectively used to disambiguate nouns. We present the unsupervised approach to Word Sense Disambiguation using the WordNet domains. The model determines the domain of the target word and the sense corresponding to this domain is taken as the correct sense. We have used the WordNet domains 3.1.as lexical database.
Possession in some Austronesian languages shows levels of elaboration far in excess of cross-linguistic norms, while in others it is strikingly unelaborated. The appearance of alienable/inalienable contrasts has been assumed to result from contact with Papuan languages, and the existence of a paradigm of indirect possessive classifiers is cited as one of the pieces of evidence for the Oceanic subgroup, while acknowledging that indirect possession constructions can be found in Malayo-Polynesian languages further west. We argue that the appearance of possessive classifiers in these languages is also the result of contact with Papuan languages west of New Guinea.
This paper describes a method of accurately projecting Propbank roles onto constituents in the CCGbank with near perfect accuracy and automatically annotating verbal categories with the semantic roles of their arguments. The current version of the CCGbank annotates arguments and adjuncts in a suboptimal way – it relies heavily on the Penn Treebank CLR tag, which is widely considered unreliable. By incorporating Propbank roles we are able to modify the derivation to better reflect linguistic reality. Tagging of nodes in the CCG derivation also permits us to annotate verbal categories with semantic roles corresponding to their syntactic arguments, which has strong implications for many NLP tasks.
The Berkeley FrameNet Project (BFN) is making an English lexical database called FrameNet, which describes syntactic and semantic properties of an English lexicon extracted from large electronic text corpora (Baker et al., 1998). Other projects dealing with Spanish, German and Japanese follow a similar approach and annotate large corpora. FrameSQL is a web-based application developed by the author, and it allows the user to search the BFN database in a variety of ways (Sato, 2003). FrameSQL shows a clear view of the headword’s grammar and combinatorial properties offered by the FrameNet database. FrameSQL has been developing and new functions were implemented for processing the Spanish FrameNet data (Subirats and Sato, 2004). FrameSQL is also in the process of incorporating the data of the Japanese FrameNet Project (Ohara et al., 2003) and that of the Saarbrücken Lexical Semantics Acquisition Project (Erk et al., 2003) into the database and will offer the same user-interface for searching these lexical data. This paper describes new functions of FrameSQL, showing how FrameSQL deals with the lexical data of English, Spanish, Japanese and German seamlessly. 1.
OBJECTIVE: This experimental, repeated-measures, crossover design study with nursing home residents examined the efficacy of reflexology in individuals with mild-to-moderate stage dementia. Specifically, the study tested whether a weekly reflexology intervention contributed to the resident outcomes of reduced physiologic distress, reduced pain, and improved affect. SETTING: The study was conducted at a large nursing home in suburban Philadelphia. SAMPLE: The sample included 21 nursing home residents with mild-to-moderate stage dementia randomly assigned to two groups. INTERVENTIONS: The first group received 4 weeks of weekly reflexology treatments followed by 4 weeks of a control condition of friendly visits. The second group received 4 weeks of friendly visits followed by 4 weeks of weekly reflexology. OUTCOME MEASURES: The primary efficacy endpoint was reduction of physiologic distress as measured by salivary alpha-amylase. The secondary outcomes were observed pain (Checklist of Nonverbal Pain Indicators) and observed affect (Apparent Affect Rating Scale). RESULTS: The findings demonstrate that when receiving the reflexology treatment condition, as compared to the control condition, the residents demonstrated significant reduction in observed pain and salivary alpha-amylase. No adverse events were recorded during the study period. CONCLUSIONS: This study provides preliminary support for the efficacy of reflexology as a treatment of stress in nursing home residents with mild-to-moderate stage dementia.
Recent parsing research has started addressing the questions a) how parsers trained on different syntactic resources differ in their performance and b) how to conduct a meaningful evaluation of the parsing results across such a range of syntactic representations. Two German treebanks, Negra and TüBa-D/Z, constitute an interesting testing ground for such research given that the two treebanks make very different representational choices for this language, which also is of general interest given that German is situated between the extremes of fixed and free word order. We show that previous work comparing PCFG parsing with these two treebanks employed PARSEVAL and grammatical function comparisons which were skewed by differences between the two corpus annotation schemes. Focusing on the grammatical dependency triples as an essential dimension of comparison, we show that the two very distinct corpora result in comparable parsing performance.
High doses of ibuprofen have been shown to inhibit muscle protein synthesis after a bout of resistance exercise. We determined the effect of a moderate dose of ibuprofen (400 mg x d(-1)) consumed on a daily basis after resistance training on muscle hypertrophy and strength. Twelve males and 6 females (approximately 24 years of age) trained their right and left biceps on alternate days (6 sets of 4-10 repetitions), 5 d x week(-1), for 6 weeks. In a counter-balanced, double-blind design, they were randomized to receive 400 mg x d(-1) ibuprofen immediately after training their left or right arm, and a placebo after training the opposite arm the following day. Before- and after-training muscle thickness of both biceps was measured using ultrasound and 1 repetition maximum (1 RM) arm curl strength was determined on both arms. Subjects rated their muscle soreness daily. There were time main effects for muscle thickness and strength (p < 0.01). Ibuprofen consumption had no effect on muscle hypertrophy (muscle thickness of biceps for arm receiving ibuprofen: pre 3.63 +/- 0.14, post 3.92 +/- 0.15 cm; and placebo: pre 3.62 +/- 0.15, post 3.90 +/- 0.15 cm) and strength (1 RM of arm receiving ibuprofen: pre 18.6 +/- 2.8, post 23.4 +/- 3.5 kg; and placebo: pre 18.8 +/- 2.8, post 22.8 +/- 3.4 kg). Muscle soreness was elevated during the first week of training only, but was not different between the ibuprofen and placebo arm. We conclude that a moderate dose of ibuprofen ingested after repeated resistance training sessions does not impair muscle hypertrophy or strength and does not affect ratings of muscle soreness.
German genitive attributes are usually tagged as such in treebanks. However, it is well known that this information is not sufficient for determining the type of relation between head nouns and attributes, as genitive attributes can express many different semantic relations. Various linguistic classifications have been worked out, but to my knowledge, nobody has so far proposed to apply this linguistic knowledge to a corpus. The challenge here is to come up with a classification that is both easy to verify and sufficiently fine-grained. Using earlier linguistic approaches as guidelines, I propose in this paper a detailed annotation scheme for German genitive attributes based on readily identifiable noun features. First insights from its application to the Smultron Treebank show that it is easy to distinguish between the proposed classes and that my classification of genitive attributes can be related to a more general semantic annotation level.
As the first holder of the first chair in computational linguistics in Sweden, Anna Sagvall Hein has played a central role in the development of computational linguistics and language technology bo...
Both impression and function are important issues for design creation. These two factors compose meanings of designed objects. Design process can be viewed as a process of development of the structure of meanings. This research presents a new design methodology by focusing on the structure of meanings. The processes of search and evaluation of meanings form the fundamental phases of this method. In order to facilitate the searching for the meanings, the WordNet lexical database and an existing visualization tool (Visuwords) are adopted. The basic tool used for evaluation process is the WordNet::Similarity software, measuring the relatedness of meanings in the database. The measures of relatedness of meanings are developed as convergence criteria for application in the processes of evaluation. In this research, the steps of the design methodology, including the search and evaluation processes involved in the development of the structure of the meanings, are elucidated by carrying out demonstrations of proposed system.
Humans are unique in being able to reflect on their own performance. For example, we are more motivated to do well on a task when we are told that our abilities are being evaluated. We set out to study the effect of self-motivation on a working memory task. By telling one group of participants that we were assessing their cognitive abilities, and another group that we were simply optimizing task parameters, we managed to enhance the motivation to do well in the first group. We matched the performance between the groups. During functional magnetic resonance imaging, the motivated group showed enhanced activity when making errors. This activity was extensive, including the anterior paracingulate cortex, lateral prefrontal and orbitofrontal cortex. These areas showed enhanced interaction with each other. The anterior paracingulate activity correlated with self-image ratings, and overlapped with activity when participants explicitly reflected upon their performance. We suggest that the motivation to do well leads to treating errors as being in conflict with one's ideals for oneself.
This paper proposes an approach using large scale case structures, which are automatically constructed from both a small tagged corpus and a large raw corpus, to improve Chinese dependency parsing. The case structure proposed in this paper has two characteristics: (1) it relaxes the predicate of a case structure to be all types of words which behaves as a head; (2) it is not categorized by semantic roles but marked by the neighboring modifiers attached to a head. Experimental results based on Penn Chinese Treebank show the proposed approach achieved 87.26% on unlabeled attachment score, which significantly outperformed the baseline parser without using case structures.
Recent work in Evolutionary Phonology (Blevins 2005, 2006, Blevins & Wedel 2008, Yu 2007, among others) has developed alternate explanations for typological universals or tendencies found across the sound systems of unrelated languages. This research emphasizes the role of patterns of language use and language change in the development of cross-linguistic patterns, rather than placing the burden of explanation on synchronic cognitive factors (i.e., Universal Grammar).
This paper introduces the infrastructure and the principles of a semantic framework used for the analysis and classification of verbs, developed with the aim of constructing a lexical database of Mandarin verbal semantics, called the Mandarin VerbNet. Distinct from most existing lexical databases that enumerate word senses without detailed grammatical considerations, the Mandarin VerbNet is designed to provide lexical semantic information based on grammatical descriptions and anchored in linguistic theories. It looks for systematic correlations between syntax and semantics and classifies verbs according to these syntax-to-semantics correspondences. The framework adopts the approach of Frame Semantics (Fillmore & Atkins 1992) in defining verb meanings in a semantic frame and building a frame-based verbal lexicon, but it differs from the structure of the English FrameNet in distinguishing different scopes of frames. Evolved and refined from previous works (Liu, Chiang & Chang 2004, Liu & Wu 2003, Liu 2002), this study summarizes the current model of the analytic framework with a detailed illustration from Mandarin statement verbs. It ultimately seeks to identify a theoretically sound and operationally effective representational scheme that bases its semantic analysis on grammatical behaviors and provides linguistic motivations for its semantic classifications.
Creation of material, technical and personal prerequisites for lexicographic work by using modern technologies; information on lexicographical bibliographical database.
Semantic Network Manual Annotation and its Evaluation The present contribution is a brief extract of (Novák, 2008). The Prague Dependency Treebank (PDT) is a valuable resource of linguistic information annotated on several layers. These layers range from morphemic to deep and they should contain all the linguistic information about the text. The natural extension is to add a semantic layer suitable as a knowledge base for tasks like question answering, information extraction etc. In this paper I set up criteria for this representation, explore the possible formalisms for this task and discuss their properties. One of them, Multilayered Extended Semantic Networks (Multi-Net), is chosen for further investigation. Its properties are described and an annotation process set up. I discuss some practical modifications of MultiNet for the purpose of manual annotation. MultiNet elements are compared to the elements of the deep linguistic layer of PDT. The tools and problems of the annotation process are presented and initial annotation data evaluated.
In the age of communication, new technological advances are made everyday in the field of information and linguistics. Deriv@ is a system that makes the most of linguistic data. All data managed by this application are connected with derivational morphology, since derived words constitute the starting point from which relationships with other linguistic elements are established. Deriv@ is a linguistic database with two models of representation: one for all Spanish words created by means of derivation; the other for the corresponding Latin words. We offer a grammatically-analysed corpus, both synchronically and diachronically, to make customised queries according to the user.s interests, and also to create a dictionary of derived words. Deriv@ allows and facilitates non-restricted access to all the information available in the two databases. This model is valid for any form derived from Spanish or from other Romance languages. It offers users the possibility of using it real time, thus letting them interact with the system and contribute to its improvement.
Robust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction).
Expression of the serotonin transporter is affected by the genotype of the 5-HTTLPR (short and long forms) as well as the genotype of the SNP rs25531 within this region. Based on the combined genotypes for these polymorphisms, we designated each allele as a high or low expressing allele according to established expression levels-resulting in HiHi, HiLo, & LoLo genotype groups for analysis. We evaluated effects of gender and the promoter genotype on induction of negative affect by intravenous infusion of L: -tryptophan (TRP). The protocol consisted of a day-1 sham saline infusion and a day-2 active TRP infusion. Models assessed 5-HTTLPR composite genotype and gender as predictors of change in ratings of negative emotion during TRP infusion. During sham infusion there were no significant changes from baseline in mood ratings. During TRP infusion all negative affect ratings increased significantly from baseline (P's <.02). The genotype x gender interaction was a significant predictor of depression-dejection (P =.013), and trended towards predicting anger-hostility (P =.084). Males in the HiHi group had greater increases in negative affect during infusion, compared to all groups except LoLo females, who also showed increased negative affect.
Modern statistical parsers are trained on large annotated corpora (treebanks). These treebanks usually consist of sentences addressing different subdomains (e.g. sports, politics, music), which implies that the statistics gathered by current statistical parsers are mixtures of subdomains of language use. In this paper we present a method that exploits raw subdomain corpora gathered from the web to introduce subdomain sensitivity into a given parser. We employ statistical techniques for creating an ensemble of domain sensitive parsers, and explore methods for amalgamating their predictions. Our experiments show that introducing domain sensitivity by exploiting raw corpora can improve over a tough, state-of-the-art baseline. 1.
This study examined the effects of appraisal of sexual stimuli on sexual arousal in women with superficial dyspareunia (n = 50) and sexually functional women (n = 25). To elicit different appraisals of an erotic film fragment, participants received an instruction prior to viewing it, with a focus on genital pain or on sexual enjoyment. A neutral instruction served as a control condition. Assignment to instruction condition was randomized. Genital arousal (vaginal pulse amplitude) and self-report ratings of affect and genital sensations were obtained in response to the erotic stimulus. As predicted, appraisal of the erotic stimulus affected genital responding, albeit marginally significant. Follow-up tests indicated that women who received the genital pain instruction responded with marginally significant lower genital arousal levels than women who received the sexual enjoyment instruction (d = 0.67). A significant instruction effect for negative affect was found, signifying that negative affect ratings were highest after the genital pain instruction and lowest after the sexual enjoyment instruction (d = 0.80). A marginally significant group by instruction interaction effect was observed for positive affect, indicating that women with dyspareunia reported significantly less positive affect than controls after the sexual enjoyment instruction (d = 1.48). Whereas women with dyspareunia reported overall marginally significant more negative affect than controls (d = 0.48), there were no differences in genital responsiveness between groups. These results provided preliminary evidence for the modulatory effects of appraisal of sexual stimuli on subsequent genital responding and affect in women with and without sexual complaints.
Both trait anger-in (managing anger through suppression) and anger-out (managing anger through direct expression) are related to pain responsiveness, but only anger-out effects involve opioid mechanisms. Preliminary work suggested that the effects of anger-out on postoperative analgesic requirements were moderated by the A118G single nucleotide polymorphism of the mu opioid receptor gene. This study further explored these potential genotypexphenotype interactions as they impact acute pain sensitivity. Genetic samples and measures of anger-in and anger-out were obtained in 87 subjects (from three studies) who participated in controlled laboratory acute pain tasks (ischemic, finger pressure, thermal). McGill Pain Questionnaire (MPQ) Sensory and Affective ratings for each pain task were standardized within studies, aggregated across pain tasks, and combined for analyses. Significant anger-outxA118G interactions were observed (p's<.05). Simple effects tests for both pain measures revealed that whereas anger-out was nonsignificantly hyperalgesic in subjects homozygous for the wild-type allele, anger-out was significantly hypoalgesic in those with the variant G allele (p's<.05). For the MPQ-Affective measure, this interaction arose both from low pain sensitivity in high anger-out subjects with the G allele and heightened pain sensitivity in low anger-out subjects with the G allele relative to responses in homozygous wild-type subjects. No genetic moderation was observed for anger-in, although significant main effects on MPQ-Affective ratings were noted (p<.005). Anger-in main effects were due to overlap with negative affect, but anger-outxA118G interactions were not, suggesting unique effects of expressive anger regulation. Results support opioid-related genotypexphenotype interactions involving trait anger-out.
While the effect of domain variation on Penn-treebank- \ntrained probabilistic parsers has been investigated in previous work, we study its effect on a Penn-Treebank-trained probabilistic generator. We show that applying the generator to data from the British National Corpus \nresults in a performance drop (from a BLEU score of 0.66 on the standard WSJ test set to a BLEU score of 0.54 on our BNC test set). We develop a generator retraining method where the domain-specific training data is automatically \nproduced using state-of-the-art parser output. The retraining method recovers a substantial portion of the performance drop, resulting in a generator which achieves a BLEU score of 0.61 on our BNC test data.
In this study, 120 males (60 sexual offenders and 60 non-sexual offenders) in psychiatric treatment while in prison were evaluated using neuropsychological, psychological, and sociological/demographic measures. All sexual offenders (N = 60) would be evaluated for potential civil commitment as sexually violent predators before prison release. Non-sexual offenders (N = 60) had not been convicted of a sexual offense. Sexual offenders demonstrated significantly more overall neuropsychological impairment suggesting diffuse brain differences, with dysfunction primarily associated with temporal and frontal brain cortexes; higher Psychopathy Checklist-Revised Factor 1 (Interpersonal/Affective) ratings and Rorschach responses indicated disordered attachment, disordered self-perception, and impulsive emotionality. Sexual offenders also were more likely to be younger and unmarried. Stepwise logistic regression analysis resulted in 80.20% accuracy of prediction of sexual offenders. Potential application of this empirically derived multidimensional description to treatment of sexual offenders is discussed. Potential limitations to generalization of this information are also discussed.
The query language in TIGERSearch is limited due to its lack of universal quantification. This restriction makes it impossible to ask simple queries like „Find sentences that do not include a certain word”. We propose an easy way to formulate such queries. We have implemented this extension to the query language in a tool that allows querying parallel treebanks including their alignment constraints. Our implementation of universal quantification relies on the view of node sets rather than single node unification. Our query tool is freely available.
In this paper we describe some technical and theoretical aspects related to a manually aligned bilingual treebank Italian (ITA) – Italian Sign Language (LIS) provided with both constituency and dependency annotation (Siena University Treebank, SUT). We briefly discuss the linguistic rationale behind the feature set and the dependency/constituency structure we adopted. Moreover we discuss the tool we used to annotate, semi-automatically, the treebank that, in the end, will be evaluated qualitatively with respect to a specific Transfer-Based Machine Translation (TB-MT) task.
Hungarian Academy of ScienceEötvös Loránd UniversityThis paper examines the Afro-Asiatic etymologies of Chadic lexical roots discussed by Olga V. Stolbova in her Chadic Lexical Database, Issue I (2005). The analysis is arranged according to the following sections: (1) Common Chadic reconstructions, (2) Isolated Chadic roots that nevertheless have Afro-Asiatic cognates. The paper represents the third part of my longer series of papers on addenda et corrigenda to Chadic lexical roots.
The paper presents a set of tools designed for the Czech syntax parser Synt. It desribes the development as well as the data used in the testing, newly created Brno Phrasal Treebank.
Our general objective is to explain how norms can emerge in complex, ambiguous situations: settings with large and complex spaces of normative options over which populations may try to agree using only limited, indirect knowledge of each others ’ currently preferred options, possibly gained through limited interaction samples. We study this process using the concrete example of agents developing and using common languages. Language can be viewed as an inherently distributed information and representation system. Because of this, it serves as a model problem for studying central issues in many kinds of distributed information systems, and norms are one such issue. Language is inherently normative. First, communicative language requires agreement—conventions—on many language aspects (such as ontological units, grammar, lexicon, morphology, etc.). Second, accurate communication is valuable. The value of successful communication translates to value for the conventionalization of language; this value in turn creates the decentralized normative force that drives agents to obey linguistic conventions. In this way, linguistic conventions become normative constraints on possible communication options, since obeying them increases communicability and its resulting communication payoffs. None of this means that convergence to linguistic norms is easy; in fact it presents novel issues not yet well understood. We show how ambiguity arises in language convergence, and describe a variety of techniques to resolve that ambiguity. We focus on one particular technique, text based learning. We show that it significantly reduces the amount of effort required for linguistic norms to emerge, and we show how it is an instance of a general norm-convergence technique.
This document gives a brief overview of the conversion of the Penn Treebank (Marcus et al., 1993, 1994) to the dependency structures used in the CoNLL-2008 SharedTask. Our dependency framework has the following properties: • single-head: every word has exactly one parent, except the root, which has no parent. • single-root: only one word in the sentence is root. • traceless: the dependency structures use no empty categories. Special arc labels are used to encode gapping. • nonprojective: some long-distance syntactic phenomena are represented in the dependency structure by means of non-local links. The conversion procedure relies on earlier work on constituent-to-dependency conversion (Magerman, 1994; Collins, 1999; Yamada and Matsumoto, 2003; Johansson and Nugues, 2007). In addition, we imported dependencies inside NPs and hyphenated words from a version of the Penn Treebank mapped into GLARF, the Grammatical and Logical Argument Representation Framework (Meyers et al., 2001). To assign dependency labels, we used the following general principles: • If there is a Treebank label other than CLR, HLN, NOM, TPC, or TTL: use this label. • If the link is inside an NP or a hyphenated word: use the label from GLARF. • Else infer a label using a set of rules. The complete set of labels is listed in Section 4.
Data-driven learning based on shift reduce parsing algorithms has emerged dependency parsing and shown excellent performance to many Treebanks. In this paper, we investigate the extension of those methods while considerably improved the runtime and training time efficiency via L2-SVMs. We also present several properties and constraints to enhance the parser completeness in runtime. We further integrate root-level and bottom-level syntactic information by using sequential taggers. The experimental results show the positive effect of the root-level and bottom-level features that improve our parser from 81.17 % to 81.41 % and 81.16 % to 81.57 % labeled attachment scores with modified Yamada’s and Nivre’s method, respectively on the Chinese Treebank. In comparison to well-known parsers, such as Malt-Parser (80.74%) and MSTParser (78.08%), our methods produce not only better accuracy, but also drastically reduced testing time in 0.07 and 0.11, respectively. 1
Grammar induction is one of attractive research areas of natural language processing. Since both supervised and to some extent semi-supervised grammar induction methods require large treebanks, and for many languages, such treebanks do not currently exist, we focused our attention on unsupervised approaches. Constituent Context Model (CCM) seems to be the state of the art in unsupervised grammar induction. In this paper, we show that the performance of CCM in free word order languages (FWOLs) such as Persian is inferior to that of fixed order languages such as English. We also introduce a novel approach, called parent-based constituent context model (PCCM), and show that by using some history notion of context and constituent information of each span's parent, the performance of CCM, especially in dealing with FWOLs, can be significantly improved.
Recently the LATL has undertaken the development of a multilingual translation system based on a symbolic parsing technology and on a transfer-based translation model. A crucial component of the system is the lexical database, notably the bilingual dictionaries containing the information for the lexical transfer from one language to another. As the number of necessary bilingual dictionaries is a quadratic function of the number of languages considered, we will face the problem of getting a large number of dictionaries. In this paper we discuss a solution to derive a bilingual dictionary by transitivity using existing ones and to check the generated translations in a parallel corpus. Our first experiments concerns the generation of two bilingual dictionaries and the quality of the entries are very promising. The number of generated entries could however be improved and we conclude the paper with the possible ways we plan to explore. 1.
Parsing is important in Linguistics and Natural Language Processing to understand the syntax and semantics of a natural language grammar. Parsing natural language text is challenging because of the problems like ambiguity and inefficiency. Also the interpretation of natural language text depends on context based techniques. A probabilistic component is essential to resolve ambiguity in both syntax and semantics thereby increasing accuracy and efficiency of the parser. Tamil language has some inherent features which are more challenging. In order to obtain the solutions, lexicalized and statistical approach is to be applied in the parsing with the aid of a language model. Statistical models mainly focus on semantics of the language which are suitable for large vocabulary tasks where as structural methods focus on syntax which models small vocabulary tasks. A statistical language model based on Trigram for Tamil language with medium vocabulary of 5000 words has been built. Though statistical parsing gives better performance through tri-gram probabilities and large vocabulary size, it has some disadvantages like focus on semantics rather than syntax, lack of support in free ordering of words and long term relationship. To overcome the disadvantages a structural component is to be incorporated in statistical language models which leads to the implementation of hybrid language models. This paper has attempted to build phrase structured hybrid language model which resolves above mentioned disadvantages. In the development of hybrid language model, new part of speech tag set for Tamil language has been developed with more than 500 tags which have the wider coverage. A phrase structured Treebank has been developed with 326 Tamil sentences which covers more than 5000 words. A hybrid language model has been trained with the phrase structured Treebank using immediate head parsing technique. Lexicalized and statistical parser which employs this hybrid language model and immediate head parsing technique gives better results than pure grammar and trigram based model.
<h3>Introduction</h3><br> Penn Discourse Treebank (PDTB) Version 3.0 is the third release in the Penn Discourse Treebank project, the goal of which is to annotate the Wall Street Journal (WSJ) section of Treebank-2 (<a href="http://catalog.ldc.upenn.edu/LDC95T7" rel="nofollow">LDC95T7</a>) with discourse relations. Penn Discourse Treebank Version 2 (<a href="../../../LDC2008T05" rel="nofollow">LDC2008T05</a>) contains over 40,600 tokens of annotated relations. In Version 3, an additional 13,000 tokens were annotated, certain pairwise annotations were standardized, new senses were included and the corpus was subject to a series of consistency checks. Details concerning the development of PDTB Version 3.0 can be found in the documentation accompanying this release. <br> Largely because the PDTB project was based on the idea that discourse relations are grounded in an identifiable set of explicit words or phrases (discourse connectives) or simply in the adjacency of two sentences, the PTDB has been used by many researchers in the natural language processing community and more recently, by researchers in psycholinguistics. It has also stimulated the development of similar resources in other languages and domains. <br> <h3>Data</h3><br> Annotations are provided in the form of separate text files (<em>standoff annotation</em>) that are byte-indexed into the raw WSJ text files in Treebank-2. The raw WSJ files are also included in this release. All text files are plain text, encoded in UTF-8. <br> This corpus contains two tools: (1) The Annotator, used for annotation and adjudication, and which can also be used for viewing the corpus; and (2) The Conversion Tool for converting Version 2 annotation files into the Version 3 format. <br> The documentation directory contains a manual describing what is new in Version 3 and how Version 3 differs from Version 2; the methods and guidelines used in annotating PDTB Version 3; and a range of statistics on the tokens, including the frequency of each connective, its sense labels and its modifiers. More information about the corpus and research carried out by the developers and others using the corpus can be found on the <a href="https://www.seas.upenn.edu/~pdtb/">PDTB website</a>. <br> <h3>Samples</h3><br> One can see samples of the annotation of different types of discourse relations, along with their visualization in the Annotator tool at: <br> <ul><br> <li><a href="desc/addenda/LDC2019T05_examples.html">Explicit relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#implicit_examples">Implicit relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#altlex_examples">Altlex and AltLexC relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#entrel_norel_examples">Entity relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#entrel_norel_examples">Hypophora relations</a></li><br> <li><a href="desc/addenda/LDC2019T05_examples.html#entrel_norel_examples">NoRel</a> (annotated only between adjacent sentences within a paragraph that are not linked to each other by a discourse relation)</li><br> </ul><br> <h3>Updates</h3><br> Experiments carried out in Fall 2019 on the intra-sentential discourse relations in the PDTB-3 revealed two problems with the corpus: (1) the final versions of two gold files of "to clause" annotation had not been loaded, and (2) several tokens were inadvertently omitted on the assumption that they were duplicates, when they were not. <br> Repairing these errors, and correcting a mis-labelled token in file wsj_1026, has added another 45 implicit intra-sentential relations to the corpus. Counts in the Annotation Manual have been adjusted to take these additional tokens into account. Specific changes/additions are recorded in the file "pdtb3-revision-jan-2020.txt". Downloads after February 3, 2020 contain the updated corpus. <br> <h3>Acknowledgment</h3><br> This work has been funded by the National Science Foundation, under grant NSF IIS 1422186 to the University of Pennsylvania and grant NSF IIS 1421067 to the University of Wisconsin, Milwaukee. The content of this publication does not necessarily reflect the position or policy of the Government, and no official endorsement should be inferred. </br> Portions © 1987-1989 Dow Jones & Company, Inc., © 2008, 2012, 2019 The Penn Discourse Treebank Group, © 2008, 2012, 2019 Trustees of the University of Pennsylvania
Abstract The semantic annotation of texts with senses from a computational lexicon is a complex and often subjective task. As a matter of fact, the fine granularity of the WordNet sense inventory [Fellbaum, Christiane (ed.). 1998. WordNet: An Electronic Lexical Database MIT Press], a de facto standard within the research community, is one of the main causes of a low inter-tagger agreement ranging between 70% and 80% and the disappointing performance of automated fine-grained disambiguation systems (around 65% state of the art in the Senseval-3 English all-words task). In order to improve the performance of both manual and automated sense taggers, either we change the sense inventory (e.g. adopting a new dictionary or clustering WordNet senses) or we aim at resolving the disagreements between annotators by dealing with the fineness of sense distinctions. The former approach is not viable in the short term, as wide-coverage resources are not publicly available and no large-scale reliable clustering of WordNet senses has been released to date. The latter approach requires the ability to distinguish between subtle or misleading sense distinctions. In this paper, we propose the use of structural semantic interconnections – a specific kind of lexical chains – for the adjudication of disagreed sense assignments to words in context. The approach relies on the exploitation of the lexicon structure as a support to smooth possible divergencies between sense annotators and foster coherent choices. We perform a twofold experimental evaluation of the approach applied to manual annotations from the SemCor corpus, and automatic annotations from the Senseval-3 English all-words competition. Both sets of experiments and results are entirely novel: structural adjudication allows to improve the state-of-the-art performance in all-words disambiguation by 3.3 points (achieving a 68.5% F1-score) and attains figures around 80% precision and 60% recall in the adjudication of disagreements from human annotators.
Many problems in Natural Language Processing (NLP) involves an efficient search for the best derivation over (exponentially) many candidates. For example, a parser aims to find the best syntactic tree for a given sentence among all derivations under a grammar, and a machine translation (MT) decoder explores the space of all possible translations of the source-language sentence. In these cases, the concept of packed forest provides a compact representation of huge search spaces by sharing common sub-derivations, where efficient algorithms based on Dynamic Programming (DP) are possible. Building upon the hypergraph formulation of forests and well-known 1-best DP algorithms, this dissertation develops fast and exact k-best DP algorithms on forests, which are orders of magnitudes faster than previously used methods on state-of-the-art parsers. We also show empirically how the improved output of our algorithms has the potential to improve results from parse reranking systems and other applications. We then extend these algorithms to approximate search when the forests are too big for exact inference. We discuss two particular instances of this new method, forest rescoring for MT decoding, and forest reranking for parsing. In both cases, our methods perform orders of magnitudes faster than conventional approaches. In the latter, faster search also leads to better learning, where our approximate decoding makes whole-Treebank discriminative training practical and results in an accuracy better than any previously reported systems trained on the Treebank. Finally, we apply the above materials to the problem of syntax-based translation and propose a new paradigm, forest-based translation. This scheme translates a packed forest of the source sentence into a target sentence, rather than just using 1-best or k -best parses as in usual practice. By considering exponentially many alternatives, it alleviates the propagation of parsing errors into translation, yet only comes with fractional overhead in running time. We also push this direction further to extract translation rules from packed forests. The combined results of forest-based decoding and rule extraction show significant improvements in translation quality with large-scale experiments, and consistently outperform the hierarchical system Hiero, one of the best performing systems to date.
Socio-economic decisions are commonly explained by rational cost versus benefit considerations, whereas person variables have not much been considered. The present study aimed at investigating the degree to which dispositional power motivation and affective states predict socio-economic decisions. The power motive was assessed both indirectly and directly using a TAT-like picture test and a power motive self-report, respectively. After 9 months, 62 students completed an affect rating and performed on a money allocation task (social values questionnaire). We hypothesized and confirmed that dispositional power should be associated with a tendency to maximize one’s profit but to care less about another party’s profit. Additionally, positive affect showed effects in the same direction. The results are discussed with respect to a motivational approach explaining socio-economic behaviour.
Traditionally, parsers are evaluated against gold standard test data. This can cause problems if there is a mismatch between the data structures and representations used by the parser and the gold standard. A particular case in point is German, for which two treebanks (TiGer and TüBa-D/Z) are available with highly different
A Treebank is a text corpus in which each sentence has been annotated with its syntactic structure. Although the construction of a treebank is an expensive task, we believe that it is indispensable for the development of real applications in the field of Natural Language Processing (NLP) and also for the development of the Information Society. At a purely linguistic level, the Treebank is an essential database for the study of a language given that it provides analyzed/annotated examples of real language. The linguistic study directly results in an improvement in the quality of several applications, such as Part-Of-Speech (POS) taggers and parsers (Collins 1997, 2000; Charniak 2000), because it provides common training and testing material allowing different algorithms to be compared and improved. In the last few years, treebank corpora such as the Penn Treebank (Marcus et al., 1993) and the Prague Dependency Treebank (Bohmova et al. 2003) have become a crucial resource for building and evaluating natural language processing tools and applications. As Abeille (2003) sets out, there are efforts underway for Czech, German, French, Japanese, Polish, Spanish and Turkish, to name just a few. In Kakkonen (2005) we can find the state of the art of dependency-based treebanks. The Basque Dependency Treebank (BDT) is actually the Reference Corpus for the Processing of Basque (EPEC) annotated at syntactic level. The EPEC is a 300,000 word corpus of standard written texts which aims to be a training corpus for the development and improvement of several NLP tools. It has been manually tagged at different levels: morphology, lemmatization and surface syntax (Aduriz et al. 2006). The next level of tagging —annotation of dependency relations— is currently being carried out in BDT. In this paper, we describe the annotation of noun phrase (henceforth, NP) constructions in detail following the Dependency Grammar theory (Tesniere 1959). For a better understanding of our work it should be noted that for us, NP is a purely descriptive term. We are not concerned with understanding the internal structure of NPs. The syntactic description of Basque NPs has been mainly developed within the generative framework by Goenaga (1980), Eguzkitza (1993), Laka (1993), Artiagoitia (2002),