Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
National audience
There is evidence that women may be less successful when attempting to quit smoking than men. One potential contributory cause of this gender difference is differential craving and stress reactivity to smoking- and negative affect/stress-related cues. The present human laboratory study investigated the effects of gender on reactivity to smoking and negative affect/stress cues by exposing nicotine dependent women (n = 37) and men (n = 53) smokers to two active cue types, each with an associated control cue: (1) in vivo smoking cues and in vivo neutral control cues, and (2) imagery-based negative affect/stress script and a neutral/relaxing control script. Both before and after each cue/script, participants provided subjective reports of smoking-related craving and affective reactions. Heart rate (HR) and skin conductance (SC) responses were also measured. Results indicated that participants reported greater craving and SC in response to smoking versus neutral cues and greater subjective stress in response to the negative affect/stress versus neutral/relaxing script. With respect to gender differences, women evidenced greater craving, stress and arousal ratings and lower valence ratings (greater negative emotion) in response to the negative affect/stressful script. While there were no gender differences in responses to smoking cues, women trended towards higher arousal ratings. Implications of the findings for treatment and tobacco-related morbidity and mortality are discussed.
This paper describes how electronic grammars can be further enhanced by adding machine-readable grammars and treebanks. We explore the potential benefits of im- plemented grammars and treebanks for descriptive linguistics, following the discursive methodology of Bird & Simons (2003) and the values and maxims identified by Nordhoff(2008). We describe the resources which we believe make implemented grammars and treebanks feasible additions to electronic descriptive grammars, with a particular focus on the Grammar Matrix grammar customization system (Bender et al. 2010) and the Fangorn treebank search application (Ghodke & Bird 2010). By presenting an ex- ample of an implemented grammar based on a descriptive prose grammar, we show one productive method of collaboration between grammar engineer and field linguist, and propose that a tighter integration could be beneficial to both, creating a virtuous cycle that could lead to more effective and informative resources.
International audience
This three-variable predictive model has excellent predictive ability in both the derivation cohort and the validation cohort. This model can identify women who are at high risk of non-initiating breastfeeding within the first hour after delivery.
International audience
A treebank may contain the annotation of different phenomena such as word order, morphological features, syntactic and semantic relations, etc., which are rather different in their nature. Quite often, the annotation of these phenomena is combined in a single structure, which leads to low-quality training results and is verifiably deficient from a theoretical (linguistic) perspective. We argue that the annotation of corpora requires a well-defined linguistic model which supports multi-level annotation, with one type of phenomenon per level. Our experience with dependency treebanks created or adjusted for surface-oriented natural language generation and based on the Meaning-Text Theory, a multi-level linguistic model, supports this argumentation.
Treebanks are language resources that provide annotations at various levels of linguistic structure starting from the word level. They typically provide syntactic constituent or dependency structures for sentences, but increasingly extend to annotation beyond syntactic structure, including semantic, pragmatic and rhetorical annotation, or go beyond a single language, as in parallel treebanks.
 Experience in building treebanks has shown that there is a close relation between formal linguistic theory and the design and practice of annotation. With increasing complexity of annotations, the design of annotation schemes becomes more and more theory-dependent. At the same time, linguistically motivated treebank annotations have become crucially important for the development of data-driven approaches to natural language processing and for linguistic research in general.
 Treebanks therefore constitute an important link between linguistic theory and computational linguistics.
 The International Workshop on Treebanks and Linguistic Theories provides a forum for researchers working on treebanks from both perspectives. The present volume presents the contents of the 10th edition of this workshop series, held in 2012 at the University of Heidelberg.
Language resources are essential for linguistic research and the development of NLP applications. Low-density languages, such as Irish, therefore lack significant research in this area. This paper describes the early stages in the development of new language resources for Irish – namely the first Irish dependency treebank and the first Irish statistical dependency parser. We present the methodology behind building our new treebank and the steps we take to leverage upon the few existing resources. We discuss language-specific choices made when defining our dependency labelling scheme, and describe interesting Irish language characteristics such as prepositional attachment, copula and clefting. We manually develop a small treebank of 300 sentences based on an existing POS-tagged corpus and report an inter-annotator agreement of 0.7902. We train MaltParser to achieve preliminary parsing results for Irish and describe a bootstrapping approach for further stages of development.
Recurrent neural network language models (RNNLMs) have recently demonstrated state-of-the-art performance across a variety of tasks. In this paper, we improve their performance by providing a contextual real-valued input vector in association with each word. This vector is used to convey contextual information about the sentence being modeled. By performing Latent Dirichlet Allocation using a block of preceding text, we achieve a topic-conditioned RNNLM. This approach has the key advantage of avoiding the data fragmentation associated with building multiple topic models on different data subsets. We report perplexity results on the Penn Treebank data, where we achieve a new state-of-the-art. We further apply the model to the Wall Street Journal speech recognition task, where we observe improvements in word-error-rate.
Individuals with autism spectrum disorders (ASD) demonstrate increased visual attention and elevated brain reward circuitry responses to images related to circumscribed interests (CI), suggesting that a heightened affective response to CI may underlie their disproportionate salience and reward value in ASD. To determine if individuals with ASD differ from typically developing (TD) adults in their subjective emotional experience of CI object images, non-CI object images and social images, 213 TD adults and 56 adults with ASD provided arousal ratings (sensation of being energized varying along a dimension from calm to excited) and valence ratings (emotionality varying along dimension of approach to withdrawal) for a series of 114 images derived from previous research on CI. The groups did not differ on arousal ratings for any image type, but ASD adults provided higher valence ratings than TD adults for CI-related images, and lower valence ratings for social images. Even after co-varying the effects of sex, the ASD group, but not the TD group, gave higher valence ratings to CI images than social images. These findings provide additional evidence that ASD is characterized by a preference for certain categories of non-social objects and a reduced preference for social stimuli, and support the dissemination of this image set for examining aspects of the circumscribed interest phenotype in ASD.
Abstract This study examines the role of the voseo, tuteo, and ustedeo in written advertising in Montevideo, Uruguay. Analysis of 133 samples shows a preference for voseo in publicity, although tú and usted are also present. While written voseo - with or without a tonic pronoun - has traditionally been stigmatized in public education (Bertolotti & Coll 2003 and Gabbiani 2000), its predominant role in advertising suggests an increased acceptance of that written form in public venues. The study examines the language of publicity within the framework of politeness strategies, which explains that written publicity uses linguistic forms that lessen social distance and power [-D, -P] for the purpose of persuading consumers to buy. As such, voseo is the preferred form of address in commercial advertising in Montevideo, appearing in 83% of the samples. Voseo also has a strong presence in non-commercial advertising, where it appears in 44% of the examples in this study, irrespective of any stigma attached to its use.
We propose a method for the extraction of a Tree Adjoining Grammar (TAG) from a dependency treebank which has some representative examples annotated with phrase structures. We show that the resulting TAG along with corresponding dependency structure can be used to convert a dependency treebank to a TAG-based phrase structure treebank.
The paper presents an integral framework for multilingual lexical databases (henceforth MLLD) based on Compreno technology. It differs from the existing approaches to MLLD in the following aspects: 1) it is based on a universal semantic hierarchy (SH) of thesaurus type filled with language-specific lexicon; 2) the position in the SH generally determines semantic and syntactic model of a word; 3) this model proposes a suite of elaborate tools to determine universal and language-specific semantic and syntactic properties and deals efficiently with problems of cross-lingual lexical, semantic and syntactic asymmetry. Currently, it includes English, Russian, German, French and Chinese and proves to be a compatible MLLD for typologically different languages that can be used as a comprehensive lexical-semantic database for various NLP applications.
We present the ongoing development of MCG, a linguistically deep and precise grammar for Mandarin Chinese together with its accompanying treebank, both based on the linguistic framework of HPSG, and using MRS as the semantic representation. We highlight some key features of our grammar design, and review a number of challenging phenomena, with comparisons to alternative linguistic treatments and implementations. One of the distinguishing characteristics of our approach is the tight integration of grammar and treebank development. The two-step treebank annotation procedure benefits from the efficiency of the discriminant-based annotation approach, while giving the annotators full freedom of producing extra-grammatical structures. This not only allows the creation of a precise and full-coverage treebank with an imperfect grammar, but also provides prompt feedback for grammarians to identify the errors in the grammar design and implementation. Preliminary evaluation and error analysis shows that the grammar already covers most of the core phenomena for Mandarin Chinese, and the treebank annotation procedure reaches a stable speed of 35 sentences per hour with satisfying quality. Keywords:Grammar Engineering, Treebank Annotation, Syntax 1.
A treebank is an important resource for developing many NLP based tools. Errors in the treebank may lead to error in the tools that use it. It is essential to ensure the quality of a treebank before it can be deployed for other purposes. Automatic (or semi-automatic) detection of errors in the treebank can reduce the manual work required to find and remove errors. Usually, the errors found automatically are manually corrected by the annotators. There is not much work reported so far on error correction tools which helps the annotators in correcting errors efficiently. In this paper, we present such an error correction tool that is an extension of the error detection method described earlier (Ambati et al., 2010; Ambati et al., 2011; Agarwal et al., 2012). Keywords:Treebank, Error Detection, Graphical User Interface
Finding coordinations provides useful infor-mation for many NLP endeavors. However, the task has not received much attention in the literature. A major reason for that is that the annotation of major treebanks does not re-liably annotate coordination. This makes it virtually impossible to detect coordinations in which two conjuncts are separated by punctu-ation rather than by a coordinating conjunc-tion. In this paper, we present an annotation scheme for the Penn Treebank which intro-duces a distinction between coordinating from non-coordinating punctuation. We discuss the general annotation guidelines as well as prob-lematic cases. Eventually, we show that this additional annotation allows the retrieval of a considerable number of coordinate structures beyond the ones having a coordinating con-junction.
We describe our method of traditional Phrase Structure Grammar (PSG) parsing in CIPS-Bakeoff2012 Task3. First, bagging is proposed to enhance the baseline performance of PSG parsing. Then we suggest exploiting another TreeBank (CTB7.0) to improve the performance further. Experimental results on the development data set demonstrate that bagging can boost the baseline F1 score from 81.33 % to 84.41%. After exploiting the data of CTB7.0, the F1 score reaches 85.03%. Our final results on the official test data set show that the baseline closed system using bagging gets the F1 score of 80.17%. It outperforms the best closed system by nearly 4 % which uses a single model. After exploiting the CTB7.0 data, the F1 score reaches 81.16%, demonstrating further increases of about 1%. 1
In this paper, we describe an ongoing research to develop an HPSG-based treebank for Persian. To this aim, we use a bootstrapping approach for the data annotation. In the first step, a set of seed rules are defined as regular expressions in the CLaRK system. Then, the data is shallow processed with this set of rules. In the next step, a human annotator completes the annotation of sentences manually. To increase automatic annotation, we extract the manual applied rules and iteratively augment the seed rules with the rules applied frequently in the manual annotation. Our experiment in building the Persian treebank which currently contains 1000 sentences shows that the proposed method reduces human intervention from 74.05% in first iterations to 39.01% in last iterations.
Minimal hepatic encephalopathy (MHE) encompasses a number of neuropsychological and neurophysiological disorders in patients suffering from liver cirrhosis, who do not display abnormalities during a medical interview or physical examination. A negative influence of MHE on the quality of life of patients suffering from liver cirrhosis was confirmed, which include retardation of ability of operating motor vehicles and disruption of multiple health-related areas, as well as functioning in the society. The data on frequency of traffic offences and accidents amongst patients diagnosed with MHE in comparison to patients diagnosed with liver cirrhosis without MHE, as well as healthy persons is alarming. Those patients are unaware of their disorder and retardation of their ability to operate vehicles, therefore it is of utmost importance to define this group. The term minimal hepatic encephalopathy (formerly "subclinical" encephalopathy) erroneously suggested the unnecessity of diagnostic and therapeutic procedures in patients with liver cirrhosis. Diagnosing MHE is an important predictive factor for occurrence of overt encephalopathy - more than 50% of patients with this diagnosis develop overt encephalopathy during a period of 30 months after. Early diagnosing MHE gives a chance to implement proper treatment which can be a prevention of overt encephalopathy. Due to continuing lack of clinical research there exist no commonly agreed-upon standards for definition, diagnostics, classification and treatment of hepatic encephalopathy. This article introduces the newest findings regarding the importance of MHE, scientific recommendations and provides detailed descriptions of the most valuable diagnostic methods.
International audience
This paper describes a method to convert existing treebanks with syntactic information into banks of meaning representations. The central component is a system of evaluation for a small formal language with respect to an information state. Inputs to the evaluation system are formal language expressions obtained from the conversion of parsed representations conforming to (Penn Treebank Project) guidelines. Outputs from the evaluation system are Davidsonian (higher-order) predicate logic meaning representations. Having a system of evaluation as the basis for generating meaning representations makes possible accepting input with minimal conversion from existing treebanks and from the tools used to construct treebanks. Results of having built corresponding banks of meaning representations from available treebanks are discussed.
We present a number of experiments on parsing the Ancient Greek Dependency Treebank (AGDT), i.e. the largest syntactically annotated corpus of Ancient Greek currently available (350k words ca). Although the AGDT is rather unbalanced and far from being representative of all genres and periods of Ancient Greek, no attempt has been made so far to perform automatic dependency parsing of Ancient Greek texts. By testing and evaluating one probabilistic dependency parser (MaltParser), we focus on how to improve the parsing accuracy and how to customize a feature model that fits the distinctive properties of Ancient Greek syntax. Also, we prove the impact of genre and author diversity on parsing performances.
Sanskrit since many thousands of years has been the oriental language of India. It is the base for most of the Indian Languages. Ambiguity is inherent in the Natural Language sentences. Here, one word can be used in multiple senses. Morphology process takes word in isolation and fails to disambiguate correct sense of a word. Part-Of-Speech Tagging (POST) takes word sequences in to consideration to resolve the correct sense of a word present in the given sentence. Efficient POST have been developed for processing of English, Japanese, and Chinese languages but it is lacking for Indian languages. In this paper our work present simple rule-based POST for Sanskrit language. It uses rule based approach to tag each word of the sentence. These rules are stored in the database. It parses the given Sanskrit sentence and assigns suitable tag to each word automatically. We have tested this approach for 15 tags and 100 words of the language this rule based tagger gives correct tags for all the inflected words in the given sentence.
Most of the reliable language resources are developed via human supervision. Developing supervised annotated data is hard and tedious, and it will be very time consuming when it is done totally manually; as a result, various types of annotated data, including treebanks, are not available for many languages. Considering that a portion of the language is regular, we can define regular expressions as grammar rules to recognize the strings which match the regular expressions, and reduce the human effort to annotate further unseen data. In this paper, we propose an incremental bootstrapping approach via extracting grammar rules when no treebank is available in the first step. Since Persian suffers from lack of available data sources, we have applied our method to develop a treebank for this language. Our exper-iment shows that this approach significantly decreases the amount of manual effort in the annotation process while enlarging the treebank. Keywords:Treebank Development, Bootstrapping Approach, Grammar Rule Extraction, the Persian Language 1.
The aim here is to create a dependency treebank from a phrase-structure treebank for Arabic. Arabic has a number of characteristics,described below, which make it particularly challenging to any natural language processing (NLP) applications. We describe an encouraging semi-automatic technique for converting phrase-structure trees to dependency trees by using a head percolation table.One of the most significant challenges here is the determination of the head of each subtree. We therefore examined different versionsof the head percolation table to find the best priority list for each entry in the table. Given that there is no absolute measure of the‘correctness’ of a conversion of a phrase structure tree to dependency form, we tested the various transformations by seeing how well astate-of-the-art dependency parser learnt the generalisations that were embodied by the converted trees.
Conversion between different grammar frameworks is of great importance to comparative performance analysis of the parsers developed based on them and to discover the essential nature of languages. This paper presents an approach that converts Combinatory Categorial Grammar (CCG) derivations to Penn Treebank (PTB) trees using a maximum entropy model. Compared with previous work, the presented technique makes the conversion practical by eliminating the need to develop mapping rules manually and achieves state-of-the-art results.
Treebanks are a linguistic resource: a large database where the morphological, syntactic and lexical information for each sentence has been explicitly marked. The critical requirements of treebanks for various NLP activities (research and application) are well known. This also implies that treebanks need to be as error free as possible. However, manual validation of a treebank is very costly, both in terms of time and money. This paper describes an approach to automatically detect errors in a treebank after a complete manual annotation. Over and above improving an earlier error detection tool (Ambati et al. (2011)) for a Hindi treebank. We also present a user study to show that our system reduces the validation time significantly while detecting 81.49% of the errors at the dependency level.
This work describes how derivation tree fragments based on a variant of Tree Adjoining Grammar (TAG) can be used to check treebank consistency. Annotation of word sequences are compared both for their internal structural consistency, and their external relation to the rest of the tree. We expand on earlier work in this area in three ways. First, we provide a more complete description of the system, showing how a naive use of TAG structures will not work, leading to a necessary refinement. We also provide a more complete account of the processing pipeline, including the grouping together of structurally similar errors and their elimination of duplicates. Second, we include the new experimental external relation check to find an additional class of errors. Third, we broaden the evaluation to include both the internal and external relation checks, and evaluate the system on both an Arabic and English treebank. The evaluation has been successful enough that the internal check has been integrated into the standard pipeline for current English treebank construction at the
In this paper, we propose a scheme for anaphora annotation in Hindi Dependency Treebank. The goal is to identify and handle the challenges that arise in the annotation of reference relations in Hindi. We identify some of the issues related to anaphora annotation specific to Hindi such as distribution of markable span, sequential annotation, representation format, annotation of multiple referents etc. The scheme hence incorporates some characteristics specific to these issues in order to achieve a consistent annotation. Most significant among these characteristics is the head-modifier separation in referent selection. The modifier-modified dependency relations inside a markable is utilized for this headmodifier distinction. A part of the Hindi Dependency Treebank, of around 2500 sentences has been annotated with anaphoric relations and an inter-annotator study was carried out which shows a significant agreement over selection of the head referent using the proposed scheme as compared to MUC annotation format. The current annotation is done for a limited set of pronominal categories.
A method is presented for transferring dependency treebanks between similar languages by using a bilingual lexicon, aiming to improve dependency parsing accuracy on the target language. It is illustrated by transferring the Slovene Dependency Treebank to Croatian by using a GIZA++ bilingual lexicon constructed from the Croatian-Slovene 1984 parallel corpus from the Multext East project. The transferred treebank is merged with the Croatian Dependency Treebank and the merged treebank is used to train and test two graph-based dependency parsers. MSTParser and CroDep accuracy on parsing the 1984 fictional text shows a statistically significant increase and a similar decrease on parsing the Croatian Dependency Treebank newspaper text.
We describe a shared task on parsing web text from the Google Web Treebank. Participants were to build a single parsing system that is robust to domain changes and can handle noisy text that is commonly encountered on the web. There was a constituency and a dependency parsing track and 11 sites submitted a total of 20 systems. System combination approaches achieved the best results, however, falling short of newswire accuracies by a large margin. The best accuracies were in the 80-84% range for F1 and LAS; even part-ofspeech accuracies were just above 90%.
International audience
We consider linguistic database summaries in the sense of Yager (1982), in an implementable form proposed by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth) and their extensions as a very general tool for a human consistent summarization of large data sets. We advocate the use of the concept of a protoform (prototypical form), vividly advocated by Zadeh and shown by Kacprzyk & Zadrozny (2005) as a general form of a linguistic data summary. Then, we present an extension of our interactive approach to fuzzy linguistic summaries, based on fuzzy logic and fuzzy database queries with linguistic quantifiers. We show how fuzzy queries are related to linguistic summaries, and that one can introduce a hierarchy of protoforms, or abstract summaries in the sense of latest Zadeh’s (2002) ideas meant mainly for increasing deduction capabilities of search engines. We show an implementation for the summarization of Web server logs.
We present two approaches (rule-based and statistical) for automatically annotating intra-chunk dependencies in Hindi. The intra-chunk dependencies are added to the dependency trees for Hindi which are already annotated with inter-chunk dependencies. Thus, the intra-chunk annotator finally provides a fully parsed dependency tree for a Hindi sentence. In this paper, we first describe the guidelines for marking intra-chunk dependency relations. Although the guidelines are for Hindi, they can easily be extended to other Indian languages. These guidelines are used for framing the rules in the rule-based approach. For the statistical approach, we use MaltParser, a data driven parser. A part of the ICON 2010 tools contest data for Hindi is used for training and testing the MaltParser. The same set is used for testing the rule-based approach.
We have created a bilingual treebank for 99% of the sentences in the FraCaS test suite. The treebank is built together with an associated bilingual English-Swedish lexicon written in the Grammatical Framework Resource Grammar. The original FraCaS sentences are English, and we have tested the multilinguality of the Resource Grammar by analysing the grammaticality and naturalness of the Swedish translations. 86% of the sentences are grammatically and semantically correct and sound natural. About 10% can probably be fixed by adding new lexical items or grammatical rules, and only a small amount are considered to be difficult to cure.\n
This article presents an online dictionary environment, with enhanced sorting and searching functionalities and a text to speech feature, for hearing the pronunciation of the words. The online dictionary environment has been developed as part of the ‘Syntychies’ research program. ‘Syntychies’ online environment is a pioneering webservice for Greek dialectal lexicography and it is the first of its kind for Cypriot Greek.
To facilitate future research in unsupervised induction of syntactic structure and to standardize best-practices, we propose a tagset that consists of twelve universal part-of-speech categories. In addition to the tagset, we develop a mapping from 25 different treebank tagsets to this universal set. As a result, when combined with the original treebank data, this universal tagset and mapping produce a dataset consisting of common parts-of-speech for 22 different languages. We highlight the use of this resource via two experiments, including one that reports competitive accuracies for unsupervised grammar induction without gold standard part-of-speech tags.
The Latvian Treebank is being developed since 2010. In this paper we describe the latest developments of this project and the problems currently faced. We examine several gaps in our annotation scheme like determinant, ellipsis and insertion annotation and describe solutions we have chosen.
In the article, the currently existing codified orthographic norms are presented with reference to four selected problem topics and paralleled with certain other norms significantly influencing the synchronous usage of written language. The expression “norm” or “norms” is to be understood in its broadest sense: at one end of the continuum as a formalized set of rules and regulations which, if ignored, may result in sanctions, and, at the other end, as a set of principles and guidelines organized according to unified rules and affecting, or regulating, any one sphere of human activity and behaviour. The discussion includes other possible reasons for the discrepancies between orthographically prescribed use and use deviating from it. Despite the weight of these reasons, the non-linguistic norms presented in this article appear to be an influential factor that can by no means be overlooked by contemporary normative linguistics.
The focus of this article is on the creation of a collection of sentences manually annotated with respect to their sentence structure. We show that the concept of linear segments—linguistically motivated units, which may be easily detected automatically—serves as a good basis for the identification of clauses in Czech. The segment annotation captures such relationships as subordination, coordination, apposition and parenthesis; based on segmentation charts, individual clauses forming a complex sentence are identified. The annotation of a sentence structure enriches a dependency-based framework with explicit syntactic informa- tion on relations among complex units like clauses. We have gathered a collection of 3,444 sentences from the Prague Dependency Treebank, which were annotated with respect to their sentence structure (these sentences comprise 10,746 segments forming 6,341 clauses). The main purpose of the project is to gain a development data—promising results for Czech NLP tools (as a dependency parser or a machine translation system for related languages) that adopt an idea of clause segmentation have been already reported. The collection of sentences with annotated sentence structure provides the possibility of further improvement of such tools.
With english speaking population expanding rapidly due to increasing international communication, english is no longer spoken by native speakers alone. English is spoken in different regions around the world, developing into different variants reflecting local language and culture. When speakers in an international conference speak in non-native english, interpreters, unfamiliar to such varieties, may be challenged as unfamiliarity to specific variants is more likely to present difficulties in intelligibility and comprehensibility. Previous studies indicate that conference interpreters found unfamiliar accents challenging and difficult to interpret. This paper discusses stress factors for English-Korean interpreters presented by non-native english speakers, following the concept of ‘World Englishes,’ which refers to different varieties developed around the world reflecting cultural and linguistic norms of the region or society of speakers. This paper aims to identify previous researches that point to the challenges of interpreting performance in general, and world englishes, in particular, to develop a theoretical framework for analyzing specific difficulties experienced by English-Korean conference interpreters at the level of phonology, syntax, and processing effort