Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 55-66. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
Parser disambiguation with precision grammars generally takes place via statistical ranking of the parse yield of the grammar using a supervised parse selection model. In the standard process, the parse selection model is trained over a hand-disambiguated treebank, meaning that without a significant investment of effort to produce the treebank, parse selection is not possible. Furthermore, as treebanking is generally streamlined with parse selection models, creating the initial treebank without a model requires more resources than subsequent treebanks. In this work, we show that, by taking advantage of the constrained nature of these HPSG grammars, we can learn a discriminative parse selection model from raw text in a purely unsupervised fashion. This allows us to bootstrap the treebanking process and provide better parsers faster, and with less resources. 1
Internet dating is now ranked third as the way people meet, behind meeting at work or school, and through a friend or family member. This study researches the use of social and linguistic norms in online dating advertisements. Previous research has posed that social groups create unique identities and group members will selectively present themselves in ways consistent with these identities. Using Craigslist to assess the similarities and differences between genders and sexualities in online personal postings, an online quiz-like survey was created. This research reports on people's abilities to predict the sexual orientation and gender of the writer based on linguistic cues.
We investigate parsing accuracy on the Korean Treebank 2.0 with a number of different grammars. Comparisons among these grammars and to their English counterparts suggest different aspects of Korean that contribute to parsing difficulty. Our results indicate that the coarseness of the Treebank’s nonterminal set is a even greater problem than in the English Treebank. We also find that Korean’s relatively free word order does not impact parsing results as much as one might expect, but in fact the prevalence of zero pronouns accounts for a large portion of the difference between Korean and English parsing scores. 1
Respondents can vary strongly in the way they use rating scales. Specifically, respondents can exhibit a variety of response styles, which threatens the validity of the responses. The purpose of this article is to investigate how response style and content of the items affect rating scale responses. The authors develop a novel model that accounts for different types of response styles, content of items, and background characteristics of respondents. By imposing a bilinear parameter structure on a multinomial logit model, the authors graphically distinguish the effects on the response behavior of the characteristics of a respondent and the content of an item. The authors combine this approach with finite mixture modeling, yielding two segmentations of the respondents: one for response style and one for item content. They apply this latent-class bilinear multinomial logit model to the well-known List of Values in a cross-national context. The results show large differences in the opinions and the response styles of respondents and reveal previously unknown response styles. Some response styles appear to be valid communication styles, whereas other response styles often concur with inconsistent opinions of the items and seem to be response bias.
Habituation is a fundamental form of learning manifested by a decrement of neuronal responses to repeated sensory stimulation. In addition, habituation is also known to occur on the behavioral level, manifested by reduced emotional reactions to repeatedly presented affective stimuli. It is, however, not clear which brain areas show a decline in activity during repeated sensory stimulation on the same time scale as reduced valence and arousal experience and whether these areas can be delineated from other brain areas with habituation effects on faster or slower time scales. These questions were addressed using functional magnetic resonance imaging acquired during repeated stimulation with piano melodies. The magnitude of functional responses in the laterobasal amygdala and in related cortical areas and that of valence and arousal ratings, given after each music presentation, declined in parallel over the experiment. In contrast to this long-term habituation (43 min), short-term decreases occurring within seconds were found in the primary auditory cortex. Sustained responses that remained throughout the whole investigated time period were detected in the ventrolateral prefrontal cortex extending to the dorsal part of the anterior insular cortex. These findings identify an amygdalocortical network that forms the potential basis of affective habituation in humans.
In this paper we describe FragmentSeeker, a tool which is capable to identify all those tree constructions which are recurring multiple times in a large Phrase Structure treebank. The tool is based on an efficient kernel-based dynamic algorithm, which compares every pair of trees of a given treebank and computes the list of fragments which they both share. We describe two different notions of fragments we will use, i.e. standard and partial fragments, and provide the implementation details on how to extract them from a syntactically annotated corpus. We have tested our system on the Penn Wall Street Journal treebank for which we present quantitative and qualitative analysis on the obtained recurring structures, as well as provide empirical time performance. Finally we propose possible ways our tool could contribute to different research fields related to corpus analysis and processing, such as parsing, corpus statistics, annotation guidance, and automatic detection of argument structure. 1.
Discontinuities occur especially frequently in languages with a relatively free word order, such as German. Generally, due to the longdistance dependencies they induce, they lie beyond the expressivity of Probabilistic CFG, i.e., they cannot be directly reconstructed by a PCFG parser. In this paper, we use a parser for Probabilistic Linear Context-Free Rewriting Systems (PLCFRS), a formalism with high expressivity, to directly parse the German NeGra and TIGER treebanks. In both treebanks, discontinuities are annotated with crossing branches. Based on an evaluation using different metrics, we show that an output quality can be achieved which is comparable to the output quality of PCFG-based systems. In most constituency treebanks, sentence annotation is restricted to having the shape of trees without crossing branches, and the non-local dependencies induced by the discontinuities are modeled by an additional mechanism. In the Penn Treebank (PTB) (Marcus et al., 1994), e.g., this mechanism is a combination of special labels and empty nodes, establishing implicit additional edges. In the German TüBa-D/Z (Telljohann et al., 2006), additional edges are established by a combination of topological field annotation and special edge labels. As an example, Fig. 1 shows a tree from TüBa-D/Z with the annotation of (1). Note here the edge label ON-MOD on the relative clause which indicates that the subject of the sentence (alle Attribute) is modified. 1
Studies of discourse relations have not, in the past, attempted to characterize what serves as evidence for them, beyond lists of frozen expressions, or markers, drawn from a few well-defined syntactic classes. In this paper, we describe how the lexicalized discourse relation annotations of the Penn Discourse Treebank (PDTB) led to the discovery of a wide range of additional expressions, annotated as AltLex (alternative lexicalizations) in the PDTB 2.0. Further analysis of AltLex annotation suggests that the set of markers is open-ended, and drawn from a wider variety of syntactic types than currently assumed. As a first attempt towards automatically identifying discourse relation markers, we propose the use of syntactic paraphrase methods.
In this paper, we present several ways to measure and evaluate the annotation and annotators, proposed and used during the building of the Czech part of the Prague Czech-English Dependency Treebank. At first, the basic principles of the treebank annotation project are introduced (division to three layers: morphological, analytical and tectogrammatical). The main part of the paper describes in detail one of the important phases of the annotation process: three ways of evaluation of the annotators- inter-annotator agreement, error rate and performance. The measuring of the inter-annotator agreement is complicated by the fact that the data contain added and deleted nodes, making the alignment between annotations non-trivial. The error rate is measured by a set of automatic checking procedures that guard the validity of some invariants in the data. The performance of the annotators is measured by a booking web application. All three measures are later compared and related to each other. 1.
Proceedings of the Workshop on Annotation and \nExploitation of Parallel Corpora AEPC 2010. \nEditors: Lars Ahrenberg, Jörg Tiedemann and Martin Volk. \nNEALT Proceedings Series, Vol. 10 (2010), 1-13. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15893.
The insula has been implicated as a component of central networks subserving evaluative and affective processes. This study examined evaluative valence and arousal ratings in response to picture stimuli in patients with lesions of the insula and two contrast groups: a control-lesion group (the primary contrast group) and an amygdala-lesion group. Patients rated the positivity and negativity of picture stimuli (from very unpleasant to very pleasant) and how emotionally arousing they found the pictures to be. Compared with patients in the control-lesion group, patients with insular lesions reported reduced arousal in response to both unpleasant and pleasant stimuli, as well as marked attenuation of valence ratings. In contrast, the arousal ratings of patients with amygdala lesions were selectively attenuated for unpleasant stimuli, and these patients' positive and negative valence ratings did not differ from those of the control-lesion group. Results support the view that the insular cortex may play a broad role in integrating affective and cognitive processes, whereas the amygdala may have a more selective role in affective arousal, especially for negative stimuli.
Cytotoxic T cell (CTL) covers several subtypes, which are CD8+, CD4 and CD4-CD8-. CTL derives from T cell repertoire in lymphoid hematopoietic stem cells. It matures in thymus and is activated in peripheral lymphoid tissues. Effector CTL kills the target cells by 2 ways. One is apoptotic effect mediated by FasL-Fas pathway and the other one is cytolytic effect mediated by granzymes. CTL has aroused great attention due to its significance in anti-tumor and anti-virus.
In this paper, we present a system that automatically extracts lexicalized tree adjoining grammars (LTAG) from treebanks. We first discuss in detail extraction algorithms and compare them to previous works. We then report the first LTAG extraction result for Vietnamese, using a recently released Vietnamese treebank. The implementation of an open source and language independent system for automatic extraction of LTAG grammars is also discussed. 1
The creation of language resources for less-resourced languages like the historical ones benefits from the exploitation of language-independent tools and methods developed over the years by many projects for modern languages. Along these lines, a number of treebanks for historical languages started recently to arise, including treebanks for Latin. Among the Latin treebanks, the Index Thomisticus Treebank is a 68,000 token dependency treebank based on the Index Thomisticus by Roberto Busa SJ, which contains the opera omnia of Thomas Aquinas (118 texts) as well as 61 texts by other authors related to Thomas, for a total of approximately 11 million tokens. In this paper, we describe a number of modifications that we applied to the dependency parser DeSR, in order to improve the parsing accuracy rates on the Index Thomisticus Treebank. First, we adapted the parser to the specific processing of Medieval Latin, defining an ad-hoc configuration of its features. Then, in order to improve the accuracy rates provided by DeSR, we applied a revision parsing method and we combined the outputs produced by different algorithms. This allowed us to improve accuracy rates substantially, reaching results that are well beyond the state of the art of parsing for Latin. 1.
Proceedings of the Ninth International Workshop \non Treebanks and Linguistic Theories. \nEditors: Markus Dickinson, Kaili Müürisep and Marco Passarotti. \nNEALT Proceedings Series, Vol. 9 (2010), 127-138. \n© 2010 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/15891.
This paper proposes a method of correcting annotation errors in a treebank. By using a synchronous grammar, the method transforms parse trees containing annotation errors into the ones whose errors are corrected. The synchronous grammar is automatically induced from the treebank. We report an experimental result of applying our method to the Penn Treebank. The result demonstrates that our method corrects syntactic annotation errors with high precision. 1
The Prague Dependency Treebank (henceforth PDT) is a large collection of texts in Czech. It contains several layers of rich annotation, ranging from morphology to deep syntax. It is unique in its size and theoretical background, especially for a language like Czech, which can be, with regard to the number of its speakers, considered a small language. In this article, we use PDT 2.0 to demonstrate that within real NLP systems, complex annotations may cut both ways. We present several issues that might pose problems when extracting data from PDT, and complex structures in general, and hint on possible solutions.
Latent variable grammars take an observed (coarse) treebank and induce more fine-grained grammar categories, that are better suited for modeling the syntax of natural languages. Estimation can be done in a generative or a discriminative framework, and results in the best published parsing accuracies over a wide range of syntactically divergent languages and domains. In this paper we highlight the commonalities and the differences between the two learning paradigms. 1
Vesicle templating presents a unique opportunity to construct submicrometer hollow particles. These authors give an overview of recent developments, discussing both polymerization inside the vesicle membrane (see Figure), and growth on the outer surface of the vesicle. Successful vesicle templating requires an understanding of the interactions between the vesicle bilayer, the polymer precursor, and the growing material.
This paper deals with the issue of linguistic database summaries understood as linguistically quantified propositions. We show basic capabilities of Quantirius, our interactive system supporting the mining and the assessment of linguistic summaries in a database. The mining of linguistic summaries is realized by means of the concept of a protoform introduced by Zadeh. The assessment of validity degrees of summaries is done via Zadeh's fuzzy logic based calculus, extended by the use of cardinality patterns and triangular norms. The main part of the paper presents an idea of a further processing of the set of generated summaries. The proposed algorithm is composed of a reduction mechanism of summaries based on linguistic terms inclusion and a reduction of summaries by means of the overlapping unimodal linguistic terms. We then present an idea of the approval threshold computing for the truth degrees of summaries. Finally, we show that summaries generated by different protoforms can be joined in a master-detail-like relation, making an interesting structure of information contained in the database.
In the architecture of a natural language processing system based on linguistic knowledge, two types of component are important: the knowledge databases and the processing modules. One of the knowledge databases is the lexical database, which is responsible for providing the lexical unities and its properties to the processing modules. The systems that process two or more languages require bilingual and/or multilingual lexical databases. These databases can be constructed by aligning distinct monolingual databases. In this paper, we present the interlingua and the strategy of aligning the two monolingual databases in REBECA, which only stores concepts from the “wheeled vehicle” domain.
This paper presents an efficient approach to use small syntactic category for constructing large tree with the assist of linguistic knowledge for parsing Thai language. VB-EM algorithm is exploited to adjust parameters of any trees for selecting the best tree. Our best result is 70.62% accuracy of bracketing recovery.
The paper introduces principles of rescript and lemmatization of lexemes with German origin for Lexical database of Baroque and Humanist Czech made in the Czech Language Institute of the Academy of Sciences of the Czech Republic in Prague.
The use of images of older people in the British advertising media has been under-researched to date. Further, previous research in any country has tended to examine such images from an a priori framework of general impressions and stereotypes of older people. This study addresses these issues with British consumers' (n = 106) impressions, trait ascriptions, and similarity-between-images ratings of a representative sample of U.K. magazine advertisements featuring older characters. After a series of sorting task laboratory sessions, multidimensional scaling and hierarchical cluster analyses revealed four clearly defined groups representing types of portrayals. These types emerged from the advertisements and from the views of the consumers themselves. These emergent groupings are: (1) Frail and Vulnerable, (2) Happy and Affluent, (3) Mentors, (4) Active and Leisure-oriented older adults. These groupings seem to be a logical context-appropriate derivation from previous findings on generally held stereotypes of older persons. It is argued that the groupings have the potential to contribute to a reliable typology of advertising portrayals of older people, with potential heuristic leverage in social scientific research of intergenerational communication, lifespan concerns, and the aging process.
We present the first effort towards producing an Arabic Discourse Treebank, a news corpus where all discourse connectives are identified and annotated with the discourse relations they convey as well as with the two arguments they relate. We discuss our collection of Arabic discourse connectives as well as principles for identifying and annotating them in context, taking into account properties specific to Arabic. In particular, we deal with the fact that Arabic has a rich morphology: we therefore include clitics as connectives as well as a wide range of nominalizations as potential arguments. We present a dedicated discourse annotation tool for Arabic and a large-scale annotation study. We show that both the human iden-tification of discourse connectives and the determination of the discourse relations they convey is reliable. Our current annotated corpus encompasses a final 5651 annotated discourse connectives in 537 news texts. In future, we will release the annotated corpus to other researchers and use it for training and testing automated methods for discourse connective and relation recognition. 1.
Finding a class of structures that is rich enough for adequate linguistic representation yet restricted enough for efficient computational processing is an important problem for dependency parsing. In this paper, we present a transition system for 2-planar dependency trees – trees that can be decomposed into at most two planar graphs – and show that it can be used to implement a classifier-based parser that runs in linear time and outperforms a stateof-the-art transition-based parser on four data sets from the CoNLL-X shared task. In addition, we present an efficient method for determining whether an arbitrary tree is 2-planar and show that 99 % or more of the trees in existing treebanks are 2-planar. 1
Abstract We investigate the use of instance-based ranking methods for surface realization in natural language generation. Our approach to instance-based natural language generation (IBNLG) employs two components: a rule system that ‘overgenerates’ a number of realization candidates from a meaning representation and an instance-based ranker that scores the candidates according to their similarity to examples taken from a training corpus. We develop an efficient search technique for identifying the optimal candidate based on a novel extension of the A * algorithm. The rule system is produced automatically from a semantically annotated fragment of the Penn Treebank II containing management succession texts. We detail the annotation scheme and grammar induction algorithm and evaluate the efficiency and output of the generator. We also discuss issues such as input coverage (completeness) and fluency that are relevant to surface generation in general.
We describe a method for the automatic extraction of a Stochastic Lexicalized Tree Insertion Grammar from a linguistically rich HPSG Treebank. The extraction method is strongly guided by HPSG–based head and argument decomposition rules. The tree anchors correspond to lexical labels encoding fine–grained information. The approach has been tested with a German corpus achieving a labeled recall of 77.33% and labeled precision of 78.27%, which is competitive to recent results reported for German parsing using the Negra Treebank.
In this paper, we present a novel approach to enhance hierarchical phrase-based machine translation systems with linguistically motivated syntactic features. Rather than directly using treebank categories as in previous studies, we learn a set of linguistically-guided latent syntactic categories automatically from a source-side parsed, word-aligned parallel corpus, based on the hierarchical structure among phrase pairs as well as the syntactic structure of the source side. In our model, each X nonterminal in a SCFG rule is decorated with a real-valued feature vector computed based on its distribution of latent syntactic categories. These feature vectors are utilized at decoding time to measure the similarity between the syntactic analysis of the source side and the syntax of the SCFG rules that are applied to derive translations. Our approach maintains the advantages of hierarchical phrase-based translation systems while at the same time naturally incorporates soft syntactic constraints.
Obesity prevalence in the U.S. has increased during the last three decades with major impact on public health. Screening for obesity in a population with unknown weight status can be time- and resource-consuming, but the information is valuable for prioritizing and allocating scarce resources. The challenge remains to properly assess obesity with the available methods. Body Image Rating Scales (BIRS) have initially been developed to assess body image disturbances, but also seem useful as an alternative method in assessing obesity prevalence. Several different BIRS exists. In this project I reviewed the literature that exists regarding the use of BIRS, and its advantages and limitations for the assessment of obesity status with regards to BMI. The result yielded nine publications that examined eight different scales and their correlation with BMI, ranging from r=.59 for self-reported BMI to r=.94 for measured BMI. One concern is the lack of standardization of this method to assess obesity, given the range of different scales. While many methods for obesity assessment are available, the simplicity, ease of use and cost-effectiveness of BIRS make it very appealing. BIRS remain a potentially attractive option to assess the weight status of a large population with minimal requirements in assets and time, especially in situations where measuring instruments are not available, or when height or weight could not be recalled.
In this paper, we compare two novel methods for part of speech tagging of Arabic without the use of gold standard word segmentation but with the full POS tagset of the Penn Arabic Treebank. The first approach uses complex tags without any word segmentation, the second approach is segmention-based, using a machine learning segmenter. Surprisingly, word-based POS tagging yields the best results, with a word accuracy of 94.74%. 1
We present a system that automatically induces Selectional Preferences (SPs) for Latin verbs from two treebanks by using Latin WordNet. Our method overcomes some of the problems connected with data sparseness and the small size of the input corpora. We also suggest a way to evaluate the acquired SPs on unseen events extracted from other Latin corpora. 1
Previous investigations of somatic hypersensitivity in IBS patients have typically involved only a single stimulus modality, and little information exists regarding whether patterns of somatic pain perception vary across stimulus modalities within a group of patients with IBS. Therefore, the current study was designed to characterize differences in perceptual responses to a battery of noxious somatic stimuli in IBS patients compared to controls. A total of 78 diarrhea-predominant and 57 controls participated in the study. We evaluated pain threshold and tolerance and sensory and affective ratings of contact thermal, mechanical pressure, ischemic stimuli, and cold pressor stimuli. In addition to assessing perceptual responses, we also evaluated differences in neuroendocrine and cardiovascular responses to these experimental somatic pain stimuli. A subset of IBS patients demonstrated the presence of somatic hypersensitivity to thermal, ischemic, and cold pressor nociceptive stimuli. The somatic hypersensitivity in IBS patients was somatotopically organized in that the lower extremities that share viscerosomatic convergence with the colon demonstrate the greatest hypersensitivity. There were also changes in ACTH, cortisol, and systolic blood pressure in response to the ischemic pain testing in IBS patients when compared to controls. The results of this study suggest that a more widespread alteration in central pain processing in a subset of IBS patients may be present as they display hypersensitivity to heat, ischemic, and cold pressor stimuli.
We compare self-training with and without reranking for parser domain adaptation, and examine the impact of syntactic parser adaptation on a semantic role labeling system. Although self-training without reranking has been found not to improve in-domain accuracy for parsers trained on the WSJ Penn Treebank, we show that it is surprisingly effective for parser domain adaptation. We also show that simple self-training of a syntactic parser improves out-of-domain accuracy of a semantic role labeler. 1
Abstract This paper aims to contribute to the current debate on ‘interculturality’ (IC) by investigating the process of language socialization whereby different generations of diasporic families negotiate, construct, and renew their sociocultural values and identities through interaction. Focusing on the use of address terms and ‘talk about social, cultural, and linguistic practice,’ the paper argues that IC is not only a dynamic process through which participants make aspects of their multiple and shifting identities relevant, but also a process of developing new social and cultural identities. In effect, it serves as a direct means of language socialization for the younger generation who are developing their sociocultural roles and learning about the social and cultural appropriateness of behavior in a diasporic context, where there are potentially substantial differences in social and cultural values between the wider local community and the diasporic community. Language socialization is regarded in this study as not simply about passing social and cultural values from one generation to another, but about bringing about changes in social and cultural values. Through language socialization, the younger generations of diasporic communities not only internalize the social, cultural and linguistic norms of their community, but also play an active role in constructing and creating their own social and cultural identities as well as bringing about changes to the existing community and family norms.
This paper proposes a method of correcting annotation errors in a treebank. By using a synchronous grammar, the method transforms parse trees containing annotation errors into the ones whose errors are corrected. The synchronous grammar is automatically induced from the treebank. We report an experimental result of applying our method to the Penn Treebank. The result demonstrates that our method corrects syntactic annotation errors with high precision.