Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Abstract When survey respondents rate the quality of life (QoL) associated with a health condition, they must not only evaluate the health condition itself, but must also interpret the meaning of the rating scale in order to assign a specific value. The way that respondents approach this task depends on subjective interpretations, resulting in inconsistent results across populations and tasks. In particular, patients and non-patients often give very different ratings to health conditions, a discrepancy that raises questions about the objectivity of either groups’ evaluations. In this study, we found that the perspective of the raters (i.e., their own current health relative to the health conditions they rated) influences the way they distinguish between different health states that vary in severity. Consistent with prospect theory, a mild and a severe lung disease scenario were rated quite differently by lung disease patients whose own health falls between the two scenarios, whereas healthy non-patients, whose own health was better than both scenarios, rated the two scenarios as much more similar. In addition, we found that the context of the rating task influences the way participants distinguish between mild and severe scenarios. Both patients and non-patients gave less distinct ratings to the two scenarios when each were presented in isolation than when they were presented alongside other scenarios that provided contextual information about the possible range of severity for lung disease. These results raise continuing concerns about the reliability and validity of subjective QoL ratings, as these ratings are highly sensitive to differences between respondent groups and the particulars of the rating task.
This paper discusses a novel probabilistic synchronous TAG formalism, synchronous Tree Substitution Grammar with sister adjunction (TSG+SA). We use it to parse a language for which there is no training data, by leveraging off a second, related language for which there is abundant training data. The grammar for the resource-rich side is automatically extracted from a treebank; the grammar on the resource-poor side and the synchronization are created by handwritten rules. Our approach thus represents a combination of grammar-based and empirical natural language processing. We discuss the approach using the example of Levantine Arabic and Standard Arabic.
This study uses a spatial logit model to evaluate the statistical effect of conditions of communities on municipal bond ratings. It finds that private (non-farm) earnings in the community positively explain bond ratings with statistical significance, while earnings from personal transfers negatively affect ratings. Own source of revenues of local governments (local taxes) increase ratings and inter-governmental revenues (transfers to local governments) negatively impact ratings. Outstanding debt fails to significantly explain ratings. The composition of the local economy (e.g., the service sector) weights heavily in a rating, and proximity of the local government to areas with high municipal bond ratings increases ratings.
Each year the Conference on Computational Natural Language Learning (CoNLL) features a shared task, in which participants train and test their systems on exactly the same data sets, in order to better compare systems. The tenth CoNLL (CoNLL-X) saw a shared task on Multilingual Dependency Parsing. In this paper, we describe how treebanks for 13 languages were converted into the same dependency format and how parsing performance was measured. We also give an overview of the parsing approaches that participants took and the results that they achieved. Finally, we try to draw general conclusions about multi-lingual parsing: What makes a particular language, treebank or annotation scheme easier or harder to parse and which phenomena are challenging for any dependency parser?
Recently, the Prague Dependency Treebank 2.0 (PDT 2.0) has emerged as the largest text corpora annotated on the level of tectogrammatical representation (“linguistic meaning”) described in Sgall et al. (2004) and containing about 0.8 milion words (see Hajič (2004)). We hope that this level of annotation is so close to the meaning of the utterances contained in the corpora that it should enable us to automatically transform texts contained in the corpora to the form of knowledge base, usable for information extraction, question answering, summarization, etc. We can use Multilayered Extended Semantic Networks (MultiNet) described in Helbig (2006) as the target formalism. In this paper we discuss the suitability of such approach and some of the main issues that will arise in the process. In section 1. we introduce formalisms underlying PDT 2.0 and MultiNet, in section 2. we describe the role MultiNet can play in the system of Functional Generative Description (FGD), section 3. discusses issues of automatic conversion to MultiNet and section 4. gives some conclusions. 1.
Statistical parsers trained and tested on the Penn Wall Street Journal (WSJ) treebank have shown vast improvements over the last 10 years. Much of this improvement, however, is based upon an ever-increasing number of features to be trained on (typically) the WSJ treebank data. This has led to concern that such parsers may be too finely tuned to this corpus at the expense of portability to other genres. Such worries have merit. The standard "Charniak parser" checks in at a labeled precision-recall f-measure of 89.7% on the Penn WSJ test set, but only 82.9% on the test set from the Brown treebank corpus.This paper should allay these fears. In particular, we show that the reranking parser described in Charniak and Johnson (2005) improves performance of the parser on Brown to 85.2%. Furthermore, use of the self-training techniques described in (McClosky et al., 2006) raise this to 87.8% (an error reduction of 28%) again without any use of labeled Brown data. This is remarkable since training the parser and reranker on labeled Brown data achieves only 88.4%.
The issue of information structure in language has been studied extensively both in the Prague School of Linguistics and in the Functional Generative Description [FGD, 11, 5]. This theory of representation of linguistic meaning is the framework for a family of multi-level annotation projects, in particular the Prague Dependency Treebank for Czech [PDT, 2, 3] and the Prague Arabic Dependency Treebank [PADT, 4, 12]. Information structure — the question of ‘the given’ and ‘the new’ in an utterance and how it is expressed — is considered to contribute to the linguistic meaning, and its annotation in PDT is part of the third, the most detailed and abstract level of linguistic description. Next to determining which elements in a sentence are context-bound and which are non-bound (the elementary distinctive feature from which the topic–focus dichotomy is derived), attention is also paid to the resolution of anaphoric relations and to capturing the communicative dynamism of a proposition (annotations of coreference and deep word order, respectively) [9]. In PADT, which now consists of the morphological and the analytical levels of description of Arabic, a similar annotation of information structure is being established. In our contribution, we would like to overview the theoretical concepts we work with, and present our formal treatment of a number of prototypical, yet corpus-based, instances of linguistic phenomena that have a principal impact on the structure of information in Arabic [6, 1]. The applicability of the general approach to written as well as spoken Arabic will be the main point of our account. In FGD, the description of information structure incorporates also the notions of intonation center and stress, contrast, subjective word order, or potential ellipsis, which are directly connected to prosody. We will provide references to other related computational research, too [cf. e.g. 8, 7, 10].
We report on a series of experiments with probabilistic context-free grammars predicting English and German syllable structure. The treebank-trained grammars are evaluated on a syllabification task. The grammar used by As she evaluates the grammar only for German, we reimplement the grammar and experiment with additional phonotactic features. Using bi-grams within the syllable, we can model the dependency from the previous consonant in the onset and coda. A 10fold cross validation procedure shows that syllabification can be improved by incorporating this type of phonotactic knowledge. Compared to the grammar of Mller (2002), syllable boundary accuracy increases from 95.8% to 97.2% for English, and from 95.9% to 97.2% for German. Moreover, our experiments with different syllable structures point out that there are dependencies between the onset on the nucleus for German but not for English. The analysis of one of our phonotactic grammars shows that interesting phonotactic constraints are learned. For instance, unvoiced consonants are the most likely first consonants and liquids and glides are preferred as second consonants in two-consonantal onsets.
Computational Modeling of Bilingualism Symposium Organizer: Ping Li (pli@richmond.edu) Department of Psychology, University of Richmond Richmond, VA 23173 USA connectionist developmental lexical model (Li et al., 2004). It considers learner variables (e.g., time of L2 learning and proficiency) and input variables (word types and bilingual distance) to assess determinants of bilingual lexical acquisition. It examines the time course of acquisition, the emergence of structured lexical representations in L1 and L2, and the effect of learning history on learning plasticity. The model attempts to account for important processes such as competition, the extent to which the two lexicons compete for resources in the lexical space over time; entrenchment, the extent to which lexical structures are consolidated in L1 affects the learning of L2, and vice versa; and plasticity, the extent to which structural consolidation of L1 impacts the learning of L2. In the final talk Michael Thomas will review the recent application of connectionist models of language processing to bilingualism. Connectionist models have frequently appealed to two different architectures, localist interactive activation models and distributed processing models. These architectures have been used to explore different phenomena within bilingual language processing, including localist models of visual and auditory word recognition and distributed models of lexical and syntactic acquisition. Recent approaches employing self-organization attempt to bridge the two types of model. The range of existing models of monolingual language processing suggest clear avenues for future bilingual research to pursue, in particular focusing on dynamic aspects of (1) bilingual acquisition, (2) gradual changes in language dominance, (3) real-time switching between languages, (4) bilingual aphasia and recovery, and (5) language decay. Finally, Thomas will conclude with his recent modeling work exploring critical periods and their implications for second language acquisition. Ping Li will give an introduction and overview of the symposium at the beginning, and Yasuhiro Shirai will provide an integrative discussion at the end. Introduction Computational modeling and bilingualism have had until recently only limited interactions (see reviews in French & Jacquet, 2004; Hernandez, Li, & MacWhinney, 2005; Li & Farkas, 2002; Thomas & van Heuven, 2005). Bilingualism has been the norm rather than the exception in our globalized world, but the acquisition of two languages entails significant complexity that challenges empirical methodology. Computational modeling, because of its flexibility in parameter variation and hypothesis testing, is ideally suited for identifying mechanisms underlying bilingual language acquisition and representation. In this symposium, we propose to integrate current computational studies of bilingualism and second language acquisition. Summary of Presentations Robert French will begin by reviewing the state of the art in the study of bilingual lexical memory, pointing out crucial issues in the field. He will then present the BSRN, a bilingual simple recurrent network model. The model learns both English and French NVN strings, intermixed at the sentence level for the two languages. The simulations show that BSRN can develop distinct representations not only for individual lexical categories in each language (as in Elman, 1990), but also for the two languages in general. Thus, the model can display distinct behaviors for the bilingual’s two lexicons without invoking separate mechanisms for each language, providing evidence to the idea of “single mechanism, variable representations” from bilingualism. In the second talk Curt Burgess will present the bilingual HAL model. Using language co-occurrences to model language or memory raises a number of controversial issues on the nature of lexical and semantic representations. In addition, using lexical co-occurrence to model bilingualism introduces crucial theoretical considerations since it is incumbent on the model to account for the transformation of the different lexical codes of multiple languages into similar semantic representations. On the surface this may seem straightforward. However, since co-occurrence models function (at some point in the encoding process) by counting the number of co-occurrences between specific lexical items, one has to provide an account of how the co- occurrence vectors for L1 can merge or co-exist with the vectors for L2. Burgess will discuss this memory consolidation process, a step that is not required in high- dimensional models that encodes only one language. In the third talk Ping Li will present a self-organizing neural network model that simulates developmental stages of the bilingual lexicon. The model is based on DevLex, a References French, R., & Jacquet, M. (2004). Understanding bilingual memory: models and data. Trends in Cognitive Sciences, Hernandez, A., Li, P., & MacWhinney, B. (2005). The emergence of competing modules in bilingualism. Trends in Cognitive Sciences, 9, 220-225. Li, P., & Farkas, I. (2002). A self-organizing connectionist model of bilingual processing. In R. Heredia & J. Altarriba (eds.), Bilingual sentence processing. Elsevier. Thomas, M., & van Heuven, W. (2005). Computational models of bilingual comprehension. In J. F. Kroll & A. de Groot (eds.) Handbook of Bilingualism: Psycholinguistic Approaches. Oxford University Press.
The present dissertation addresses a set of questions about processes involved in lexical access and literacy and psycholinguistic factors (such as word learning age, word frequency etc.) that affect them in monolingual and bilingual speakers. Four experiments examined these issues for nouns in native English speakers and bilingual Hindī- English speakers with a developmental perspective. Experiments 1 and 2 were conducted with English monolinguals in San Diego. In experiment 1, age of acquisition norms were collected from college-age adults. In experiment 2, online picture naming data was collected from four age groups of English monolinguals (5-7, 8-10, 11-13 and college-age adults). Experiments 3 and 4 were conducted on Hindī-English bilinguals in India. In experiment 3, age of acquisition and word frequency norms were collected from college-age adults. In experiment 4, online picture naming and word reading data were collected from three age groups of Hindī-English bilinguals (8-10, 11-13 and college-age adults). Comparisons of performance on two lexical access tasks (on-line picture naming and word reading) in monolingual English speakers and bilingual Hindī-English speakers, were conducted. Results and discussion are aimed at addressing issues of language processing, lexical access and development. Overall, results indicate that there is developmental improvement on the lexical access tasks. In addition, the predictor- outcome relationships are generally similar for both monolinguals and bilinguals. Age of acquisition is the most consistent predictor of both picture naming and word reading behavior, in both monolingual and bilingual speaker. There are differential effects of frequency in the languages of the bilingual in the word reading task, with orthographic differences interacting with frequency effects. However, there are interesting differences that arise between the monolinguals and within the bilinguals, because of language dominance and proficiency. Results and discussion focus on quantitative analyses, examining lexical access processes in monolinguals and bilinguals, and examining the relationship between the psycholinguistic variables (such as age of acquisition, frequency, and syllable length) and performance on the language production tasks. Future directions focus on highlighting some limitations of this research, in addition to discussing the need for more in depth qualitative analyses, and extending these paradigms to clinical populations
In this article, we describe the Naturalistic University of Alberta Nonlinear Correlation Explorer (NUANCE), a computer program for data exploration and analysis. NUANCE is specialized for finding nonlinear relations between any number of predictors and a dependent value to be predicted. It searches the space of possible relations between the predictors and the dependent value by using natural selection to evolve equations that maximize the correlation between their output and the dependent value. In this article, we introduce the program, describe how to use it, and provide illustrative examples. NUANCE is written in Java, which runs on most computer platforms. We have contributed NUANCE to the archival Web site of the Psychonomic Society (www.psychonomic.org/archive), from which it may be freely downloaded.
The Leximancer system is a relatively new method for transforming lexical co-occurrence information from natural language into semantic patterns in an unsupervised manner. It employs two stages of co-occurrence information extraction—semantic andrelational—using a different algorithm for each stage. The algorithms used are statistical, but they employ nonlinear dynamics and machine learning. This article is an attempt to validate the output of Leximancer, using a set of evaluation criteria taken from content analysis that are appropriate for knowledge discovery tasks.
Multilingual dependency parsing is gaining popularity in recent years for several reasons. Dependency structures are more adequate for languages with freer word order than the traditional constituency notion. There is a growing availability of dependency treebanks for new languages. Broad coverage statistical dependency parsers are available and easily portable to new languages. Dependency parsing can provide useful contributions in areas such as information extraction, machine translation and question answering, among others. In addition, syntactic head-dependent pairs are a good interface between the traditional phrase structures and semantic theta roles. In this paper we present the learning curves of a statistical dependency parser for four languages: Arabic, Bulgarian, Italian and Slovene. We discuss issues that mostly concern the employed annotation scheme for each treebank with an emphasis on coordinated structures. Povzetek: Opisano je večjezično odvisnostno skladenjsko razčlenjevanje štirih jezikov. 1
Abstract Attempting to automatically learn to identify verb complements from natural language corpora without the help of sophisticated linguistic resources like grammars, parsers or treebanks leads to a significant amount of noise in the data. In machine learning terms, where learning from examples is performed using class-labelled feature-value vectors, noise leads to an imbalanced set of vectors: assuming that the class label takes two values (in this work complement/non-complement), one class (complements) is heavily underrepresented in the data in comparison to the other. To overcome the drop in accuracy when predicting instances of the rare class due to this disproportion, we balance the learning data by applying one-sided sampling to the training corpus and thus by reducing the number of non-complement instances. This approach has been used in the past in several domains (image processing, medicine, etc) but not in natural language processing. For identifying the examples that are safe to remove, we use the value difference metric, which proves to be more suitable for nominal attributes like the ones this work deals with, unlike the Euclidean distance, which has been used traditionally in one-sided sampling. We experiment with different learning algorithms which have been widely used and their performance is well known to the machine learning community: Bayesian learners, instance-based learners and decision trees. Additionally we present and test a variation of Bayesian belief networks, the COr-BBN (Class-oriented Bayesian belief network). The performance improves up to 22% after balancing the dataset, reaching 73.7% f-measure for the complement class, having made use only a phrase chunker and basic morphological information for preprocessing.
This paper explores several important issues in developing syntactically annotated Korean corpora for higher-level language processing, including semantic-discourse parsing, question-answering, machine translation, information retrieval, etc.In particular, we compare the Penn Korean Treebank (PKT) and the Korean Treebank of the 21st Century Sejong Project (ST) and discuss four critical issues in syntactic annotation.We argue for the use of more sophisticated morphosyntactic information, and based on our comparative study, we propose revisions in the syntactic annotation schemes of the existing Korean Treebanks in order to improve the quality of annotated corpora and their usability both for conducting theoretical research and for developing computational tools.The results of our comparative study reveal four significant issues in syntactic annotations: the syntactic analysis of verbal complexes, the hierarchical structure of noun phrases, the representation of traces, and the marking of zero elements.These factors may trigger erroneous syntactic representations for certain linguistic phenomena, and they may increase difficulties in data search and lessen reliability in computational processing.Thus, evaluating and improving the syntactic annotation of Treebanks is an important task for aspects of both theoretical and computational linguistics.
Studies on attribution in the moral domain often involve the use of specific behavior examples. To make valid comparisons across trait dimensions (such as honesty and friendliness), it is important to equate the intensities of the specific behaviors used. Pretesting specific behaviors can be a costly effort, but it is often necessary for research in social psychology. Our study provides a rich source of such pretested behaviors. Positive and negative examples of behaviors in the categories of honesty, loyalty, friendliness, charitableness, and cooperativeness were solicited from participants and then rated on the relevant trait dimension by an independent group. The result is data representing rankings, raw scores, andz-scores in an index of 500 behaviors across 10 trait categories that can be used by researchers to study moral and immoral behaviors. The full index of behaviors is available at www .psychonomic.org/archive/.
Previous articleNext article FreeCurrent ApplicationsLinguistic AnthropologyV.ChandV.Chand Search for more articles by this author PDFPDF PLUSFull Text Add to favoritesDownload CitationTrack CitationsPermissionsReprints Share onFacebookTwitterLinked InRedditEmailQR Code SectionsMoreImmigration Practices in Belgium: African AsylumSeeker DiscourseJan Blommaert, a Belgian anthropologist currently at the Institute of Education, University of London, is known in Belgium as a public campaigner on immigration issues. He became involved in African asylumseekers rights in 1998, when the death of Semira Adamu during her forced repatriation provoked a public outcry over the implementation of Belgian immigration policies, in which more than 95% of applicants for asylum are rejected. Adamu had fled Nigeria in her teens to avoid entering into a polygamous marriage with a 65yearold man. Shackled at the ankles and vigorously resisting, she was suffocated by Belgian police attempting to restrain her.Poster for a 2003 commemoration of Adamu's death.View Large ImageDownload PowerPointAdamus death was a catalyst for the formation of new action groups, and existing organizations also became involved: churches opened their doors to asylum seekers, and NGOs such as OXFAM and the League of Human Rights campaigned for asylum seekers rights. Early on, Blommaert contributed to this organic collaborative effort: I gave tons of public lectures for any audience willing to listen, wrote opeds in major newspapers, campaigned with MPs close to the government, participated in public debates on these matters, and wrote expert articles for a wider audience. Additionally, he saw that his understanding of the legal and sociolinguistic issues was directly applicable to the problem. He was motivated, he explains, by awareness of the real stakes and real people involved.African asylum seekers face circumstances not shared with those from Europe and the Gulf because they come from wartorn areas with unclear state boundaries, may be illiterate and lack documentation, have long migration paths to Belgium, and do not share a language with immigration authorities. In particular, Blommaert points out, their choices of language, background texts, and genres for storytelling affect their chances of acquiring refugee status. African asylumseeker language is stereotypically filled with language mixing, language impurities, varying degrees of literacy, and different styles of storytelling and discourse. When asylum seekers present their stories, they are unaware of how the officials are judging them on their linguistic habits. The officials note these details and find them inconsistent with Belgian expectations; they judge the asylum seekers in terms of these nave choices and reject them. In many such cases, rejection is a matter of life or death for the refugees, given the risks associated with deportation and with repatriation into the home country that originally motivated them to seek asylum.Blommaert has mobilized an informal network of Belgian academics and institutions that has produced many academically informed public statements. Additionally, by documenting and analyzing oral immigration interviews with an eye to understanding the disconnect between asylum seekers presentations and immigration officials expectations, he has improved practice in the Immigration Department. He has used his analysis to train members of the department, helping to adjust views of what can be gathered in an immigration interview, to improve interview techniques, and to promote an awareness of variability in sociocultural linguistic norms and presentation styles.He continues to work as an official expert for Belgian legal, government, and security authorities, translating and analyzing documentation and corroboratory written texts provided for immigration procedures and advising on specific issues. He has been able to influence outcomes for some individual asylum seekers. His public campaigning has raised awareness of the issues surrounding African asylum seekers, and his publications have provoked work within academia that may lead to further interventions by academics with regard to the interview process.Blommaert hopes that public campaigning by NGOs and academic scholarship will promote continued dialogue with immigration officials, eventually producing policies and procedures that take into account the politics of migration and displacement and the way in which asylum seekers frame their life stories. Previous articleNext article DetailsFiguresReferencesCited by Current Anthropology Volume 47, Number 3June 2006 Sponsored by the Wenner-Gren Foundation for Anthropological Research Article DOIhttps://doi.org/10.1086/504161 Views: 287Total views on this site Citations: 1Citations are reported from Crossref PDF download Crossref reports the following articles citing this article:Kevin D. O'Gorman, Cailein Gillespie The mythological power of hospitality leaders?, International Journal of Contemporary Hospitality Management 22, no.55 (Jul 2010): 659–680.https://doi.org/10.1108/09596111011053792
The databases record instances of deponency, which is the term we have adopted to describe mismatches between morphology and morphosyntax. The prototypical example are the deponent verbs of Latin, which involve a mismatch between passive form and active meaning. That is, a normal Latin verb had active forms such as amō 'I love' and amāvī 'I have loved', which contrasted with the passive forms amor 'I am loved' and amātus sum 'I have been loved' (in this case, with a masculine subject). A deponent verb, on the other hand, looks like the passive but functions like the active, as in mīror 'I admire', mīrātus sum 'I have admired'. In the the databases we construe deponency in an extended fashion, covering any mismatch between the apparent morphosyntactic value of a morphological form and its actual value in a given syntactic context. Two databases are housed on this site, accessible through the links above. The cross-linguistic database looks at the presence of morphological mismatches in a controlled sample of genetically and geographically diverse languages (based on the 100-language sample from the World Atlas of Language Structures). The typological database records the logical space of deponency: what features may be affected, and what are the characteristics of the resulting paradigm? Every logical combination of parameters is represented by one exemplar (or where none has been found, this is noted too). The typological database is supplemented by a set of formal analyses of examples which hold particular interest for morphological theory.
Automatic analysis of syntax is one of the core problems in natural language processing. Despite significant advances in syntactic parsing of written text, the application of these techniques to spontaneous spoken language has received more limited attention. The recent explosive growth of online, accessible corpora of spoken language interactions opens up new opportunities for the development of high accuracy parsing approaches to the analysis of spoken language. The availability of high accuracy parsers will in turn provide a platform for development of a wide range of new applications, as well as for advanced research on the nature of conversational interactions. One concrete field of investigation that is ripe for the application of such parsing tools is the study of child language acquisition. In this thesis, we describe an approach for analyzing the syntactic structure of spontaneous conversational language in parent-child interactions. Specific emphasis is placed on the challenge of accurately annotating the English corpora in the CHILDES database with grammatical relations (such as subject, objects and adjuncts) that are of particular interest and utility, to researchers in child language acquisition. This work involves rule-based and corpus-based natural language processing techniques, as well as methodology for combining results from different parsing approaches. We present novel strategies for integrating the results of different parsers into a system with improved accuracy. One practical application of this research is the automation of language competence measures used by clinicians and researchers of child language development. We present an implementation of an automatic version of one such measurement scheme. This provides not only a useful tool for the child language research community, but also a task-based evaluation framework for grammatical relation identification. Through experiments using data from the Penn Treebank, we show that several of the techniques and ideas presented in this thesis are applicable not just to analysis of parent-child dialogs, but to parsing in general.
The correct attachment of prepositional phrases (PPs) is a central disambiguation problem in parsing natural languages. This paper compares the baseline situation in English, German and Swedish based on manual PP attachments in various treebanks for these languages. We argue that cross-language comparisons of the disambiguation results in previous research is impossible because of the different selection procedures when building the training and test sets. We perform uniform treebank queries and show that English has the highest noun attachment rate followed by Swedish and German. We also show that the high rate in English is dominated by the preposition of. From our study we derive a list of criteria for profiling data sets for PP attachment experiments.
Previously, we introduced a new computational tool for nonlinear curve fitting and data set exploration: the Naturalistic University of Alberta Nonlinear Correlation Explorer (NUANCE) (Hollis & Westbury, 2006). We demonstrated that NUANCE was capable of providing useful descriptions of data for two toy problems. Since then, we have extended the functionality of NUANCE in a new release (NUANCE 3.0) and fruitfully applied the tool to real psychological problems. Here, we discuss the results of two studies carried out with the aid of NUANCE 3.0. We demonstrate that NUANCE can be a useful tool to aid research in psychology in at least two ways: It can be harnessed to simplify complex models of human behavior, and it is capable of highlighting useful knowledge that might be overlooked by more traditional analytical and factorial approaches. NUANCE 3.0 can be downloaded from the Psychonomic Society Archive of Norms, Stimuli, and Data at www.psychonomic.org/archive.
We exploit the resources in the Arabic Treebank (ATB) for the novel task of automatically creating lexical semantic verb classes for Modern Standard Arabic (MSA). Verbs are clustered into groups that share semantic elements of meaning as they exhibit similar syntactic behavior. The results of the clustering experiments are compared with a gold standard set of classes, which is approximated by using the noisy English translations provided in the ATB to create Levin-like classes for MSA. The quality of the clusters is found to be sensitive to the inclusion of information about lexical heads of the constituents in the syntactic frames, as well as parameters of the clustering algorithm. The best set of parameters yields an Fβ=1 score of 0.501, compared to a random baseline with an Fβ=1 score of 0.37.
Web searchers reformulate their queries, as they adapt to search engine behavior, learn more about a topic, or simply correct typing errors. Automatic query rewriting can help user web search, by augmenting a user’s query, or replacing the query with one likely to retrieve better results. One example of query-rewriting is spell-correction. We may also be interested in changing words to synonyms or other related terms. For Japanese, the opportunities for improving results are greater than for languages with a single character set, since documents may be written in multiple character sets, and a user may express the same meaning using different character sets. We give a description of the characteristics of Japanese search query logs and manual query reformulations carried out by Japanese web searchers. We use characteristics of Japanese query reformulations to extend previous work on automatic query rewriting in English, taking into account the Japanese writing system. We introduce several new features for building models resulting from this difference and discuss their impact on automatic query rewriting. We also examine enhancements in the form of rules which block conversion between some character sets, to address Japanese homophones. The precision/recall curves show significant improvement with the new feature set and blocking rules, and are often better than the English counterpart.
We have introduced a new methodology that maps designs to human perceptions. Perceptions are adjectives/adverbs, phrases or sentences expressed in natural language. We used the lexical database WordNet to compute semantical relationships (or distances) between these perceptions. We partitioned the set of perceptions into k clusters that represent the classes for a further classification task. We have developed a new classifier called “structural hidden Markov model” (SHMM) that combines probability and distances in a seamless way. SHMM enables to learn and predict user perceptions given object designs. We have applied this approach to Kansei engineering in order to map car external contours (shapes) to customer perceptions. The accuracy obtained using the SHMM is 90%. This model has outperformed the neural network and the k-nearest-neighbor classifiers.
AIM: (i) To determine the ability of general practitioners (GPs) and paediatricians to correctly identify children as overweight or obese by visual cues alone; (ii) to describe the current management practices of overweight and obese children by these practitioners; and (iii) to compare these with National Health and Medical Research Council (NHMRC) Clinical Practice Guidelines. METHODS: Forty-four GPs and 29 paediatricians participated in the study. Respondents completed a questionnaire based on a series of body images, rating the size of the child as acceptable weight, overweight or obese and indicating the likelihood of carrying out a series of management options. RESULTS: There was considerable variation in ability to rate images correctly with the total number of correct responses being 72% and 68%, respectively, for GPs and paediatricians. There were statistically significant differences in management between GPs and paediatricians in terms of conducting appropriate anthropometry and screening for co-morbidities, with paediatricians performing closer to the NHMRC Clinical Practice Guidelines. CONCLUSION: GPs and paediatricians have the opportunity to screen children for overweight and obesity during their everyday practice. Accurate determination of weight status cannot be performed by visualisation alone and all children should have height and weight measured and correctly interpreted. Some areas of current GP and paediatrician management of overweight and obese children fall short of the NHMRC clinical guidelines and areas for improvement are highlighted in this paper.
Purpose This paper examined employee perceptions of the rewards associated with their participation in a six sigma program. Six sigma is an approach to organizational change that incorporates elements of total quality management, business process reengineering, and employee involvement. Design/methodology/approach A survey was completed by 215 employees (34 percent response rate). Respondents rated the extent to which they felt their participation in six sigma was “instrumental” for a range of outcomes, as well as valence (desirability) of each outcome (based on the VIE concept of instrumentality). The outcomes were classified into four categories: extrinsic, intrinsic, social, and organizational. Findings Valence ratings revealed that all 12 outcomes were perceived as desirable. Instrumentality ratings showed that extrinsic outcomes were rated significantly lower than intrinsic, social, and organizational outcomes. Additional analyses revealed significant differences on all four outcome categories between participants and non‐participants in the six sigma program. Practical implications The positive valence and instrumentality ratings for participants indicate they believe their participation will lead to valued outcomes for themselves and their organizations. However, employees who choose not to get involved in six sigma do not perceive that their participation would have led to desired outcomes. The results also show that while participants value extrinsic rewards, they do not see six sigma as instrumental in their receipt. These perceptions have important implications for attracting and retaining program participants. Originality/value While much has been written about the use of reward systems in supporting a successful six sigma effort, this study empirically examines how employees actually perceive the rewards associated with their participation. It also identifies which types of rewards are most instrumental for participants and non‐participants.
Group identifications and intergroup relations among Turkish Dutch respondents What determines group identification processes among ethnic minority groups and how are these processes related to in-group and out-group evaluations? This article focuses on Turkish and Dutch identification among Turkish Dutch respondents and their feelings towards different ethnic and religious groups. The results show that Turkish identification is strong and Dutch identification rather weak, and that both group identifications are not strongly associated. Perceived socio-structural characteristics of intergroup relations (stability, legitimacy, permeability, and discrimination) affected both Turkish and Dutch identification. Group identification was positively related to (ethnic and religious) in-group evaluation, but there were few relationships with out-group evaluations. The affective ratings of Moroccans, Antilleans, Jews and non-believers were quite negative.
In recent years, research in parsing has extended in several new directions. One of these directions is concerned with parsing languages other than English. Treebanks have become available for many European languages, but also for Arabic, Chinese, or Japanese. However, it was shown that parsing results on these treebanks depend on the types of treebank annotations used [, ]. Another direction in parsing research is the development of dependency parsers. Dependency parsing profits from the non-hierarchical nature of dependency relations, thus lexical information can be included in the parsing process in a much more natural way. Especially machine learning based approaches are very successful (cf. e.g. [12, 13]). The results achieved by these dependency parsers are very competitive although comparisons are difficult because of the differences in annotation. For English, the Penn Treebank [11] has been converted to dependencies. For this version, Nivre et al. [14] report an accuracy rate of 86.3%, as compared to an F-score of 2.1 for Charniak’s parser [1]. The Penn Chinese Treebank [1 ] is also available in a constituent and a dependency representations. The best results reported for parsing experiments with this treebank give an F-score of 81.8 for the constituent version [2] and. % accuracy for the dependency version [14]. The general trend in comparisons between constituent and dependency parsers is that the dependency parser performs slightly worse than the constituent parser. The only exception occurs for German, where F-scores for constituent plus grammatical function parses range between 51.4 and 5.3, depending on the treebank, NEGRA [1 ] or TuBa-D/Z [1 ]. The dependency parser based on a converted version of Tuba-D/Z, in contrast, reached an accuracy of 3.4% [14], i.e. 12 percent points better than the best constituent analysis including grammatical functions.
This paper describes an ongoing effort to parse the Hebrew Bible. The parser consults the bracketing information extracted from the cantillation marks of the Masoetic text. We first constructed a cantillation treebank which encodes the prosodic structures of the text. It was found that many of the prosodic boundaries in the cantillation trees correspond, directly or indirectly, to the phrase boundaries of the syntactic trees we are trying to build. All the useful boundary information was then extracted to help the parser make syntactic decisions, either serving as hard constraints in rule application or used probabilistically in tree ranking. This has greatly improved the accuracy and efficiency of the parser and reduced the amount of manual work in building a Hebrew treebank.
We present the implementation of a system which extracts not only lexicalized grammars but also feature-based lexicalized grammars from Korean Sejong Treebank. We report on some practical experiments where we extract TAG grammars and tree schemata. Above all, full-scale syntactic tags and well-formed morphological analysis in Sejong Treebank allow us to extract syntactic features. In addition, we modify Treebank for extracting lexicalized grammars and convert lexicalized grammars into tree schemata to resolve limited lexical coverage problem of extracted lexicalized grammars.
The main purpose of this study was to compare the validity of two methodologies in measuring the change of a destination image resulting from a small-scale community festival. One method compared same respondents’ rating of image attributes during and after their event participation, while the other method asked respondents to directly report their image change, and chose which factors contributed to the change. Interestingly, the two separate measures employed within this study provided contrasting information on how the festival impacted the image of the host community. The more subjective method indicated that the vast majority (77.9%) of respondents felt that the festival improved their image of the hosting community, while the more objective measure identified decreased image ratings between on-site and mail-back measurement periods. Potential reasons of the conflicting results are proposed and implications of the findings are discussed.
Grammars can be induced from treebanks for a potential variety (e.g., parsing), but this task faces two major problems. One is the existence of annotation er-rors in the treebank which can affect the final system performance (cf. [4]). The other is that the sheer number of rules is overwhelming for most applications. As
This document describes the information used for summarization-inspired temporal-relation extraction [Dorr and Gaasterland, 2007]. We present a set of tense/aspect extraction templates that are applied to a Penn Treebank-style analysis of the input sentence. We also present an analysis of tense-pair combinations for different temporal connectives based on a corpus analysis of complex tense structures in Treebank-3. Finally, we include analysis charts and temporal relation tables for all combinations of intervals/points for each legal BTS combinations.
Abstract. The present paper focuses on representation of morphological meanings on the underlying syntactic level. The concept of semantic counterparts of morphological meanings, the so-called grammatemes, was introduced in Functional Generative Description in the 1960’s. We suggest an elaborated system of these grammatemes, which have become a part of the tectogrammatical level of the Prague Dependency Treebank.
Modern Chinese is a highly developed language that has norms phonetically, lexically, and grammatically. The norms are neither invariable nor variable unconditionally outside language context or without the restrictions of communication or styles. Standardization of modern Chinese should be viewed in a multi - perspective dimension. The chimerical views of standardization that deviates from language context, ignores stylistic characteristics, overlooking pragmatic functions, and blindly pursues standardization for its own sake should be opposed to.
We present a treebank conversion method by which we construct an RMRS bank for HPSG parser evaluation from the TIGER Dependency Bank. Our method effectively performs automatic RMRS semantics construction from functional dependencies, following the semantic algebra of (Copestake et al., 2001). We present the semantics construction mechanism, and focus on some special phenomena. Automatic conversion is followed by manual validation. First evaluation results yield high precision of the automatic semantics construction rules. 1
Incremental parsing gains its importance in natural language processing and psycholinguistics because of its cognitive plausibility. Modeling the associated cognitive data structures, and their dynamics, can lead to a better understanding of the human parser. In earlier work, we have introduced a recursive neural network (RNN) capable of performing syntactic ambiguity resolution in incremental parsing. In this paper, we report a systematic analysis of the behavior of the network that allows us to gain important insights about the kind of information that is exploited to resolve different forms of ambiguity. In attachment ambiguities, in which a new phrase can be attached at more than one point in the syntactic left context, we found that learning from examples allows us to predict the location of the attachment point with high accuracy, while the discrimination amongst alternative syntactic structures with the same attachment point is slightly better than making a decision purely based on frequencies. We also introduce several new ideas to enhance the architectural design, obtaining significant improvements of prediction accuracy, up to 25% error reduction on the same dataset used in previous work. Finally, we report large scale experiments on the entire Wall Street Journal section of the Penn Treebank. The best prediction accuracy of the model on this large dataset is 87.6%, a relative error reduction larger than 50% compared to previous results.
Dieser Beitrag beschäftigt sich mit sprachlichen Mischformen im frankophonen Kanada. Im Mittelpunkt steht die hybride Varietät des Chiac, eine in Moncton/Acadie gesprochene urbane Mischvarietät, die aus dem Sprachkontakt von Englisch und Französisch entstanden ist und auf einer französischen Grammatik mit englischen Lexikelementen basiert. Gegenstand der Betrachtung ist die in den späten 1960er Jahren entstandene littérature acadienne und ihr Umgang mit sprachlicher Variation im Spannungsfeld von gesellschaftlicher Norm und Formen ihrer Transgression. Am Beispiel der zeitgenössischen Autorin France Daigle und ihrem Umgang mit der Varietät des Chiac soll verdeutlicht werden, in welcher Weise hybride Formen von Sprache, die existente Sprach- und Kulturgrenzen in Frage stellen, dazu beitragen, individuelle und soziale Widersprüche zu erfassen und zu bearbeiten.