Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Purpose This paper examined employee perceptions of the rewards associated with their participation in a six sigma program. Six sigma is an approach to organizational change that incorporates elements of total quality management, business process reengineering, and employee involvement. Design/methodology/approach A survey was completed by 215 employees (34 percent response rate). Respondents rated the extent to which they felt their participation in six sigma was “instrumental” for a range of outcomes, as well as valence (desirability) of each outcome (based on the VIE concept of instrumentality). The outcomes were classified into four categories: extrinsic, intrinsic, social, and organizational. Findings Valence ratings revealed that all 12 outcomes were perceived as desirable. Instrumentality ratings showed that extrinsic outcomes were rated significantly lower than intrinsic, social, and organizational outcomes. Additional analyses revealed significant differences on all four outcome categories between participants and non‐participants in the six sigma program. Practical implications The positive valence and instrumentality ratings for participants indicate they believe their participation will lead to valued outcomes for themselves and their organizations. However, employees who choose not to get involved in six sigma do not perceive that their participation would have led to desired outcomes. The results also show that while participants value extrinsic rewards, they do not see six sigma as instrumental in their receipt. These perceptions have important implications for attracting and retaining program participants. Originality/value While much has been written about the use of reward systems in supporting a successful six sigma effort, this study empirically examines how employees actually perceive the rewards associated with their participation. It also identifies which types of rewards are most instrumental for participants and non‐participants.
We study unsupervised methods for learning refinements of the nonterminals in a treebank. Following Matsuzaki et al. (2005) and Prescher (2005), we may for example split NP without supervision into NP[0] and NP[1], which behave differently. We first propose to learn a PCFG that adds such features to nonterminals in such a way that they respect patterns of linguistic feature passing: each node's nonterminal features are either identical to, or independent of, those of its parent. This linguistic constraint reduces runtime and the number of parameters to be learned. However, it did not yield improvements when training on the Penn Treebank. An orthogonal strategy was more successful: to improve the performance of the EM learner by treebank preprocessing and by annealing methods that split nonterminals selectively. Using these methods, we can maintain high parsing accuracy while dramatically reducing the model size.
Abstract Facial masculinity may be used as a cue in female mate choice, as it reflects the success of the male genotype in its developmental environment. Women may maximize reproductive success by using a conditional strategy favoring highly masculine facial features for short‐term relationships and feminized facial features in men for long‐term relationships. Three studies examine reactions to masculinized and feminized male facial composites. Properties of the original composite image affect ratings of critical attributes and the magnitude of the differences in ratings between versions undergoing identical processes of geometric manipulation (Study 1). Both men and women attribute personality, behavior, and mating strategies consistent with predictions derived from the good genes and mating trade‐off hypotheses (Study 2). Participants accurately grouped behavioral tendencies related to high mating effort/risky strategies and high parenting effort/risk adverse strategies and associated mating effort more so with masculinized faces and parenting effort more so with feminized faces (Study 3). These results indicate that male facial masculinity serves as a visual cue for inferring personality and reproductive strategy.
OBJECTIVE: We examined whether affect ratings predicted regional cerebral responses to high and low-calorie foods. METHOD: Thirteen normal-weight adult women viewed photographs of high and low-calorie foods while undergoing functional magnetic resonance imaging (fMRI). Regression analysis was used to predict regional activation from positive and negative affect scores. RESULTS: Positive and negative affect had different effects on several important appetite-related regions depending on the calorie content of the food images. When viewing high-calorie foods, positive affect was associated with increased activity in satiety-related regions of the lateral orbitofrontal cortex, but when viewing low-calorie foods, positive affect was associated with increased activity in hunger-related regions including the medial orbitofrontal and insular cortex. The opposite pattern of activity was observed for negative affect. CONCLUSION: These findings suggest a neurobiologic substrate that may be involved in the commonly reported increase in cravings for calorie-dense foods during heightened negative emotions.
In this paper we describe the structure and development of the Brandeis Semantic Ontology (BSO), a large generative lexicon ontology and lexical database. The BSO has been designed to allow for more widespread access to Generative Lexicon-based lexical resources and help researchers in a variety of computational tasks. The specification of the type system used in the BSO largely follows that proposed by the SIMPLE specification (Busa et al., 2001), which was adopted by the EU-sponsored SIMPLE project (Lenci et al., 2000). 1.
With increasing numbers of Web users, there is a necessity to improve their Web site navigation experience over the Internet and a range of Web applications have emerged recently for this purpose. Many researchers have stressed the importance of identifying semantic relatedness of Web pages in such Web applications as Web site navigation, automatic tour generation and adaptive Web applications. One approach to identifying semantic relatedness between documents is to use lexical databases and lexical chains. For example, an approach using lexical chains has been proposed by Green for identifying paragraph similarity in a document [1]. However, due to the unacceptable length of time needed for lexical chaining and the difficulty of global representation of documents, Green used synset weight vectors to compare semantic relatedness between two documents. But his approach to identifying paragraph similarities can be extended to identify semantic similarities between documents. In this study, an approach to identifying semantic similarity between Web pages incorporating weighted lexical chains (SRWLC) and document properties based on reiteration, density, length and semantic distance is proposed. The two approaches (the proposed approach and Green’s approach - SR Green ) were empirically compared by determining the semantic relatedness of Web pages using human subjects. The research hypothesis of this research is that the proposed approach identifies significantly more semantically related pages with a higher precision than the approach that has been proposed by Green. The null hypothesis is that there is no significant different in identification of semantically related pages between the two approaches. In this context precision is defined as the proportion of retrieved pages that are relevant. Web pages belonging to the Department of Computer Science, Keele University are used for the empirical evaluation of the two methods. The semantic relatedness of all pages was identified using both approaches and a Web-based page categorisation exercise using human subjects was carried out for this empirical evaluation. The evaluation is Web based, and therefore can be carried out on the subjects’ preferred Web browser at his or her preferred time & place. Therefore the distortion effects are minimised and the results of the evaluation are realistic and also can be reliably generalised to some extent. An invitation e-mail was sent to twenty subjects during the first week of October 2004 giving a link and guidelines for them to start the experiment. Once the link on the email was clicked, subjects were shown the initial Web page, giving an introduction to the experiment and instructions on how to continue. Twelve out of twenty invited subjects completed the experiment. Two subjects attempted the experiments but couldn’t finish because of network problems; another three subjects couldn’t finish because of time restrictions, and the other three subjects did not respond at all. Therefore responses from only twelve subjects are used for experimental evaluation. The Wilcoxon signed ranks test returns a p value of 0.004, indicating that the null hypothesis can be rejected, and that there is evidence to suggest that the SR WLC approach identifies significantly more semantically-related pages with higher precision than the SR Green approach. The SRWLC approach should be evaluated further using different Web site contents and language styles (e.g. American/British English). It would be interesting to use more subjects from different backgrounds to do the evaluation. This would determine whether the results of the evaluation are influenced by the human subject’s background, such as their status (student, staff, or other) his familiarity with the pages of the test database, gender differences or the level of English knowledge. Most importantly the approach is believed to be valid for more general browsing environments than a computer science Website and a wider study is desirable.
Recent research on causal learning found (a) that causal judgments reflect either the current predictive value of a conditional stimulus (CS) or an integration across the experimental contingencies used in the entire experiment and (b) that postexperimental judgments, rather than the CS's current predictive value, are likely to reflect this integration. In the current study, the authors examined whether verbal valence ratings were subject to similar integration. Assessments of stimulus valence and contingencies responded similarly to variations of reporting requirements, contingency reversal, and extinction, reflecting either current or integrated values. However, affective learning required more trials to reflect a contingency change than did contingency judgments. The integration of valence assessments across training and the fact that affective learning is slow to reflect contingency changes can provide an alternative interpretation for researchers' previous failures to find an effect of extinction training on verbal reports of CS valence.
In this paper, we present a semiautomatic approach for annotating semantic information in biomedical texts. The information is used to construct a biomedical proposition bank called BioProp. Like PropBank in the newswire domain, BioProp contains annotations of predicate argument structures and semantic roles in a treebank schema. To construct BioProp, a semantic role labeling (SRL) system trained on PropBank is used to annotate BioProp. Incorrect tagging results are then corrected by human annotators. To suit the needs in the biomedical domain, we modify the Prop-Bank annotation guidelines and characterize semantic roles as components of biological events. The method can substantially reduce annotation efforts, and we introduce a measure of an upper bound for the saving of annotation efforts. Thus far, the method has been applied experimentally to a 4,389-sentence treebank corpus for the construction of Bio-Prop. Inter-annotator agreement measured by kappa statistic reaches.95 for combined decision of role identification and classification when all argument labels are considered. In addition, we show that, when trained on BioProp, our biomedical SRL system called BIOSMILE achieves an F-score of 87%.
This paper describes a featurized functional dependency corpus automatically derived from the Penn Treebank. Each word in the corpus is associated with over three dozen features describing the functional syntactic structure of a sentence as well as some shallow morphology. The corpus was created for use in probabilistic surface generation, but could also be useful as a resource for the study of English and the development of other NLP applications. 1.
We evaluate the accuracy of an unlexicalized statistical parser, trained on 4K treebanked sentences from balanced data and tested on the PARC DepBank. We demonstrate that a parser which is competitive in accuracy (without sacrificing processing speed) can be quickly tuned without reliance on large in-domain manually-constructed treebanks. This makes it more practical to use statistical parsers in applications that need access to aspects of predicate-argument structure. The comparison of systems using DepBank is not straightforward, so we extend and validate DepBank and highlight a number of representation and scoring issues for relational evaluation schemes.
The Czech Academic Corpus was created during the 1970s and 1980s at the Czech Lan- guage Institute under the supervision of Marie Těsitelova. The main motivation to build it (a total of 540 thousand word tokens) was to obtain the quantitative characteristics of contemporary Czech. The corpus is structurally annotated on two levels - the morphological level and the syntactical-ana- lytical level. The original stochastic experiments in morphological tagging of Czech were performed using the corpus at the beginning of the 1990s. Given this, the corpus-based processing of Czech was launched. At the end of 1990s, work on the Prague Dependency Treebank had started (independently from the corpus) and its first edition was published in 2001. In considering future released versions of the treebank, we have decided to convert the corpus into the treebank-like format. This article focuses on the twenty-year history of the Czech Academic Corpus. Special attention is devoted to thus far un- published facts about the corpus annotation. The conversion steps resulting in the first version of the Czech Academic Corpus are described in detail.
Natural language processing researchers currently have access to a wealth of information about words and word senses. This presents problems as well as resources, as it is often difficult to search through and coordinate lexical information across various data sources. We have approached this problem by creating a shared environment for various lexical resources. This browser, BULB (Brandeis Unified Lexical Browser) and its accompanying front-end provides the NLP researcher with a coordinated display from many of the available lexical resources, focusing, in particular, on a newly developed lexical database, the Brandeis Semantic Ontology (BSO). BULB is a module-based browser focusing on the interaction and display of modules from existing NLP tools. We discuss the BSO, PropBank, FrameNet, WordNet, and CQP, as well as other modules which will extend the system. We then outline future extensions to this work and present a release schedule for BULB. 1.
Introduction: Stimulation with affective stimuli elicits event-related potentials (ERP), the area under the curve (AUC) correlating with „arousal“-values of these stimuli. The International Affective Picture System (IAPS) contains 716 pictures with different „arousal“-values (1 to 9), most of them identical (i.e. sports, food), some different (erotic pictures) for men and women. The aim of this study was to examine sex differences in the generation of ERP, elicited by pictures with large differences in the arousal-ratings in both sexes.
The PropBank primarily adds semantic role labels to the syntactic constituents in the parsed trees of the Treebank. The goal is for automatic semantic role labeling to be able to use the domain of locality of a predicate in order to find its arguments. In principle, this is exactly what is wanted, but in practice the PropBank annotators often make choices that do not actually conform to the Treebank parses. As a result, the syntactic features extracted by automatic semantic role labeling systems are often inconsistent and contradictory. This paper discusses in detail the types of mismatches between the syntactic bracketing and the semantic role labeling that can be found, and our plans for reconciling them.
We investigate generalizations of the all-subtrees "DOP" approach to unsupervised parsing. Unsupervised DOP models assign all possible binary trees to a set of sentences and next use (a large random subset of) all subtrees from these binary trees to compute the most probable parse trees. We will test both a relative frequency estimator for unsupervised DOP and a maximum likelihood estimator which is known to be statistically consistent. We report state-of-the-art results on English (WSJ), German (NEGRA) and Chinese (CTB) data. To the best of our knowledge this is the first paper which tests a maximum likelihood estimator for DOP on the Wall Street Journal, leading to the surprising result that an unsupervised parsing model beats a widely used supervised model (a treebank PCFG).
Treebank data have been utilized as data sources for a wide range of tasks in computational linguistics, including statistical parsing, anaphora resolution, induction of valence lexica, etc. More recently, researchers have experimented with extracting semantic information from syntactically annotated data. Here, treebank data
Traditionally, orthographic variants have been modelled as different ways of spelling the same word - described at the level of the lexeme. But when inflection is taken into account, this runs into a problem: different citation forms have different inflectional paradigm - and orthographic variation does not merely affect the citation form, but the entire paradigm. The MorDebe database therefore models orthographic variation as a relation between distinct, yet still token-identical lexemes. This paper discusses the advantage of that approach, and the full set of practical problems that arose during the structural treatment of orthographic variation in the MorDebe database.
We introduce Talbanken05, a Swedish treebank based on a syntactically annotated corpus from the 1970s, Talbanken76, converted to modern formats. The treebank is available in three different formats, besides the original one: two versions of phrase structure annotation and one dependency-based annotation, all of which are encoded in XML. In this paper, we describe the conversion process and exemplify the available formats. The treebank is freely available for research and educational purposes. 1.
The paper presents new lexicon of verb valencies for the Czech language named VerbaLex. VerbaLex is based on three valuable language resources for Czech, three independent electronic dictionaries of verb valency frames. The first resource, Czech WordNet valency frames dictionary, was created during the Balkanet project and contains semantic roles and links to the Czech WordNet semantic network. The other resource, VALLEX 1.0, is a lexicon based on the formalism of the Functional Generative Description (FGD) and was developed during the Prague Dependency Treebank (PDT) project. The third source of information for VerbaLex is the syntactic lexicon of verb valencies denoted as BRIEF, which originated at FI MU Brno in 1996. The resulting lexicon, VerbaLex, comprehends all the information found in these resources plus additional relevant information such as verb aspect, verb synonymity, types of use and semantic verb classes based on the VerbNet project.
While syntactically annotated corpora known as treebanks have been available for many years, along with a variety of customized tools for querying these annota-
WordNet, a lexical database for English that is extensively used by computational linguists, has not previously distinguished hyponyms that are classes from hyponyms that are instances. This work describes an attempt to draw this distinction and reports the way in which the results were incorporated in the last version (2.1) of WordNet.
While the processing of verbal and psychophysiological indices of emotional arousal have been investigated extensively in relation to the left and right cerebral hemispheres, it remains poorly understood how both hemispheres normally function together to generate emotional responses to stimuli. Drawing on a unique sample of nine high-functioning subjects with complete agenesis of the corpus callosum (AgCC), we investigated this issue using standardized emotional visual stimuli. Compared to healthy controls, subjects with AgCC showed a larger variance in their cognitive ratings of valence and arousal, and an insensitivity to the emotion category of the stimuli, especially for negatively-valenced stimuli, and especially for their arousal. Despite their impaired cognitive ratings of arousal, some subjects with AgCC showed large skin-conductance responses, and in general skin-conductance responses discriminated emotion categories and correlated with stimulus arousal ratings. We suggest that largely intact right hemisphere mechanisms can support psychophysiological emotional responses, but that the lack of interhemispheric communication between the hemispheres, perhaps together with dysfunction of the anterior cingulate cortex, interferes with normal verbal ratings of arousal, a mechanism in line with some models of alexithymia.
The aim of this research was to study the influence of both the emotional content and the physical characteristics of affective stimuli on the psychophysiological, behavioral and cognitive indexes of the emotional response. We selected 54 pictures from the IAPS, depicting unpleasant, neutral, and pleasant contents, and used two picture sizes as experimental conditions (120 x 90 cm and 52 x 42 cm). Sixty-one subjects were randomly assigned to each experimental condition. We recorded the startle blink reflex, skin conductance response, heart rate, free viewing time, and picture valence and arousal ratings. In line with previous research (e.g., Bradley, Codispoti, Cuthbert, and Lang, 2001), our data showed an effect of the affective content on all the measurements recorded. Importantly, effects of the size of the affective pictures on emotional responses were not found, indicating that the emotional content is more important than the formal properties of the stimuli in evoking the emotional response.
WordNet, a lexical database for English that is extensively used by computational linguists, has not previously distinguished hyponyms that are classes from hyponyms that are instances. This note describes an attempt to draw that distinction and proposes a simple way to incorporate the results into future versions of WordNet.
Supervised text classification is the task of automatically assigning a category label to a previously unlabeled text document. We start with a collection of pre-labeled examples whose assigned categories are used to build a predictive model for each category. In previous research, incorporating semantic features from the WordNet lexical database is one of many approaches that have been tried to improve the predictive accuracy of text classification models. The intuition is that words in the training set alone may not be extensive enough to enable the generation of a universal model for a category, but through Word-Net expansion (i.e., incorporating words defined by various relationships in WordNet), a more accurate model may be possible. In this paper, we report preliminary results obtained from a comprehensive study where WordNet features, part of speech tags, and term weighting schemes are incorporated into two-category text classification models generated by both a Naive Bayes text classifier and an SVM text classifier. We characterize the behaviour of these classifiers on fifteen document collections extracted from the Reuters-21578, USENET, DigiTrad, and 20-Newsgroups text corpora. Experimental results show that incorporating WordNet features, utilizing part of speech tags during WordNet expansion, and term weighting schemes have no positive effect on the accuracy of the Naive Bayes and SVM classifiers.
In the present contribution we claim that corpus annotation serves, among other things, as an invaluable test for linguistic theories standing behind the annotation schemes, and as such represents an irreplaceable resource of linguistic information for the build-up of grammars (Sect. 1.). To support this claim we present four linguistic phenomena for the study and relevant description of which in grammar a deep layer of corpus annotation as introduced in the Prague Dependency Treebank has brought important observations, namely the information structure of the sentence (Sect. 2.), condition of projectivity and word order (Sect. 3.), types of dependency relations (Sect. 4.) and textual coreference (Sect. 5.). 1. Introductory remarks 1.1. Annotation of corpus It has been already commonly accepted in computational and corpus linguistics that grammatical (or lexicalsemantic, etc.) annotation does not ‘spoil ’ a corpus, since the annotation is done ‘in addition ’ to the raw corpus. Thus, on the contrary, annotation may and should bring an additional value to the corpus. Necessary conditions for this aim are: • its scenario is carefully (i.e. systematically and consistently) designed, and • it is based on a sound linguistic theory. This view is corroborated by the existence of annotated corpora of various languages such as Penn Treebank (English), its successors as PropBank or Penn Discourse Treebank,
CT resulted in variable functional and structural changes in dementia, and conclusions are limited by heterogeneity and study quality. Larger, more robust studies are required to correlate these findings with clinical benefits from CT.
We present a two stage parser that recovers Penn Treebank style syntactic analyses of new sentences including skeletal syntactic structure, and, for the first time, both function tags and empty categories. The accuracy of the first-stage parser on the standard Parseval metric matches that of the (Collins, 2003) parser on which it is based, despite the data fragmentation caused by the greatly enriched space of possible node labels. This first stage simultaneously achieves near state-of-the-art performance on recovering function tags with minimal modifications to the underlying parser, modifying less than ten lines of code. The second stage achieves state-of-the-art performance on the recovery of empty categories by combining a linguistically-informed architecture and a rich feature set with the power of modern machine learning methods.
We present an automatic approach to tree annotation in which basic nonterminal symbols are alternately split and merged to maximize the likelihood of a training treebank. Starting with a simple X-bar grammar, we learn a new grammar whose nonterminals are subsymbols of the original nonterminals. In contrast with previous work, we are able to split various terminals to different degrees, as appropriate to the actual complexity in the data. Our grammars automatically learn the kinds of linguistic distinctions exhibited in previous work on manual tree annotation. On the other hand, our grammars are much more compact and substantially more accurate than previous work on automatic annotation. Despite its simplicity, our best grammar achieves an F1 of 90.2% on the Penn Treebank, higher than fully lexicalized systems.
Since ancient linguistics, the studies of Indo-European word order work with the conception of universal natural word order (ordo naturalis) – an order of the verb-dependent constituents in the linear organization of a clause. The description of the natural word order is usually based on occasional (and in some degree random) observations of clauses in a certain language. In Czech linguistics, the idea of the natural word order was formulated in a more precise way as the hypothesis of the systemic ordering (Sgall, Hajicova and Buraňova, 1980). According to the authors, the contextually non-bound participants and adverbials are ordered as follows (o. c., page 77):
We present the implementation of a system which extracts not only lexicalized grammars but also feature-based lexicalized grammars from Korean Sejong Treebank. We report on some practical experiments where we extract TAG grammars and tree schemata. Above all, full-scale syntactic tags and well-formed morphological analysis in Sejong Treebank allow us to extract syntactic features. In addition, we modify Treebank for extracting lexicalized grammars and convert lexicalized grammars into tree schemata to resolve limited lexical coverage problem of extracted lexicalized grammars.
This presentation reports the methodology followed and the results attained on an on-going project aiming at building a large lexical database of corpus-extracted multiword (MW) expressions for the Portuguese language. MW expressions were automatically extracted from a balanced 50 million word corpus compiled for this project, furthermore statistically interpreted using lexical association measures and are undergoing a manual validation process. The lexical database covers different types of MW expressions, from named entities to lexical associations with different degrees of cohesion, ranging from totally frozen idioms to favoured co-occurring forms, like collocations. We aim to achieve two main objectives with this resource: to build on the large set of data of different types of MW expressions to revise existing typologies of collocations and to integrate them in a larger theory of MW units; to use the extensive hand-checked data as training data to evaluate existing statistical lexical association measures.