Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Among the variety of proposals currently making the dependency perspective on grammar more concrete, there are several treebanks whose annotation exploits some form of Relational Structure that we can consider a generalization of the fundamental idea of dependency at various degrees and with reference to different types of linguistic knowledge. The paper describes the Relational Structure as the common underlying representation of treebanks which is motivated by both theoretical and task-dependent considerations. Then it presents a system for the annotation of the Relational Structure in treebanks, called Augmented Relational Structure, which allows for a systematic annotation of various components of linguistic knowledge crucial in several tasks. Finally, it shows a dependency-based annotation for an Italian treebank, i.e. the Turin University Treebank, that implements the Augmented Relational Structure. 1
Despite much research, the distinctive personality characteristics of entrepreneurs are yet to be established and the influence of personality on entrepreneurial behaviour remains unclear. This is particularly evident in our understanding of the personal response of entrepreneurs to business failure. In this thesis the Life Story Model of Identity proposed by McAdams' (1993; McAdams & Pals, 2006) narrative theory of personality formed the main theoretical approach to investigating these two related aspects in the psychological understanding of entrepreneurs. This model overcomes some of the limitations of previous personality research by permitting investigation of personality within the entrepreneurial environment and provides a wholistic and complex view of personality as expressed in the entrepreneurs' own words. The model's qualitative methodology and theoretical emphasis on personal meaning making also rendered it most suitable for exploring entrepreneurs' personal response to business failure. McAdams' (1993; McAdams & Pals, 2006) Life Story Interview was employed to explore the self-narrative identities of 40 highly successful entrepreneurs (39 males, one female). Participants were managing directors of businesses sourced from two lists of the fastest growing small to medium companies in Australia, as compiled by the Australian business magazine, the 'Business Review Weekly'. Participants were the founders of their businesses, and had been pursuing entrepreneurship for at least five years. A broad range of business sectors were represented, including computer services, manufacturing, engineering and communications. Prior to interviews, participants completed the Life Story Interview Questionnaire (LSIQ), which was an adapted form of the Life Story Interview that requested written responses to open-ended questions about the content of participants' life stories. A section requesting affective ratings for key events, derived from Herman’s (Hermans & Hermans-Jansen, 1995) Extended List of Affect Terms, was included to further the exploration of life story themes. A second questionnaire, comprised of measures of personality and a measure of psychological symptoms was also completed. During interviews, participants' responses to the LSIQ were discussed, concentrating on further investigation of the key events that defined their life stories. Findings revealed a prototypical life story of the entrepreneur, highlighting distinctive, commonly shared personality characteristics, with much of their selfnarrative identity grounded in experiences within the entrepreneurial environment. The prototypical life story contained a core theme with an agentic-type emphasis on strengthening the self, and a lesser theme with a communion-type emphasis on valuing relationships. Each of these themes comprised two further themes. The selfstrengthening theme included a redemptive theme of overcoming difficulties in a way that left the protagonist feeling stronger and more able to influence their environment, and a positively toned theme of drawing strength and confidence in one's abilities from achievements and successes. The relational theme included a redemptive theme of responding to private relationship difficulties and losses in one area by strengthening other private relationships, and a negatively toned, sometimes contaminated theme, of experiencing either private or professional relational difficulties and losses as irresolvable. The resulting prototypical life story of the entrepreneur was a story centred upon overcoming adversity and celebrating personal achievement, of confirming and boosting confidence in one’s abilities and a sense of personal power to influence their environment. Running parallel to this main storyline was a less prominent plot involving the importance of relationships, with difficulties and losses sometimes redeemed and sometimes left unresolved. To investigate the impact of business failure, participants were asked to describe their experience of business failure as a key life story event during the Life Story Interview. Additional open-ended questions explored important elements of their critical and retrospective responses. A commonly shared personal response was evidenced. Despite being strongly identified with their business at the time, the failure was evaluated in business rather than personal terms. The causes were most often attributed to a combination of internal and external factors, but business recovery was attributed exclusively to their own actions. The self-narrative meanings given to this event centred upon overcoming the business failure in a self-strengthening way as either: mastering business conflict situations; learning entrepreneurial skills; or affirming entrepreneurial self-confidence. In the midst of the failure, most entrepreneurs remained highly optimistic about their chances of success in the future, based largely upon confidence in their ability to bring about business recovery. Practised coping skills were used to manage negative feelings arising from the failure, and an active problem-solving approach was adopted. When reflecting upon their experience, there was an absence of rumination and regret about the business failure. Instead, it was regarded as an inevitable and even welcome event that provided valuable entrepreneurial learning. In making sense of the failure in relation to the rest of their self-narrative identity, most entrepreneurs were able to integrate its meaning within their larger self-story. This was done by relating the business’ recovery to a story of overcoming obstacles, or by containing the business’ failure within a story of either repeated success or sustained self-confidence in one's ability. It was concluded that these entrepreneurs shared a particular type of selfnarrative identity that was conducive to the pursuit of entrepreneurship; positively influencing their behaviour within the entrepreneurial environment and having particular relevance to how they personally responded to business failure. These findings advance understanding of the personality of entrepreneurs, and begin to inform what constitutes a constructive personal response to the event of business failure.
A statistical estimator attempts to guess an unknown probability distribution by analyzing a sample from this distribution. One desirable property of an estimator is that its guess is increasingly likely to get arbitrarily close to the actual distribution as the sample size increases. This property is called consistency. Data Oriented Parsing (DOP) employs all fragments of the trees in a training treebank, including the full parse-trees themselves, as the rewrite rules of a probabilistic tree-substitution grammar. Since the most popular DOP-estimator (DOP1) was shown to be inconsistent, there is an outstanding theoretical question concerning the possibility of DOP-estimators with reasonable statistical properties. This question constitutes the topic of the current paper. First, we show that, contrary to common wisdom, any unbiased estimator for DOP is futile because it will not generalize over the training treebank. Subsequently, we show that a consistent estimator that generalizes over the treebank should involve a local smoothing technique. This exposes the relation between DOP and existing memory-based models that work with full memory and an analogical function such as k-nearest neighbor, which is known to implement backoff smoothing. Finally, we present a new consistent backoff-based estimator for DOP and discuss how it combines the memory-based preference for the longest match with the probabilistic preference for the most frequent match.
The present work falls in the line of activities promoted by the European Languguage Resource Association (ELRA) Production Committee (PCom) and raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources. A two-fold purpose lies behind this experiment. The first aim is to investigate the feasibility, define methods and procedures for combining two Italian lexical resources that have incompatible formats and complementary information into a Unified Lexicon (UL). The adopted strategy and the procedures appointed are described together with the driving criterion of the merging task, where a balance between human and computational efforts is pursued. The coverage of the UL has been maximized, by making use of simple and fast matching procedures. The second aim is to exploit this newly obtained resource for implementing the phonological and morphological layers of the CLIPS lexical database. Implementing these new layers and linking them with the already exisitng syntactic and semantic layers is not a trivial task. The constraints imposed by the model, the impact at the architectural level and the solution adopted in order to make the whole database ‘speak ’ efficiently are presented. Advantages vs. disadvantages are discussed. 1. Background and Motivations The work described here raises issues in methods, procedures and tools for the reusability, creation, and management of Language Resources (LRs) and has been
Word-to-word dependency structures are useful for consistent representation and comparable evaluation of parsing results. However, most large-scale treebanks contain various variants of phrase structure trees, since automatic parsers usually produce constituent struc-tures. We present a freely available extensible tool for converting phrase structure to dependencies automatically, and discuss its appli-cation to the NEGRA treebank of German. 1.
Abstract This study attempted to demonstrate an elevated disgust sensitivity in bulimia nervosa. Eleven bulimic patients and 12 control subjects underwent a functional magnetic resonance imaging (fMRI) study in which they were presented with alternating blocks of 40 disgust‐inducing, 40 fear‐inducing and 40 affectively neutral scenes. Each scene was shown for 1.5 s. After completion of all blocks, affective ratings were then determined. The viewing of the disgusting pictures, which had been rated as highly repulsive by the bulimic females, was associated with an activation of the left amygdala and the occipito‐temporal visual cortex. The subjective and brain‐physiological responses did not differ from those of the healthy control subjects. This held true for the fear‐inducing scenes as well. Thus, bulimic patients are not characterized by an increased global disgust sensitivity and they do not show any indication of an altered central processing of generally disgust and fear‐inducing visual stimuli. Copyright © 2004 John Wiley & Sons, Ltd and Eating Disorders Association.
INTRODUCTION Several authors described cases of dissociated impairment in naming nouns and verbs. There are four accounts of this dissociation: (i) patients may have purely lexical damage, which selectively affects verbs or nouns at a late stage of the linguistic processing (phonological or orthographic lexicons) (Rapp & Caramazza, 2002); (ii) the damage affects a lexical device, either at an ortographic-phonological modality-specific level (the lexeme; Levelt et al., 1999) or at a unitary lexical-syntactic level (the lemma) (Berndt et al., 1997); (iii) N-V dissociation arises from a semantic damage (Bird, Howard & Franklin, 2000); (iv) N-V dissociation is due to syntactic damage (Friedmann & Grodzinsky, 2000). To disentangle imageability and grammatical class effects, a new task was developed allowing to elicit nouns and verbs with identical imageability ratings in a sentence context. The results obtained will permit to address the following three questions: Does imageability play a role in determining N-V dissociation? If so, is imageability the unique cause of dissociation? If there is additional damage, at which level of linguistic processing does it take place? METHODS Twelve Italian aphasic patients and 11 normal controls participated in the study. Nouns and Verbs Retrieval in a Sentence Context (NVR-SC). Forty-five pairs of sentences denoting the same event, either using a noun or the corresponding verb (e.g. the evasion/to evade) were used. The first sentence was presented in complete form, while a gap was left in the second sentence, to be completed with the target word. For each pair of sentences two different conditions were employed, one triggering a verb and one triggering a noun. E.g. V-N: The prisoner was dreaming to evade The prisoner was dreaming the......... N-V: The prisoner was dreaming the evasion The prisoner was dreaming to......... The performance in the NVR-SC task was compared to that obtained on a classic picture naming task (50 nouns and 50 verbs). Statistical methods. Logistic regression analysis (LRA) was applied to the profiles of the patients, making it possible to study the effects of the lexical-semantic variables in univariate and multivariate linear models. RESULTS Picture naming task All patients with predominant verb deficit also have an imageability effect. In eight patients the grammatical class effect was no longer significant after introducing imageability in the statistical design (bivariate LRA). NVR-SC task Two of the verb-impaired patients in the picture naming task maintained the predominant verb deficit also in the NVR-SC task. In eight patients, the difference between nouns and verbs was no longer significant. In two patients, a paradoxical dissociation (V>N) emerged. Group analysis: performance on nouns and verbs across naming tasks (Figure 1) Patients named actions in the NVR-SC task better than in the picture naming task (58% correct versus 37%; p<.001). On the contrary, the naming of objects in the picture naming task was better than in the NVR-SC task (76% versus 61%; p<.05). DISCUSSION As expected, imageability effect is highly associated with noun-superiority. This result may have two explanations. (a) Since nouns are generally more imaginable than verbs, the imageability effect may cause noun-superiority. But, imageability alone cannot completely account for predominant verb impairment. In fact, four patients have predominant verb impairment even when imageability has been partialled out (bivariate LRA) and two patients are still dissociated in the NVR-SC task (where imageability of nouns and verbs was perfectly matched). Therefore, additional damage must be hypothesized. This would account for the fact that nouns are named better in the picture naming task, and verbs in the NVR-SC task. The former result is explained by the reduced imageability ratings; the latter may only be explained by localizing the additional damage at a central, lexical-syntactic level (i.e. the lemma). This explanation would account for the better performance on verbs in the NVR-SC task: the sentence frame and the correspondent noun may provide the patients with those information lost with the lemma damage, i.e. number of arguments and thematic roles. Nonetheless, this may account also for the fact that noun retrieval is not enhanced in the NVR-SC task: the thematic grid is not useful to retrieve nouns, since it is not a crucial aspect of noun lexical representation. In addition, aphasic patients may use a compensatory strategy to face their deficit. Since the thematic grid may be inferred from a mental image, this strategy will probably rely on visual representations of actions. Thus, the effectiveness of the compensatory strategy increases in relation to imageability (Luzzatti & Chierchia, 2002). (b) When left hemisphere language areas are completely damaged, lexical representations located in the right hemisphere emerge. This emergence is limited to high-frequency concrete nouns (Coltheart, 2000).
One morning each of us received a phone call from Ed Hovy. Are you sitting down? he asked. He told us that as a way to combat conference overload, and to promote interaction among communities, a joint conference had been proposed to combine HLT and NAACL. A diverse oversight committee had been formed, and according to Ed, this committee had been able to agree on two people -- and only two people -- as program co-chairs, because together we represented all of the vested interests. Marti was meant to represent the standards and tastes of the NAACL and the SIGIR crowds, and Mari the speech community, and both have been working on research contracts with HLT funders. Ed told us that if either of us said no, the entire enterprise would come crashing down. There are few better ways to convince busy people to become program co-chairs. Throughout the process, Ed provided the vision for and the drive behind this conference. We salute him for making this idea a reality, and for his enthusiastic and energetic phone calls that kept everything going. This is an exciting time for research in human language technologies. After years of relative calm, the field seems suddenly to be moving by leaps and bounds. Evidence of this can be found in our conference panel on Preparing for a Surprise Language (and as embodied in the short paper Desperately Seeking Cebuano). This panel will discuss the experiences of several groups of researchers, who at the behest of DARPA, acquired and developed language resources for an entirely new language within a span of only 10 days. This experiment took place in March of 2003, and the language in question was Cebuano, a language spoken in the Philippines. Participants successfully collected a large body of lexical and textual resources and developed a range of tools, including stemmers and POS taggers. (In June, DARPA will announce a new surprise language.) The existence of a variety of language resources, combined with advances in statistical analysis and modeling techniques, is resulting in fast-paced improvements in the field. parsers can now produce syntax trees for long sentences with high accuracy and great speed. Advances are starting to be made in automated semantic analysis. Great strides are being made in the sophistication and coverage of question answering systems. Speech recognition systems have achieved suficiently high accuracy that it is now possible to do retrieval, information extraction and topic tracking on spoken documents. Large and growing collections of text and speech corpora -- and the promise of much more from the web -- have enabled many of these advances. New developments in weakly supervised and unsupervised learning algorithms are critical for taking advantage of many new data sources, and hence this was chosen as a special theme of the conference. Lexical resources such as FrameNet, WordNet, PropBank, MeSH, and the Penn TreeBank also play prominent roles in HLT advances. As a field, human language technologies research should use, as motivation and guide, an understanding of the linguistic and cognitive bases of language. The invited talk by Dr. Elissa Newport, entitled Statistical language learning: Mechanisms for language acquisition in human learners, should help enlighten the community by informing us about the latest in psycholinguistic research. We received 162 submissions for full papers, of which 37 were accepted, resulting in a highly competitive acceptance rate of 22%. For the short (late-breaking) papers track, we received 80 submissions, of which 41 were accepted (2 later withdrawn). Some of these will be presented as short talks, and others as posters. Seventeen demonstrations will be shown. We were fortunate to be able to accept 15 papers that addressed the conference theme of unsupervised and weakly supervised methods. We also encouraged papers that described techniques that cross over or combine NLP, speech and/or IR, and several of the papers demonstrate this kind of crossover. The full paper reviewing was done using a two-tier system. First, two first-tier reviewers read every paper. Then a third reviewer, known as the meta-reviewer, wrote their own review. Finally, the meta-reviewer summarized these reviews and introduced additional comments. In some cases, the meta-reviewer instigated discussion among the first-tier reviewers to work out controversial issues. The meta-reviewers also attended the program committee meeting in which all the papers were discussed and acceptances were decided. For the short papers, each short paper received at least two reviews. Those papers whose reviewers disagreed, or which received middling scores, were subsequently reviewed by a member of the program committee and the program co-chairs. Paper submission and reviewing was done online using Marti's conference reviewing software (Conga), which she updated for this conference. Marti also maintained the conference website.
Recent work in machine translation and information extraction has demonstrated the utility of a level that represents the predicate-argument structure. It would be especially useful for machine translation to have two such Proposition Banks, one for each language under consideration. A Proposition Bank for English has been developed over the last few years, and we describe here our development of a tool for facilitating the development of a Chinese Proposition Bank. We also discuss some issues specific to the Chinese Treebank that complicate the matter of mapping syntactic representation to a predicate-argument level, and report on some preliminary evaluation of the accuracy of the semantic tagging tool. 1
Many extensions to text-based, data-intensive knowledge management approaches, such as Information Retrieval or Data Mining, focus on integrating the impressive recent advances in language technology. For this, they need fast, robust parsers that deliver linguistic data which is meaningful for the subsequent processing stages. This paper introduces such a parsing system. Its output is a hierarchical structure of syntactic relations, functional dependency structures.
Startle probe modulation during affective picture viewing was assessed in a Spanish prison population. As for North American inmates, psychopaths failed to display normal blink potentiation during unpleasant slides even though their evaluative judgments and autonomic reaction to affective stimuli paralleled those of other inmate and noninmate participants. The results suggest that diminished defense activation characterizes psychopaths despite cultural differences.
This paper discusses an automatic, data-driven approach to treebank error detection. The approach adapts the use of so-called variation n-grams as defined in Dickinson and Meurers (2003) for the detection of inconsistent part-of-speech annotations to syntactic annotation. The underlying idea is to define a consistency test for the mapping from recurring strings to their syntactic annotation. The paper illustrates with a case study based on the WSJ treebank that the method successfully detects inconsistencies in syntactic category annotation. Since such inconsistencies are typically introduced by humans, our method works best for large corpora that have been annotated manually or semi-automatically, which is generally the case for current syntactic and other high-level annotation. \n \nOur work serves two main purposes for treebank improvement. It is a means for finding erroneous variation in a corpus, which can then be corrected. And it provides feedback for the development of empirically adequate standards for syntactic annotation, showing which distinctions are difficult to maintain over an entire corpus. Additionally, as a method for comparing syntactic annotation, our work could have uses for interannotator agreement testing and parser evaluation.
The main idea of the dynamic lexical chain algorithm is presented. The algorithm relies on the Wordnet thesaurus as a lexical database without requiring full semantic interpretation on an arbitrary text. The lexical chain extracted from the text can produce a summary of the original text.
Acronyms are a very dynamic area of the lexicon of many languages. A hybrid, modular methodology for the acquisition of acronyms is presented, which uses an existing acronym-expansion matching component, and machine learning in two separate phases for the identification of long-distance acronym definition patterns.The resulting system, using Support Vector Machines (SVM) is trained on 600 news stories from the Wall Street Journal component of the Penn Treebank corpus using a number of lexical, syntactic, and acronym-expansion matching features. Statistical cooccurrence information for acronym-expansion pairs is extracted from search engine hit counts.The system achieves Fβ=1=92.38% on 400 news stories from the same source and has good asymptotic efficiency, making it adequate for the automatic extraction of acronyms even from noisy sources, such as newspaper text.
We present a neural-network-based statistical parser, trained and tested on the Penn Treebank. The neural network is used to estimate the parameters of a generative model of left-corner parsing, and these parameters are used to search for the most probable parse. The parser's performance (88.8% F-measure) is within 1% of the best current parsers for this task, despite using a small vocabulary size (512 inputs). Crucial to this success is the neural network architecture's ability to induce a finite representation of the unbounded parse history, and the biasing of this induction in a linguistically appropriate way.
We explore learning prepositionalphrase attachment in Dutch, to use it as a filter in prosodic phrasing. From a syntactic treebank of spoken Dutch we extract instances of the attachment of prepositional phrases to either a governing verb or noun. Using cross-validated parameter and feature selection, we train two learning algorithms, IB1 and RIPPER, on making this distinction, based on unigram and bigram lexical features and a cooccurrence feature derived from WWW counts. We optimize the learning on noun attachment, since in a second stage we use the attachment decision for blocking the incorrect placement of phrase boundaries before prepositional phrases attached to the preceding noun. On noun attachment, IB1 attains an F-score of 82; RIPPER an F-score of 78. When used as a filter for prosodic phrasing, using attachment decisions from IB1 yields the best improvement on precision (by six points to 71) on phrase boundary placement.
We investigate the performance of the Structured Language Model (SLM) in terms of perplexity (PPL) when its components are modeled by connectionist models. The connectionist models use a distributed representation of the items in the history and make much better use of contexts than currently used interpolated or back-off models, not only because of the inherent capability of the connectionist model in fighting the data sparseness problem, but also because of the sublinear growth in the model size when the context length is increased. The connectionist models can be further trained by an EM procedure, similar to the previously used procedure for training the SLM. Our experiments show that the connectionist models can significantly improve the PPL over the interpolated and back-off models on the UPENN Treebank corpora, after interpolating with a baseline trigram language model. The EM training procedure can improve the connectionist models further, by using hidden events obtained by the SLM parser.
The purpose of this study was to investigate preservice teacher intensity, time on and off task, and effectiveness in relation to in-service teacher retention/attrition. Senior music education majors (N = 150) made a videotape of their “best” teaching during the final weeks of student teaching. Excerpts were observed, evaluated, and teacher time on/off task was recorded using the Simple Recording Interface for Behavioral Observation (Duke & Farra, 1996). Each subject's teaching status as K—college music educator in 1995 and 2001 was then determined. No differences were found for intensity, effectiveness, or time on / off task in relation to teaching status assessed, suggesting that preservice teaching performance does not indicate continuing participation in music education as an in-service music teacher. Further analysis revealed that mean interval duration of teacher off task affects ratings of intensity and effectiveness more than do frequency and duration of on task. Application to teacher training and suggestions for further areas of research are discussed.
This article explores the possibilities of automatic extraction of both surface and valency frames of Czech verbs. First, it is clearly documented that the data from Prague Dependency Treebank is not sufficient for collecting enough examples of verb frames to build a large scale lexicon. As a solution, an approach to pick nice examples of sentences from any texts is suggested and thoroughly described. A new scripting language to simplify the selection of sentences based on linguistic criteria was implemented and its main concepts are presented here, too. Also the problems of extracting surface and valency frames from the collected data are addressed and illustrated on real corpus data.
We aim at finding the minimal set of fragments that achieves maximal parse accuracy in Data Oriented Parsing (DOP). Experiments with the Penn Wall Street Journal (WSJ) treebank show that counts of almost arbitrary fragments within parse trees are important, leading to improved parse accuracy over previous models tested on this treebank. We isolate a number of dependency relations which previous models neglect but which contribute to higher accuracy. We show that the history of statistical parsing models displays a tendency towards using more and larger fragments from training data.
Abstract. This paper explores the use of initial Stochastic Context-Free Grammars (SCFG) obtained from a treebank corpus for the learning of SCFG by means of estimation algorithms. A hybrid language model is defined as a combination of a word-based n-gram, which is used to capture the local relations between words, and a category-based SCFG with a word distribution into categories, which is defined to represent the long-term relations between these categories. Experiments on the UPenn Treebank corpus are reported. These experiments have been carried out in terms of the test set perplexity and the word error rate in a speech recognition experiment. 1
INTRODUCTION: Although evidence suggests that interpersonal psychotherapy may be an efficacious treatment for eating disorders, there is surprisingly little systematic knowledge about the interpersonal world of these patients. METHOD: SASB self-image ratings were used to explore interpersonal profiles in a large heterogeneous sample of eating disorders (N = 830), matched normal controls (N = 105) and a small group of controls with subclinical depression (N = 26). RESULTS: Eating disorder patients clearly presented with significantly more negative interpersonal profiles compared to controls. Within the eating disorder group, anorexics were characterized by high self-control, self-blame and self-attack. Patients with binge eating disorder expressed the least negative self-image, and were significantly more self-affirming than bulimics and less self-controlling than patients with atypical eating disorders. CONCLUSIONS: Eating disorder patients may have distinct interpersonal profiles that increase the risk of negative therapeutic reaction. Better knowledge of interpersonal processes in eating disorders may help to improve both diagnostic assessment and treatment.
We describe an algorithm for recovering non-local dependencies in syntactic dependency structures. The pattern-matching approach proposed by Johnson (2002) for a similar task for phrase structure trees is extended with machine learning techniques. The algorithm is essentially a classifier that predicts a non-local dependency given a connected fragment of a dependency structure and a set of structural features for this fragment. Evaluating the algorithm on the Penn Treebank shows an improvement of both precision and recall, compared to the results presented in (Johnson, 2002).
We investigate the performance of the Structured Language Model when one of its components is modeled by a connectionist model. Using a connectionist model and a distributed representation of the items in the history makes the component able to use much longer contexts than possible with currently used interpolated or backoff models, both because of the inherent capability of the connectionist model to fight the data sparseness problem, and because of the only sub-linear growth in the model size when increasing the context length. Experiments show significant improvement in perplexity and moderate reduction in word error rate over the baseline SLM results on the UPENN treebank and Wall Street Journal (WSJ) corpora respectively. The results also show that the probability distribution obtained by our model is much less correlated to regular N-grams than the baseline SLM model.
This paper reports on the use of two distinct evaluation metrics for assessing a stochastic parsing model consisting of a broad-coverage Lexical-Functional Grammar (LFG), an efficient constraint-based parser and a stochastic disambiguation model. The first evaluation metric measures matches of predicate-argument relations in LFG f-structures (henceforth the LFG annotation scheme) to a gold standard of manually annotated f-structures for a subset of the UPenn Wall Street Journal treebank. The other metric maps predicate-argument relations in LFG f-structures to dependency relations (henceforth DR annotations) as proposed by Carroll et al. (Carroll et al., 1999). For evaluation, these relations are matched against Carroll et al.&apos;s gold standard which was manually annnotated on a subset of the Brown corpus. The parser plus stochastic disambiguator gives an F-measure of 79% (LFG) or 73% (DR) on the WSJ test set. This shows that the two evaluation schemes are similar in spirit, although accuracy is impaired systematically by mapping one annotation scheme to the other. A systematic loss of accuracy is incurred also by corpus variation: Training the stochastic disambiguation model on WSJ data and testing on Carroll et al.&apos;s Brown corpus data yields an F-score of 74% (DR) for dependency-relation match. A variant of this measure comparable to the measure reported by Carroll et al. yields an F-measure of 76%. We examine divergences between annotation schemes aiming at a future improvement of methods for assessing parser quality.
We have developed an example-based machine translation (EBMT) system that uses the World Wide Web for two different purposes: First, we populate the system's memory with translations gathered from rule-based MT systems located on the Web. The source strings input to these systems were extracted automatically from an extremely small subset of the rule types in the Penn-II Treebank. In subsequent stages, the source, target translation pairs obtained are automatically transformed into a series of resources that render the translation process more successful. Despite the fact that the output from on-line MT systems is often faulty, we demonstrate in a number of experiments that when used to seed the memories of an EBMT system, they can in fact prove useful in generating translations of high quality in a robust fashion. In addition, we demonstrate the relative gain of EBMT in comparison to on-line systems. Second, despite the perception that the documents available on the Web are of questionable quality, we demonstrate in contrast that such resources are extremely useful in automatically postediting translation candidates proposed by our system.
The primary purpose of this study was to assess the cross-cultural invariance of job performance ratings. A secondary purpose was to examine potential cross-cultural differences in correlates of performance ratings (i.e., ratee sex, age, tenure; supervisor's opportunity to observe ratee). Fast-food supervisors from Canada, South Korea, and Spain rated employees on their technical proficiency, customer service, and teamwork. Results show that these ratings demonstrate a basic level of measurement invariance, although the error variances of the ratings and pattern of construct variances and covariances were largely culture-specific. This suggests that supervisors across cultures may use and interpret the ratings similarly, but perceive differences in performance. Furthermore, age, tenure, and the supervisor's opportunity to observe the ratee were found to affect ratings differently across cultures. Overall, this study suggests that although job performance ratings are at least partially invariant across cultures, latent performance may not be, and we present some preliminary data as to why latent invariance may not exist.
In this paper we will present work carried out lately on the 50,000 words Italian Spontaneous Speech Corpus called AVIP, under national project API, made available for free download from the website of the coordinator, the University of Naples. We will concentrate on the tuning of the parser for Italian which had been previously used to parse 100,000 words corpus of written Italian within the National Treebank initiative coordinated by ILC in Pisa. The parser receives as input the adequately transformed orthographic transcription of the dialogues making up the corpus, in which pauses, hesitations and other disfluencies have been turned into most likely corresponding punctiation marks, interjections or truncation of the word underlying the uttered segment.\nThe most interesting phenomenon we will discuss is without any doubts "overlapping", i.e. a speech event in which two people speak at the same time by uttering actual words or in some cases nonwords, when one of the speakers, usually the one which is not the current turntaker, interrupts the current speaker.\nThis phenomenon takes place at a certain point in time where it has to be anchored to the speech signal but in order to be fully parsed and subsequently semantically interpreted, it needs to be referred semantically to a following turn.
El estudio experimental de la emoción requiere de estímulos que evoquen en una forma confiable reacciones psicológicas y fisiológicas que varien sistemáticamente sobre el rango de emociones de acuerdo a las dimensiones de valencia (agradable o desagradable), activación (excitado o calmado) y dominancia (alta y baja) (Lang, Bradley, Cuthbert, 1999). A pesar de que los correlatos neurales de las emociones básicas han sido investigados, la organización neural de las "emociones morales" en el cerebro humano no se conocen bien. El objetivo de la presente investigación fue obtener un grupo de estimulos diferenciados (fotografías) y caracterizarlos en términos de su valencia afectiva, activación, dominancia, y contenido moral, en una población mexicana. Se seleccionaron fotografías que representan escenas con una carga emocional amplia como violaciones morales (escenas de guerra, asaltos físicos, etc), escenas aversivas sin connotación moral (tumores, cuerpos mutilados) y escenas naturales (toallas, mesas, puertas, etc. ). Los sujetos evaluaron cada fotografía de acuerdo a su valencia, activación, dominancia y contenido moral (ausente o extremo). Para la evaluación, se utilizó la Escala Internacional Self-Assessment Maniki Affective Rating System desarrollada por Lang (1980). Se discute las implicaciones de los datos, para el estudio de las emociones y del juicio moral.
After many successes, statistical approaches that have been popular in the parsing community are now making headway into Natural Language Generation (NLG). These systems are aimed mainly at surface realization, and promise the same advantages that make statistics valuable for parsing: robustness, wide coverage and domain independence. A recent experiment aimed to empirically verify the linguistic coverage for such a statistical surface realization component by generating transformed sentences from the Penn TreeBank corpus. This article presents the empirical results of a similar experiment to evaluate the coverage of a purely symbolic surface realizer. We present the problems facing a symbolic approach on the same task, describe the results of its evaluation, and contrast them with the results of the statistical method to help quantitatively determine the level of coverage currently obtained by NLG surface realizers. 1
BACKGROUND: Individual differences in neural circuitry that regulate emotional reactivity may be associated with alcoholism and antisocial personality disorder (ASPD), a common comorbid condition. The emotion-modulated startle reflex was used to investigate emotional reactivity among alcohol-dependent (AD) men with and without ASPD. METHODS: Sixty-two men were tested: (1) AD (n = 24), (2) AD-ASPD (n = 17), and (3) non-AD, non-ASPD controls (n = 21). Participants completed self-report instruments and clinical interviews and had eye-blink electromyograms measured in response to acoustic startle probes while viewing color photographs rated as affectively pleasant, neutral, and unpleasant. RESULTS: Startle blink magnitudes were larger during unpleasant as compared with pleasant slides for control and AD groups, resulting in significant linear trend effects (p < 0.001) and nonsignificant quadratic trend effects. In contrast, AD-ASPD did not show a significant difference in blink magnitude during unpleasant and pleasant slides and did not show a significant linear valence trend or quadratic trend effect (p > 0.6). Subjective valence and arousal ratings of the photographs were similar across groups. CONCLUSIONS: Adult male alcoholics with ASPD have abnormal emotional responsiveness to both pleasant and unpleasant stimuli relative to alcoholics without ASPD and to controls.
Treebanks are a valuable resource for the training of parsers that perform automatic annotation of unseen data. It has been shown that changes in the representation of linguistic annotation have an impact on the performance of a certain annotation task. We focus on the task of Topological Field Parsing for German using Probabilistic Context-Free Grammars in the present research. We investigate an iterative algorithm for tuning the label set of a given treebank to this task and show that the number of parses proposed by a context-free grammar is reduced considerably in addition to an increase in labeled precision and recall for the annotation of node labels. We also show that the optimal refinement can be achieved with a relatively small number of changes to the treebank. 1
Agents seeking to discover and compose needed Web services may face knowledge sharing interoperability problems due to differing ontologies. In practice, agents may not have a global consensus ontology that will facilitate knowledge sharing and integration of required services. We investigate a method for agents to develop local consensus ontologies to aid in the communication within a multi-agent system of business-tobusiness (B2B) agents. We compare variations of syntactic and semantic similarity matching to form local consensus ontologies with and without the use of a lexical database.
Since Eloise Jelinek has been interested in the issues of negation, focus and information structure, to the research of which she has contributed substantially, we want to use this nice occasion and present here partial results of an analysis of the Topic-Focus articulation (TFA) of Czech and of the impact of these results on inquiries into coreferrence in coherent discourse. In Czech linguistics, TFA has been systematically explored thanks to the classical Prague School of functional and structural linguistics. As reflecting the ‘given – new ’ strategy in discourse, TFA has been considered to belong to the main objects of linguistic study. Continuing the results gained by V. Mathesius, J. Firbas and others since the 1920s, the explicit linguistic descriptive framework characterized in Sgall et al. (1986), Hajičová (1993), Hajičová E., Partee B. and P. Sgall (1998) includes a possibility to describe TFA not only as concerning the intrinsic dynamics of the process of communication, patterned in the utterance (sentence occurrence), but also as constituting the structure of the sentence itself, i.e. grammar. Within this framework, TFA is understood as one of the basic aspects of (underlying) sentence structure, which characterizes the sentence as a unit of the interactive system of language; TFA thus is seen as a manifestation of the sentence being anchored in the context.
In this paper, we propose a method for analyzing word-word dependencies using deterministic bottom-up manner using Support Vector machines. We experimented with dependency trees converted from Penn treebank data, and achieved over 90 % accuracy of word-word dependency. Though the result is little worse than the most up-to-date phrase structure based parsers, it looks satisfactorily accurate considering that our parser uses no information from phrase structures. 1
Past research has found that individual differences in both attitudinal and situational variables may be associated with males’ likelihood of acquaintance rape (LAR). The present research was conducted to examine the predictive value of both attitudinal and situational factors on males’ likelihood of forcing a female acquaintance to have non-consensual sexual intercourse. In Study 1, male and female respondents (Rs) were presented with a scenario depicting a hypothetical sexual interaction between the respondent and a newly acquainted member of the opposite sex. As the encounter progressed from one sexual activity to the next, Rs made three ratings regarding their own and partner’s intent to engage in each successive activity. The scenario ended with the female refusing further activity and males’ affect ratings, adherence to attitudes conducive of rape, and LAR were measured. Males’ initial perceptions of female sexual intent (to later engage in sexual intercourse) best predicted LAR. Study 2 was conducted to examine the role of female sexual communication on perceptions of consent to sexual intercourse. Rs were presented with a scenario similar to that in Study 1, but at each stage were requested to rate the extent to which the female had consented to engage in each sexual activity. Males completed the same affect and attitudinal measures. The results of Study 2 again suggested that males’ initial perception of female consent to (later) engage in sexual intercourse best predicted LAR. The present research suggests that further investigation into the role of situational factors and males’ initial perception of sexual intent and consent in the aetiology of acquaintance rape is required.
The linguistic annotation of natural language corpora is one of the main areas of computational linguistics. Much energy has been devoted to building large syntactically annotated corpora, which are also called treebanks after the phrase-structure trees they contain. For some time now, functional information, i.e. information on whether a constituent functions as e. g. subject, object or adverbial, has also been included in the annotation. Yet, while at first glance this may not seem to be a venture too complicated, matters are not always as easy as they seem.
Abstract This paper describes a framework for building story traces (compact global views of a narrative) and story projections (selections of key elements of a narrative) and their applications in text understanding and classification. Word and sense properties are extracted from documents using the WordNet lexical database enhanced with Prolog inference rules and a number of lexical transformations. Inference rules are based on navigation in various WordNet relation chains (hypernyms, meronyms, entailment and causality links, etc.) and derived relations expressed as Boolean combinations of node and edge properties used to direct the navigation. The resulting abstract story traces provide a compact view of the underlying narrative's key content elements and a means for automated indexing and classification of text collections. Ontology driven projections act as a kind of “semantic lenses” and provide a means to select a subset of a narrative whose key sense elements are subsumed by a set of concepts, predicates and properties expressing the focus of interest of a user. Finally, we discuss applications of these techniques in text understanding, classification of text collections and answering questions about a text.
This paper deals with the problem of how to interrelate theory-specific treebanks and how to transform one treebank format to another. Currently, two approaches to achieve these goals can be differentiated. The first creates a mapping algorithm between treebank formats. Categories of a source format are transformed into a target format via a given set of general or language-specific mapping rules. The second relates treebanks via a transformation to a general model of linguistic categories, for example based on the EAGLES recommendations for syntactic annotations of corpora, or relying on the HPSG framework. This paper proposes a new methodology as a solution for these desiderata.