Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We show that jointly parsing a bitext can substantially improve parse quality on both sides. In a maximum entropy bitext parsing model, we define a distribution over source trees, target trees, and node-to-node alignments between them. Features include monolingual parse scores and various measures of syntactic divergence. Using the translated portion of the Chinese treebank, our model is trained iteratively to maximize the marginal likelihood of training tree pairs, with alignments treated as latent variables. The resulting bitext parser outperforms state-of-the-art monolingual parser baselines by 2.5 F1 at predicting English side trees and 1.8 F1 at predicting Chinese side trees (the highest published numbers on these corpora). Moreover, these improved trees yield a 2.4 BLEU increase when used in a downstream MT evaluation.
Recently, most of the research in NLP has concentrated on the creation of applications coping with textual entailment. However, there still exist very few resources for the evaluation of such applications. We argue that the reason for this resides not only in the novelty of the research field but also and mainly in the difficulty of defining the linguistic phenomena which are responsible for inference. As the TSNLP project has shown test suites provide optimal diagnostic and evaluation tools for NLP applications, as contrary to text corpora they provide a deep insight in the linguistic phenomena allowing control over the data. Thus in this paper, we present a test suite specifically developed for studying inference problems shown by English adjectives. The construction of the test suite is based on the deep linguistic analysis and following classification of entailment patterns of adjectives and follows the TSNLP guidelines on linguistic databases providing a clear coverage, systematic annotation of inference tasks, large reusability and simple maintenance. With the design of this test suite we aim at creating a resource supporting the evaluation of computational systems handling natural language inference and in particular at providing a benchmark against which to evaluate and compare existing semantic analysers. 1.
We have built a parallel treebank that includes word and phrase alignment. The alignment information was manually checked using a graphical tool that allows the annotator to view a pair of trees from parallel sentences. We found the compilation of clear alignment guidelines to be a difficult task. However, experiments with a group of students have shown that we are on the right track with up to 89% overlap between the student annotation and our own. At the same time these experiments have helped us to pin-point the weaknesses in the guidelines, many of which concerned unclear rules related to differences in grammatical forms between the languages.
Probabilistic modeling of lexicalized grammars is difficult because these grammars exploit complicated data structures, such as typed feature structures. This prevents us from applying common methods of probabilistic modeling in which a complete structure is divided into sub-structures under the assumption of statistical independence among sub-structures. For example, part-of-speech tagging of a sentence is decomposed into tagging of each word, and CFG parsing is split into applications of CFG rules. These methods have relied on the structure of the target problem, namely lattices or trees, and cannot be applied to graph structures including typed feature structures. This article proposes the feature forest model as a solution to the problem of probabilistic modeling of complex data structures including typed feature structures. The feature forest model provides a method for probabilistic modeling without the independence assumption when probabilistic events are represented with feature forests. Feature forests are generic data structures that represent ambiguous trees in a packed forest structure. Feature forest models are maximum entropy models defined over feature forests. A dynamic programming algorithm is proposed for maximum entropy estimation without unpacking feature forests. Thus probabilistic modeling of any data structures is possible when they are represented by feature forests. This article also describes methods for representing HPSG syntactic structures and predicate-argument structures with feature forests. Hence, we describe a complete strategy for developing probabilistic models for HPSG parsing. The effectiveness of the proposed methods is empirically evaluated through parsing experiments on the Penn Treebank, and the promise of applicability to parsing of real-world sentences is discussed.
In this paper, we give a description of the machine translation (MT) system developed at DCU that was used for our third participation in the evaluation campaign of the International Workshop on Spoken Language Translation (IWSLT 2008). In this participation, we focus on various techniques for word and phrase alignment to improve system quality. Specifically, we try out our word packing and syntax-enhanced word alignment techniques for the Chinese–English task and for the English–Chinese task for the first time. For all translation tasks except Arabic–English, we exploit linguistically motivated bilingual phrase pairs extracted from parallel treebanks. We smooth our translation tables with out-of-domain word translations for the Arabic–English and Chinese–English tasks in order to solve the problem of the high number of out of vocabulary items. We also carried out experiments combining both in-domain and out-of-domain data to improve system performance and, finally, we deploy a majority voting procedure combining a language modelbased method and a translation-based method for case and punctuation restoration. We participated in all the translation tasks and translated both the single-best ASR hypotheses and the correct recognition results. The translation results confirm that our new word and phrase alignment techniques are often helpful in improving translation quality, and the data combination method we proposed can significantly improve system performance.
This paper presents recent advances in an established treebank annotation framework comprising of an abstract XML-based data format, fully customizable editor of tree-based annotations, a toolkit for all kinds of automated data processing with support for cluster computing, and a work-in-progress database-driven search engine with a graphical user interface built into the tree editor.
The mere exposure effect is the commonly observed increase in pleasantness ratings of stimuli that have been given prior exposure. According to the fluency attribution account of the mere exposure effect, repeated presentations of a stimulus lead to increased ease of processing, which in turn is attributed to pleasantness. If so, processing fluency manipulated by means other than repetition should influence liking. In the present experiment, processing fluency was manipulated using a negative priming procedure, and its influence on affective judgement was examined. Previously ignored stimuli were responded to slower (negative priming) and were rated as less pleasant than controls. It was concluded that decreased processing fluency decreases liking of previously ignored stimuli.
We present the first results on parsing the SynTagRus treebank of Russian with a data-driven dependency parser, achieving a labeled attachment score of over 82% and an unlabeled attachment score of 89%. A feature analysis shows that high parsing accuracy is crucially dependent on the use of both lexical and morphological features. We conjecture that the latter result can be generalized to richly inflected languages in general, provided that sufficient amounts of training data are available.
The need for syntactically annotated data for use in natural language processing has increased dramatically in recent years. This is true especially for parallel treebanks, of which very few exist. The ones that exist are mainly hand-crafted and too small for reliable use in data-oriented applications. In this paper we introduce a novel platform for fast and robust automatic generation of parallel treebanks. The software we have developed based on this platform has been shown to handle large data sets. We also present evaluation results demonstrating the quality of the derived treebanks and discuss some possible modifications and improvements that can lead to even better results. We expect the presented platform to help boost research in the field of data-oriented machine translation and lead to advancements in other fields where parallel treebanks can be employed.
We describe a parsing approach that makes use of the perceptron algorithm, in conjunction with dynamic programming methods, to recover full constituent-based parse trees. The formalism allows a rich set of parse-tree features, including PCFG-based features, bigram and trigram dependency features, and surface features. A severe challenge in applying such an approach to full syntactic parsing is the efficiency of the parsing algorithms involved. We show that efficient training is feasible, using a Tree Adjoining Grammar (TAG) based parsing formalism. A lower-order dependency parsing model is used to restrict the search space of the full model, thereby making it efficient. Experiments on the Penn WSJ treebank show that the model achieves state-of-the-art performance, for both constituent and dependency accuracy.
Specific phobias are the most common anxiety disorder and are characterized by avoidance behavior. Avoidance behavior impacts daily function and is proposed to impair extinction learning. However, despite its prevalence, its objective assessment remains a challenge. To this end, we developed a fully automated experimental procedure using immersive virtual reality. The procedure contained a behavioral search, forced-choice, and an approach task with varying degrees of freedom and task relevance of the stimuli. In this study, we examined the sensitivity and feasibility of these tasks to assess avoidance behavior in patients with specific phobia. We adapted the tasks by replacing the originally conditioned stimuli with a spider and a neutral animal and investigated 31 female participants composed of 15 spider-phobic and 16 non-phobic participants. As the non-phobics were quite heterogeneous in terms of their Fear of Spiders Questionnaire (FSQ) scores, we subdivided them into six "fearfuls" that had elevated FSQ scores, and 10 "non-fearfuls" that had no fear of spiders. The phobics successfully managed to complete the procedure and showed consistent avoidance behavior across all behavioral tasks. Compared to the non-fearfuls, which did not show any avoidance behavior at all, the phobics looked at the spider much more often and clearly directed their body toward it in the search task. In the approach task, they hesitated most when they were close to the spider, and their difficulty to touch the spider was reflected in a strong increase in right hand acceleration changes. The fearfuls showed avoidance behavior depending on the tasks: strongest in the search task and weakest in the approach task. Additionally, we identified subjective valence ratings of the spider as the main influence on both objective avoidance behavior and subjective well-being after exposure, mediating the effect of the FSQ. In summary, the behavioral tasks are well suited to assess avoidance behavior in phobic participants and provide detailed insights into the process of avoidance.
We consider the value of replacing and/or combining string-based methods with syntax-based methods for phrase-based statistical machine translation (PBSMT), \nand we also consider the relative merits of using constituency-annotated vs. dependency-annotated training data. We automatically derive two subtree-aligned treebanks, \ndependency-based and constituency-based, from a parallel English–French corpus and extract syntactically motivated word- and phrase-pairs. We automatically measure PB-SMT quality. The results show that combining string-based and syntax-based word- and phrase-pairs can improve translation quality irrespective of the type of syntactic annotation. Furthermore, using dependency annotation yields greater translation quality than constituency annotation for PB-SMT.
Computational, descriptive, and theoretical linguistics use both phrase (PS) structure and dependency structure (DS) to represent syntax. We believe that the next-generation treebank should be multi-representational, designed for both representations with an automatic conversion. In this paper, we highlight the assumptions made by existing PS-to-DS and DS-to-PS conversion algorithms and show the limitations of these algorithms. We then propose a new DS-to-PS conversion algorithm that outperforms existing algorithms and allows more flexibility. Our experiments and error analysis show that high-quality DS-to-PS conversion is possible if the conversion process is performed at the designing stage of treebank construction to ensure that all information we wish to represent in PS is provided in DS. 1
In this paper, we describe the TeP 2.0 -- Electronic Thesaurus for Brazilian Portuguese -- which stores sets of synonym and antonym word forms. Specifically, we present the lexical database and the Web interface of TeP 2.0.
Drawing on motivational approaches to emotion, the authors propose that the perceived change in spatial distance to pictures that arouse negative emotions exerts an influence on the significance of these pictures. Two experiments induced the illusion that affective pictures approach toward the observer, recede from the observer, or remain static. To determine the motivational significance of the pictures, emotional valence and arousal ratings as well as startle responses were assessed. Approaching unpleasant pictures were found to exert an influence on both the valence and the arousal elicited by the pictures. Furthermore, movement of pleasant or neutral pictures did not influence startle responses, while the second experiment showed that approaching unpleasant pictures elicited enhanced startle responses compared to receding unpleasant pictures. These findings support the view that a change of spatial distance influences motivational significance and thereby shapes emotional responses.
We propose a cascaded linear model for joint Chinese word segmentation and partof-speech tagging. With a character-based perceptron as the core, combined with realvalued features such as language models, the cascaded model is able to efficiently utilize knowledge sources that are inconvenient to incorporate into the perceptron directly. Experiments show that the cascaded model achieves improved accuracies on both segmentation only and joint segmentation and part-of-speech tagging. On the Penn Chinese Treebank 5.0, we obtain an error reduction of 18.5 % on segmentation and 12 % on joint segmentation and part-of-speech tagging over the perceptron-only baseline. 1
The studies on lexicographic definitions connected with the French tradition take charge eminently of typology and leave aside the question of metalanguage. So, in lexicography, the metalinguistic definition is often considered in the typological frame. This is because the above-mentioned studies are mostly based upon definitions either of nouns or verbs. In my presentation I shall attempt to demonstrate, from defining statements of the syncategorematic words drawn from the Tresor de la langue francaise, that the metalinguistic definition is indeed a category of the definitions but that, when compared to the other categories, it requires a different criteria of analysi, due to its nature. In order to do this, I shall present, first, the different nature of this issue from a typological approach on one side and a metalinguistic approach on the other. I shall expose, then, the main typological studies-in particular the unpublished document which is stored in the archives of the Laboratory ATILF [.Pour un nouveau cahier de normes...., 1979] as well as Martin (1983) and Rey-Debove (1998)-in which the question of the metalanguage is dealt with inside and following the example of typology to demonstrate that, if a definition such as aiguillette-nom populaire de l.orphie-is metalinguistic and a definition such as chaise - siege a dossier sans bras - is perifrastic, nom et siege are both hyperonyms, so that the typological criteria are not enough to distinguish between mealinguistic and perifrastic definition. Thus, I will establish, in accordance with Rey-Debove (1997), in which the definition is considered from a metalinguistic point of view-according to the sintactic relation between a lexical entry and its lexicographical definition, the principles which govern the metalinguistic analysis. The results will lead to three different categories of metalinguistic definitions of the syncategorematic words: 1. the definition refers to both infralinguistic and extralinguistic reality-in this case two sub-categories are possible: a) the hyperonym refers to the infralinguistic reality while the specific semes refer to the extralinguistic reality; b) the hyperonym refers to the infralinguistic reality while the specific semes, among which there is at least an autonym with 'schize' (cf. Rey-Debove 1997: 116-118), refer to the extralinguistic reality; 2. the definition refers to the only infralinguistic reality; 3. the definition refers to the only extralinguistic reality.
This paper deals with a multilingual relational lexical database of proper name, Prolexbase, a free resource available on the CNRTL website. The Prolex model is based on two main concepts: firstly, a language independent pivot and, secondly, the prolexeme (the projection of the pivot onto particular language), that is a set of lemmas (names and derivatives). These two concepts model the variations of proper name: firstly, independent of language and, secondly, language dependent by morphology or knowledge. Variation processing is very important for NLP: the same proper name can be written in different instances, maybe in different parts of speech, and it can also be replaced by another one, a lexical anaphora (that reveals semantic link). The pivot represents different referent's points of view, i.e. language independent variations of name. Pivots are linked by three semantic relations (quasi-synonymy, partitive relation and associative relation). The prolexeme is a set of variants (aliases), quasi-synonyms and morphosemantic derivatives. Prolexemes are linked to classifying contexts and reliability code.
\n Dans cette proposition, nous plaidons pour une meilleure mutualisation des résultats de recherche sur le lexique à travers loutil informatique que représente le Web. Après avoir analysé les conditions de réussite dune telle mutualisation et limportance des normes et standards en ce domaine, nous montrons quelques exemples de réussite dune telle mutualisation tant en lexicographie contemporaine, à travers le Trésor de la langue française informatisé, quen lexicographie historique. Ainsi à travers le DMF (Dictionnaire du Moyen Français), nous explicitons le concept nouveau de lexicographie évolutive et montrons quelques exemples de résultats de recherche qui nauraient pas vu le jour sans sappuyer sur la richesse dexploitation inégalée, rendue possible grâce à son informatisation: \n - en lexicologie, par exemple sur la datation dapparition de sens nouveaux dun lexème dans la langue,\n - en pragmatique, à travers lexemple dune anté-datation de près de deux siècles de lusage de enfin énumératif,\n - ou en morphologie constructionnelle, à travers létude des formations en inr- qui pour certaines furent ensuite abandonnées au profit de formation en irr- (tel inrégulier versus irrégulier).\nToujours dans le domaine de la lexicographie historique, nous montrons lintérêt de mutualiser nos connaissances sur létymologie, tel quil se pratique dans le projet TLF-Etym, ou sur les « mots fantômes », pseudo lexèmes disposant à tort dun statut lexicographique (« ces mots qui nexistent pas »), et les lemmatisations erronées qui se trouvent encore trop souvent dans les dictionnaires historiques et étymologiques français de référence.\nNous terminons enfin par la présentation dun exemple dintégration et de valorisation de données lexicographiques et lexicales au sein du portail lexical du Centre National de Ressources Textuelles et Lexicales (CNRTL, www.cnrtl.fr) qui à travers les quelque 300 000 requêtes quil sert par jour est aujourdhui une magnifique vitrine des résultats de recherche en lexicographie, morpho-syntaxe, étymologie, synonymie, antonymie.\n\n
Morphological processes in Semitic languages deliver space-delimited words which introduce multiple, distinct, syntactic units into the structure of the input sentence. These words are in turn highly ambiguous, breaking the assumption underlying most parsers that the yield of a tree for a given sentence is known in advance. Here we propose a single joint model for performing both morphological segmentation and syntactic disambiguation which bypasses the associated circularity. Using a treebank grammar, a data-driven lexicon, and a linguistically motivated unknown-tokens handling technique our model outperforms previous pipelined, integrated or factorized systems for Hebrew morphological and syntactic processing, yielding an error reduction of 12% over the best published results so far. 1
Parser self-training is the technique of taking an existing parser, parsing extra data and then creating a second parser by treating the extra data as further training data. Here we apply this technique to parser adaptation. In particular, we self-train the standard Charniak/Johnson Penn-Treebank parser using unlabeled biomedical abstracts. This achieves an f-score of 84.3% on a standard test set of biomedical abstracts from the Genia corpus. This is a 20% error reduction over the best previous result on biomedical data (80.2% on the same test set).
CONTEXT: Cognitive decline, mood, behavioral and sleep disturbances, and limitations of activities of daily living commonly burden elderly patients with dementia and their caregivers. Circadian rhythm disturbances have been associated with these symptoms. OBJECTIVE: To determine whether the progression of cognitive and noncognitive symptoms may be ameliorated by individual or combined long-term application of the 2 major synchronizers of the circadian timing system: bright light and melatonin. DESIGN, SETTING, AND PARTICIPANTS: A long-term, double-blind, placebo-controlled, 2 x 2 factorial randomized trial performed from 1999 to 2004 with 189 residents of 12 group care facilities in the Netherlands; mean (SD) age, 85.8 (5.5) years; 90% were female and 87% had dementia. INTERVENTIONS: Random assignment by facility to long-term daily treatment with whole-day bright (+/- 1000 lux) or dim (+/- 300 lux) light and by participant to evening melatonin (2.5 mg) or placebo for a mean (SD) of 15 (12) months (maximum period of 3.5 years). MAIN OUTCOME MEASURES: Standardized scales for cognitive and noncognitive symptoms, limitations of activities of daily living, and adverse effects assessed every 6 months. RESULTS: Light attenuated cognitive deterioration by a mean of 0.9 points (95% confidence interval [CI], 0.04-1.71) on the Mini-Mental State Examination or a relative 5%. Light also ameliorated depressive symptoms by 1.5 points (95% CI, 0.24-2.70) on the Cornell Scale for Depression in Dementia or a relative 19%, and attenuated the increase in functional limitations over time by 1.8 points per year (95% CI, 0.61-2.92) on the nurse-informant activities of daily living scale or a relative 53% difference. Melatonin shortened sleep onset latency by 8.2 minutes (95% CI, 1.08-15.38) or 19% and increased sleep duration by 27 minutes (95% CI, 9-46) or 6%. However, melatonin adversely affected scores on the Philadelphia Geriatric Centre Affect Rating Scale, both for positive affect (-0.5 points; 95% CI, -0.10 to -1.00) and negative affect (0.8 points; 95% CI, 0.20-1.44). Melatonin also increased withdrawn behavior by 1.02 points (95% CI, 0.18-1.86) on the Multi Observational Scale for Elderly Subjects scale, although this effect was not seen if given in combination with light. Combined treatment also attenuated aggressive behavior by 3.9 points (95% CI, 0.88-6.92) on the Cohen-Mansfield Agitation Index or 9%, increased sleep efficiency by 3.5% (95% CI, 0.8%-6.1%), and improved nocturnal restlessness by 1.00 minute per hour each year (95% CI, 0.26-1.78) or 9% (treatment x time effect). CONCLUSIONS: Light has a modest benefit in improving some cognitive and noncognitive symptoms of dementia. To counteract the adverse effect of melatonin on mood, it is recommended only in combination with light. TRIAL REGISTRATION: controlled-trials.com/isrctn Identifier: ISRCTN93133646.
In the context of crises in which emergency services or the general population are of different languages, effective interoperability requires not only that translations of messages and alerts be done rapidly but also, being safety critical, that there be no errors. We have developed a methodology based on linguistic norms and a supporting mathematical model for the construction of a single source controlled language to be machine translated to specific target controlled languages. In this paper we discuss in particular the architecture of our machine translation system which is based on the `canonical¿ case where there are no language divergences (identical source and target languages), and the `variant¿ cases encompassing the divergences between each target controlled language and our source controlled language. We explain the way that we classify and organize the divergences in a declarative manner so as to be incorporated in the machine translation process.
Disordered gambling stigma was examined. University students (117 male, 132 female) rated vignettes describing males with five health conditions (schizophrenia, alcohol dependence, disordered gambling, cancer, and a no diagnosis control with subclinical problems) on a measure of attitudinal social distance. A mixed ANOVA revealed that, in keeping with hypotheses, disordered gambling was more stigmatized than the cancer and control conditions. Interactions suggested that stigma may be influenced by context (i.e., order of vignette appearance) and participant characteristics (i.e., sex and ethnicity), although follow–up analyses revealed this was not the case for disordered gambling. Perceived dangerousness attributions and familiarity (previous experience with a disordered gambler) were also examined. As predicted, perceived dangerousness was positively correlated with social distance scores. Familiarity ratings were unrelated to social distance.
The paper describes the treatment of some specific syntactic constructions in two treebanks of Latin according to a common set of annotation guidelines. Both projects work within the theoretical framework of Dependency Grammar, which has been demonstrated to be an especially appropriate framework for the representation of languages with a moderately free word order, where the linear order of constituents is broken up with elements of other constituents. The two projects are the first of their kind for Latin, so no prior established guidelines for syntactic annotation are available to rely on. The general model for the adopted style of representation is that used by the Prague Dependency Treebank, with departures arising from the Latin grammar of Pinkster, specifically in the traditional grammatical categories of the ablative absolute, the accusative + infinitive, and gerunds/gerundives. Sharing common annotation guidelines allows us to compare the datasets of the two treebanks for tasks such as mutually checking annotation consistency, diachronically studying specific syntactic constructions, and training statistical dependency parsers.
Parsing algorithms that process the input from left to right and construct a single derivation have often been considered inadequate for natural language parsing because of the massive ambiguity typically found in natural language grammars. Nevertheless, it has been shown that such algorithms, combined with treebank-induced classifiers, can be used to build highly accurate disambiguating parsers, in particular for dependency-based syntactic representations. In this article, we first present a general framework for describing and analyzing algorithms for deterministic incremental dependency parsing, formalized as transition systems. We then describe and analyze two families of such algorithms: stack-based and list-based algorithms. In the former family, which is restricted to projective dependency structures, we describe an arc-eager and an arc-standard variant; in the latter family, we present a projective and a non-projective variant. For each of the four algorithms, we give proofs of correctness and complexity. In addition, we perform an experimental evaluation of all algorithms in combination with SVM classifiers for predicting the next parsing action, using data from thirteen languages. We show that all four algorithms give competitive accuracy, although the non-projective list-based algorithm generally outperforms the projective algorithms for languages with a non-negligible proportion of non-projective constructions. However, the projective algorithms often produce comparable results when combined with the technique known as pseudo-projective parsing. The linear time complexity of the stack-based algorithms gives them an advantage with respect to efficiency both in learning and in parsing, but the projective list-based algorithm turns out to be equally efficient in practice. Moreover, when the projective algorithms are used to implement pseudo-projective parsing, they sometimes become less efficient in parsing (but not in learning) than the non-projective list-based algorithm. Although most of the algorithms have been partially described in the literature before, this is the first comprehensive analysis and evaluation of the algorithms within a unified framework.
Although punctuation is pervasive in written text, their treatment in parsers and corpora is often second-class. We examine the treatment of commas in CCGbank, a wide-coverage corpus for Combinatory Categorial Grammar (CCG), reanalysing its comma structures in order to eliminate a class of redundant rules, obtaining a more consistent treebank. We then eliminate these rules from C&C, a wide-coverage statistical CCG parser, obtaining a 37 % increase in parsing speed on the standard CCGbank test set and a considerable reduction in memory consumed, without affecting parser accuracy. 1
We present the LFG PARSEBANKER, a comprehensive toolkit for interactive incremental construction of a treebank as a parsed corpus. This web-based toolkit offers an environment for batch and interactive parsing, versioning, inspection of structures, discriminant-based disambiguation, and statistics. It has recently been extended with a structural search facility.
Expression of the serotonin transporter is affected by the genotype of the 5-HTTLPR (short and long forms) as well as the genotype of the SNP rs25531 within this region. Based on the combined genotypes for these polymorphisms, we designated each allele as a high or low expressing allele according to established expression levels-resulting in HiHi, HiLo, & LoLo genotype groups for analysis. We evaluated effects of gender and the promoter genotype on induction of negative affect by intravenous infusion of L: -tryptophan (TRP). The protocol consisted of a day-1 sham saline infusion and a day-2 active TRP infusion. Models assessed 5-HTTLPR composite genotype and gender as predictors of change in ratings of negative emotion during TRP infusion. During sham infusion there were no significant changes from baseline in mood ratings. During TRP infusion all negative affect ratings increased significantly from baseline (P's <.02). The genotype x gender interaction was a significant predictor of depression-dejection (P =.013), and trended towards predicting anger-hostility (P =.084). Males in the HiHi group had greater increases in negative affect during infusion, compared to all groups except LoLo females, who also showed increased negative affect.
Robust spoken language understanding (SLU) is a key component of spoken dialogue systems. Recent statistical approaches to this problem require additional resources (e.g. gazetteers, grammars, syntactic treebanks) which are expensive and time-consuming to produce and maintain. However, simple datasets annotated only with slot-values are commonly used in dialogue systems development, and are easy to collect, automatically annotate, and update. We show that it is possible to reach state-of-the-art performance using minimal additional resources, by using Markov logic networks (MLNs). We also show that performance can be further improved by exploiting long distance dependencies between slot-values. For example, by representing such features in MLNs, but without using a gazetteer, we outperform the hidden vector state (HVS) model of He and Young 2006 (1.26% improvement, a 13% error reduction).
We present a new English→Czech machine translation system combining linguistically motivated layers of language description (as defined in the Prague Dependency Treebank annotation scenario) with statistical NLP approaches.
The use of authentic text has been argued to increase learner awareness of lexical form, function, and meaning (for example, Willis 1990; Johns 1994). The Web provides ready-made material and tools for both learner-centred reading and vocabulary tasks. This study reports on the results of a project in which Japanese university EFL students made use of the Web as a living corpus to investigate the specific contexts and collocative properties of lexis. Using an online database, students created a communal dictionary composed of lexis and example sentences culled from web sources, along with examples of their own devising. The language database was then used to facilitate peer teaching of lexis. Work produced indicates that learners paid attention to lexical form, function, and meaning when composing.
Purpose Customer satisfaction is seen to be one of the main determinants of loyalty. However, the relationship between customer satisfaction and loyalty does not seem to be linear, many researchers have reported doubts about the predictability of loyalty solely due to customer satisfaction ratings which ignore image as predictor of loyalty. This paper aims to address the issues. Design/methodology/approach The authors report a study of ski resorts where they first established a causal model of customer satisfaction and image predicting customer loyalty, and then map the scores in a four‐fields‐grid. Additionally the authors conducted a moderator analysis to assess the relative importance of image and satisfaction for loyalty intentions between two different groups (first‐time‐visitors, and regular guests). Findings The results show that those ski resorts with the highest satisfaction ratings and the highest image ratings have the highest loyalty scores. Among first‐time‐visitors overall satisfaction is more important than image, with increasing number of repeat visits the importance of overall satisfaction declines and that of image relatively augments. Practical implications Besides measuring customer satisfaction, managers must assess also image ratings in order to get a realistic view of the loyalty intentions of their customer base. The scores can than be mapped together with the ratings of other ski resorts, and serve as a benchmark study. Originality/value Second order analysis of image (comprising three different dimensions), the image‐satisfaction‐grid, moderating effect of experience to relative importance of satisfaction and image on loyalty.
Natural language processing modules such as part-of-speech taggers, named-entity recognizers and syntactic parsers are commonly evaluated in isolation, under the assumption that artificial evaluation metrics for individual parts are predictive of practical performance of more complex language technology systems that perform practical tasks. Although this is an important issue in the design and engineering of systems that use natural language input, it is often unclear how the accuracy of an end-user application is affected by parameters that affect individual NLP modules. We explore this issue in the context of a specific task by examining the relationship between the accuracy of a syntactic parser and the overall performance of an information extraction system for biomedical text that includes the parser as one of its components. We present an empirical investigation of the relationship between factors that affect the accuracy of syntactic analysis, and how the difference in parse accuracy affects the overall system.
We present procedures which pool lexical information estimated from unlabeled data via the Inside-Outside algorithm, with lexical information from a treebank PCFG. The procedures produce substantial improvements (up to 31.6% error reduction) on the task of determining subcategorization frames of novel verbs, relative to a smoothed Penn Treebank-trained PCFG. Even with relatively small quantities of unlabeled training data, the re-estimated models show promising improvements in labeled bracketing f-scores on Wall Street Journal parsing, and substantial benefit in acquiring the subcategorization preferences of low-frequency verbs.
Periods in the development of the lexical database in the Czech Language Institute, programmes.
Despite recent calls for establishing a paremiological minimum for the United States, there has been little systematic empirical attention devoted to the issue of proverb familiarity in this country or to the possibility that geographic differences might preclude establishing a single paremiological minimum that would be equally representative of proverb familiarity in diverse regions around the country. This article reviews the call for establishing a paremiological minimum for the United States and the relevant efforts to date. The article then summarizes the results of a study of proverb familiarity among college students in four different regions of the United States, presenting results from both a proverb familiarity rating task and a proverb generation task. These data make it possible to determine the relative familiarity of several hundred proverbs and to assess the possibility that familiarity with these proverbs varies across regions. The success of this approach is evidence that this type of nomothetic and quantitative analysis of folkloric material can be a useful complement to performance-oriented research. For example, the results suggest some limits to the verisimilitude of people’s intuitions as to the currency and familiarity of proverbial texts, as proverbs often listed as “common” were sometimes shown to be relatively unfamiliar (at least to contemporary college students). Further, although statistical analyses of the data did reveal that there may be some difference in absolute levels of familiarity across regions, it was also clear that the relative familiarity of various proverbs nonetheless shows strong stability from one region to another. In general, then, the data presented here provide clear evidence that proverbs familiar in one region of the United States can generally be expected to be familiar in other regions as well; this finding suggests that a truly national paremiological minimum may well be achievable.
We present a simple and effective semisupervised method for training dependency parsers. We focus on the problem of lexical representation, introducing features that incorporate word clusters derived from a large unannotated corpus. We demonstrate the effectiveness of the approach in a series of dependency parsing experiments on the Penn Treebank and Prague Dependency Treebank, and we show that the cluster-based features yield substantial gains in performance across a wide range of conditions. For example, in the case of English unlabeled second-order parsing, we improve from a baseline accuracy of 92:02% to 93:16%, and in the case of Czech unlabeled second-order parsing, we improve from a baseline accuracy of 86:13% to 87:13%. In addition, we demonstrate that our method also improves performance when small amounts of training data are available, and can roughly halve the amount of supervised data required to reach a desired level of performance.
Description of lexicographical procedures used during the work with corpus material, programmes.
This paper discusses the role of annotated corpora as works of reference for grammatical translation problems. Within this context, the English-German CroCo Corpus and its multi-layer alignment and annotation are introduced. It is described how the corpus is exploited as interactive resource to display translation solutions for typologically problematic constructions. Additionally, the Penn and TiGer Treebanks are used as comparable corpora for English and German. The linguistic enrichment of the treebanks, i.e. their syntactic annotation, is described and corpus query techniques relevant for translation problems are shown. Relevant structures are extracted from the treebanks and translation candidates are displayed and discussed. The advantage of this technique is that translation solutions are extracted from published translations, i.e. language in use. Consequently, they are more comprehensive and inventive than dictionary entries or descriptions in grammars are. Treebanks could thus be used as an interactive reference grammar in translation education and practice.