Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The basic colour terms for black and white are studied in four archaic and two contemporary linguistic norms of the Chinese language. It is presented that studied Chinese linguistic norms use a common term for white and three different terms for black. It is suggested that the different basic colour terms for black might originate from different source languages. The study supports a panchronic language development instead of a diachronic one, and includes introductions to histories of the Chinese linguistic norms
International audience
OBJECTIVES: Understanding the relationship between the menstrual cycle and pain can contribute significantly to our knowledge of pain processing in women. Many early studies suggested that pain sensitivity was enhanced during the luteal phase of the menstrual cycle relative to the follicular phase; however, these studies were often limited by small sample sizes, lack of ovulation verification, focus on a single pain modality, inadequate assessment of menstrual cycle regularity, and low-powered statistical methods. The current study was designed to address these limitations and examine the difference in pain processing between the mid-follicular (days 5 to 8) and late-luteal (days 1 to 6 preceding menses) phases. METHODS: Forty-one healthy, regularly cycling women attended testing sessions that measured pain sensitivity from mechanical pain threshold, electrocutaneous pain threshold/tolerance, and ischemia pain threshold/tolerance, as well as McGill Pain Questionnaire qsensory and affective ratings of electric and ischemic stimuli. Electrocutaneous stimulation was also used to assess nociceptive flexion reflex threshold, a physiological measure of spinal nociception. RESULTS: When analyses were limited to data collected only in the targeted menstrual phases (N=30), results indicated no menstrual phase effect on any pain outcome (all P's>0.05), with the exception of lower electrocutaneous pain thresholds during the late-luteal phase. No outcomes differed by menstrual phase in the full sample (N=41). This indicates nociceptive responding varies little between the mid-follicular and late-luteal phases. DISCUSSION: The present study suggests that experimental pain processing does not significantly differ between the mid-follicular and late-luteal phases of the menstrual cycle in healthy women. This implies hormonal variation across these 2 phases (ie, progesterone) has a minimal effect on subjective and physiological responses to pain.
Implicit discourse relation classification is a challenge task due to missing discourse connective. Some work directly adopted machine learning algorithms and linguistically informed features to address this task. However, one interesting solution is to automatically predict implicit discourse connective. In this paper, we present a novel two-step machine learning-based approach to implicit discourse relation classification. We first use machine learning method to automatically predict the discourse connective that can best express the implicit discourse relation. Then the predicted implicit discourse connective is used to classify the implicit discourse relation. Experiments on Penn Discourse Treebank 2.0 (PDTB) and Biomedical Discourse Relation Bank (BioDRB) show that our method performs better than the baseline system and previous work.
This chapter explores the potential and challenges of basing pedagogical, bilingual electronic dictionaries on WordNet – a comprehensive machine-readable lexical database of the English language. It introduces the specifics of WordNet’s complex lexical architecture; explores the interfaces between WordNet and humanistic, pedagogical and bilingual lexicography; and highlights the Transpoetika Dictionary – an innovative, bilingualised Serbian-English pedagogical dictionary. The chapter shows that the potential of relational lexical databases by far exceeds their exclusive application in language engineering and automatic text processing. The author argues that WordNet can not only serve as a general lexicographic framework for building humanistic eDictionaries, but also pave the way for other collaborative projects in Natural Language Processing (NLP) and pedagogic lexicography. The author also discusses strategies for expanding the scope of WordNet-based humanistic dictionaries through web services and social media platforms such as Flickr and Twitter.
Decoding pain in others is of high individual and social benefit in terms of harm avoidance and demands for accurate care and protection. The processing of facial expressions includes both specific neural activation and automatic congruent facial muscle reactions. While a considerable number of studies investigated the processing of emotional faces, few studies specifically focused on facial expressions of pain. Analyses of brain activity and facial responses elicited by the perception of facial pain expressions in contrast to other emotional expressions may unravel the processing specificities of pain-related information in healthy individuals and may contribute to explaining attentional biases in chronic pain patients. In the present study, 23 participants viewed short video clips of neutral, emotional (joy, fear), and painful facial expressions while affective ratings, event-related brain responses, and facial electromyography (Musculus corrugator supercilii, M. orbicularis oculi, M. zygomaticus major, M. levator labii) were recorded. An emotion recognition task indicated that participants accurately decoded all presented facial expressions. Electromyography analysis suggests a distinct pattern of facial response detected in response to happy faces only. However, emotion-modulated late positive potentials revealed a differential processing of pain expressions compared to the other facial expressions, including fear. Moreover, pain faces were rated as most negative and highly arousing. Results suggest a general processing bias in favor of pain expressions. Findings are discussed in light of attentional demands of pain-related information and communicative aspects of pain expressions.
In the present paper, we describe in detail and evaluate the process of semi-automatic annotation of intra-sentential discourse relations in the Prague Dependency Treebank, which is a part of the project of otherwise mostly manual annotation of all (intra- and inter-sentential) discourse relations with explicit connectives in the treebank. Our assumption that some syntactic features of a sentence analysis (in a form of a deepsyntax dependency tree) correspond to certain discourse-level features proved to be correct, and the rich annotation of the treebank allowed us to automatically detect the intra-sentential discourse relations, their connectives and arguments in most of the cases. TITLE AND ABSTRACT IN CZECH Poloautomatická anotace vnitrovětných diskurzních vztahů v PDT ABSTRAKT V tomto článku nabízíme detailní popis a evaluaci procesu poloautomatické anotace vnitrovětných textových vztahů v Pražském závislostním korpusu jako součást projektu jinak především manuální anotace všech (vnitro- a mezivětných) textových vztahů s explicitním konektorem v tomto korpusu. Potvrdil se náš předpoklad, že některé syntaktické vlastnosti analýzy věty (ve formě závislostního stromu hloubkové syntaxe) odpovídají jistým vlastnostem na úrovni analýzy textových vztahů (diskurzu). Bohatá anotace korpusu nám ve většině případů umožnila automaticky detekovat vnitrovětné vztahy, jejich konektory a argumenty.
PCFGs can grow exponentially as additional annotations are added to an initially simple base grammar. We present an approach where multiple annotations coexist, but in a factored manner that avoids this combinatorial explosion. Our method works with linguisticallymotivated annotations, induced latent structure, lexicalization, or any mix of the three. We use a structured expectation propagation algorithm that makes use of the factored structure in two ways. First, by partitioning the factors, it speeds up parsing exponentially over the unfactored approach. Second, it minimizes the redundancy of the factors during training, improving accuracy over an independent approach. Using purely latent variable annotations, we can efficiently train and parse with up to 8 latent bits per symbol, achieving F1 scores up to 88.4 on the Penn Treebank while using two orders of magnitudes fewer parameters compared to the naïve approach. Combining latent, lexicalized, and unlexicalized annotations, our best parser gets 89.4 F1 on all sentences from section 23 of the Penn Treebank. 1
We describe a transformation-based learning method for learning a sequence of mono-lingual tree transformations that improve the agreement between constituent trees and word alignments in bilingual corpora. Using the manually annotated English Chinese Transla-tion Treebank, we show how our method au-tomatically discovers transformations that ac-commodate differences in English and Chi-nese syntax. Furthermore, when transforma-tions are learned on automatically generated trees and alignments from the same domain as the training data for a syntactic MT system, the transformed trees achieve a 0.9 BLEU im-provement over baseline trees. 1
In the following paper, we discuss and evaluate the benefits that deep syntactic trees (tectogrammatics) and all the rich annotation of the Prague Dependency Treebank bring to the process of annotating the discourse structure, i.e. discourse relations, connectives and their arguments. The decision to annotate discourse structure directly on the trees contrasts with the majority of similarly aimed projects, usually based on the annotation of linear texts. Our basic assumption is that some syntactic features of a sentence analysis correspond to certain discourselevel features. Hence, we use some properties of the dependency-based large-scale treebank of Czech to help establish an independent annotation layer of discourse. The question that we answer in the paper is how much did we gain by employing this approach. TITLE AND ABSTRACT IN CZECH Pomaha tektogramatika při anotaci diskurznich vztahů?
Slang is one of the problems encountered in developing Lexical Database. There is no complete source officially for slangs in any language. This research attempts to create a list of Indonesian slangs using the proposed framework and methodology. The contents for slangs are generated by Twitter; therefore it can encounter the newest slangs available in the society.
The study is based on an experiment to measure the affective states of computer users via their use of mouse and keyboard. The experiment was replicated from a previous study by Khan et al., [5] resulting in significant correlations between the computer users pattern of interactions and their valence, arousal ratings. This study utilized the same data set from [5] and re-confirmed its validity by training Artificial Neural Networks (ANN). The data was divided into two portions for each individual. A portion to train ANN on his/her patterns of interaction and other portion to test the ANN. The study resulted in an average recognition rate of 64.72 % for valence and 61.02 % for arousal ratings. The highest recognition rates for individual participants' valence and arousal were 100% and 87% respectively. These figures suggest that ANN is a bright prospect for the measurement of affective states of individual computer users via their interaction with keyboard and mouse.
The paper analyzes 33 grammatical structures of Tsinghua University 973 Treebank from the syntactic point of view. Firstly, we explore the distribution of these structures in Tsinghua University 973 Treebank and analyze the syntactic constituents of these structures. Then, we gather statistics about these structures based on the external functional relation and the internal structural relation of its subcomponents. Finally, according to the same internal structural relations, we generate a matrix to show the ambiguity among these structures. The statistical data offers the syntactic knowledge for auto-identifying these structures in future use.
The paper presents a small empirical study into emotion and affect recognition based on auditory and visual features, which was performed in the context of the Audio-Visual Emotion Challenge (AVEC) 2012. The goal of this competition is to predict continuous-valued affect ratings based on the provided auditory and visual features, e.g., local binary pattern (LBP) features extracted from aligned face images, and spectral audio features.
Statistical machine translation has been remarkably successful for the world’s well-resourced languages, and much effort is focussed on creating and exploiting rich resources such as treebanks and wordnets. Machine translation can also support the urgent task of documenting the world’s endangered languages. The primary object of statistical translation models, bilingual aligned text, closely coincides with interlinear text, the primary artefact collected in documentary linguistics. It ought to be possible to exploit this similarity in order to improve the quantity and quality of documentation for a language. Yet there are many technical and logistical problems to be addressed, starting with the problem that – for most of the languages in question – no texts or lexicons exist. In this position paper, we examine these challenges, and report on a data collection effort involving 15 endangered languages spoken in the highlands of
An investigation of the relationship between background knowledge and reading comprehension performance on standardized reading tests (the California STAR Test) was conducted with sixth, seventh, and eighth-grade ethnic minority children from low-income backgrounds (N = 68). Predictor variables examined included perceived background knowledge (overall and topic-specific), GPA, basic literacy skills, reading self-concept, race and ethnicity, language background, gender, and grade level. Research questions addressed participants' familiarity with topics discussed in STAR test reading passages and about the predictive nature of participant rankings and ratings of passages, as measured by the Topic Familiarity Ranking Measure and the Topic Familiarity Rating Scale. Results indicated that background knowledge of passage topics had a significant positive association (p <.05) with reading comprehension performance for 30% of the CST passages, for seventh and eighth-grade participants. Hierarchical regression analyses conducted on three of the passages showed that between 7% and 16% of the variance in reading comprehension performance was accounted for by background knowledge, as measured by the Topic Familiarity Rating Scale.
the effect of pathological aging on explicit memory is very well documented, but relatively few studies have addressed this issue in the musical domain. To examine learning and consolidation of melodies, we designed a melodic recognition task involving immediate and delayed recognition of 16 target melodies (8 familiar and 8 unfamiliar). Seventeen patients with mild to moderate Alzheimer's disease (AD) and 17 age-matched controls were tested. During the initial presentation of the targets, the participant had to decide whether or not the melody was familiar. Recognition was tested after one and three presentations of the target melodies using a yes/no recognition paradigm. Delayed recognition was tested after 24 hours to evaluate consolidation. In keeping with the findings of Bartlett, Halpern, and Dowling (1995), age-matched controls showed better recognition of familiar than unfamiliar melodies. Controls also showed improved performance with multiple presentations for both familiar and unfamiliar melodies, without forgetting after 24-hour delay. In contrast, patients with AD showed impaired learning and recognition of both unfamiliar and familiar melodies with no benefit of familiarity on recognition. Nevertheless, the familiarity decision-based ratings of patients was in keeping with controls. These findings suggest that musical recognition memory is impaired in AD, but the musical lexicon (as assessed by familiarity ratings) is preserved. These findings highlight the need to use both familiar and unfamiliar music in experimental tasks to study the different processes underlying recognition memory.
Dependency parsing has attracted considerable interest from researchers and developers in natural language processing. However, to obtain a high‐accuracy dependency parser, supervised techniques require a large volume of hand‐annotated data, which are extremely expensive. This paper presents a simple and effective approach for improving dependency parsing with subtrees derived from unannotated data, which are easy to obtain. First, we use a baseline parser to parse large‐scale unannotated data. Then, we extract subtrees from dependency parse trees in the auto‐parsed data. Next, the extracted subtrees are classified into several sets according to their frequency. Finally, we design new features based on the subtree sets for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we conduct experiments on the English Penn Treebank and Chinese Penn Treebank. The results show that our approach significantly outperforms baseline systems. It also achieves the best accuracy for the Chinese data and an accuracy competitive with the best known systems for the English data.
The paper concentrates on which language means may be included into the annotation of discourse relations in the Prague Dependency Treebank (PDT) and tries to examine the so called alternative lexicalizations of discourse markers (AltLex’s) in Czech. The analysis proceeds from the annotated data of PDT and tries to draw a comparison between the Czech AltLex’s from PDT and English AltLex’s from PDTB (the Penn Discourse Treebank). The paper presents a lexico-syntactic and semantic characterization of the Czech AltLex’s and comments on the current stage of their annotation in PDT. In the current version, PDT contains 306 expressions (within the total 43,955 of sentences) that were labeled by annotators as being an AltLex. However, as the analysis demonstrates, this number is not final. We suppose that it will increase after the further elaboration, as AltLex’s are not restricted to a limited set of syntactic classes and some of them exhibit a great degree of variation. Key words: alternative lexicalization of discourse markers (AltLex); discourse connectives; discourse relations 1.
The common use of a single de facto standard annotation scheme for dependency treebank creation leaves the question open to what extent the performance of an application trained on a treebank depends on this annotation scheme and whether a linguistically richer scheme would imply a decrease of the performance of the application. We investigate the effect of the variation of the number of grammatical relations in a tagset on the performance of dependency parsers. In order to obtain several levels of granularity of the annotation, we design a hierarchical annotation scheme exclusively based on syntactic criteria. The richest annotation contains 60 relations. The more coarse-grained annotations are derived from the richest. As a result, all annotations and thus also the performance of a parser trained on different annotations remain comparable. We carried out experiments with four state-of-the-art dependency parsers. The results support the claim that annotating with more fine-grained syntactic relations does not necessarily imply a significant loss of accuracy. We also show the limits of this approach by giving details on the fine-grained relations that do have a negative impact on the performance of the parsers.
We show that orthographic cues can be helpful for unsupervised parsing. In the Penn Treebank, transitions between upper- and lower-case tokens tend to align with the boundaries of base (English) noun phrases. Such signals can be used as partial bracketing constraints to train a grammar inducer: in our experiments, directed dependency accuracy increased by 2.2% (average over 14 languages having case information). Combining capitalization with punctuation-induced constraints in inference further improved parsing performance, attaining state-of-the-art levels for many languages.
Title: Valency of verbs in the Prague Dependency Treebank Author: PhDr. Zdeňka Urešová Department: Institute of Formal and Applied Linguistics MFF UK Supervisor: Prof. PhDr. Eva Hajičová, DrSc. Abstract: This dissertation describes PDT-Vallex, a valency lexicon of Czech verbs, and its relation to the annotation of the Prague Dependency Treebank (PDT). The PDT-Vallex lexicon was created during the an- notation of the PDT and it is a valuable source of verbal valency information available both for linguistic research and for computer- ized natural language processing. In this thesis, we describe not only the structure and design of the lexicon (which is closely related to the notion of valency as developed in the Functional Generative De- scription of language) but also the relation between the PDT-Vallex and the PDT. The explicit and full-coverage linking of the lexicon to the treebank prompted us to pay special attention to diatheses; we propose formal transformation rules for diatheses to handle their surface realization even when the canonical forms of verb arguments as captured in the lexicon do not correspond to the forms of these arguments actually appearing in the corpus.
This paper presents a theoretical discussion about the use of Frame Semantics as corpora annotation paradigm. The objective of this paper is to evaluate the applicability of Frame Semantics theory and FrameNet paradigm for the semantic annotation of legal texts. The work presented in this paper is an initial step in the construction of a treebank for the Brazilian legal language.
Drawing from social identity theory, this research examines scarce gender representation as a contextual condition that inhibits same‐gender supervisors' support. Survey results in S tudy 1 found that when women were proportionally underrepresented, they reported feeling less supported by female supervisors than male supervisors. S tudy 2 showed that women who perceived they were gender tokens in their organization were less likely to support an outstanding female subordinate than an identical male. S tudy 3 experimentally tested social mobility as a mechanism for the effects of tokenism on same‐gender supervisor support. Results suggest that social mobility and group composition jointly affect ratings of same‐gender targets. Perceptions of gender‐based social mobility appear to be one mechanism through which tokenism influences same‐gender relations at work.
We present a detailed error analysis of a transition-based dependency parser trained on a Hindi dependency treebank. Parser error analysis has not been systematically examined from the point of view of treebanking before and this work intends to contribute in this area. We address two main questions in this paper: Can the parsing of certain structures be made easier by using alternative analyses for these structures? Are there certain linguistic cues implicit (or missing) in the current treebank that can be made explicit (or added) in order to make the parsing of complex constructions easier? These questions will guide us in examining the potential benefits of parser error analysis during treebanking. Through our experiments and analysis we were able to shed light on the causes of errors and subsequently have been able to improve the performance of the parser.
and determiner errors with spelling correction as a pre-processing step. The result shows that spelling correction improves the Detection, Correction, and Recognition F-scores for preposition errors. With regard to preposition error correction, F-scores were not improved when using the training set with correction of all but preposition errors. As for determiner error correction, there was an improvement when the constituent parser was trained with a concatenation of treebank and modified treebank where all the articles appearing as the first word of an NP were removed. Our system ranked third in preposition and fourth in determiner error corrections. 1
The majority of fear conditioning studies in humans have focused on fear acquisition rather than fear extinction. For this reason only a few functional imaging studies on fear extinction are available. A large number of animal studies indicate the medial prefrontal cortex (mPFC) as neuronal substrate of extinction. We therefore determined mPFC contribution during extinction learning after a discriminative fear conditioning in 34 healthy human subjects by using functional near-infrared spectroscopy. During the extinction training, a previously conditioned neutral face (conditioned stimulus, CS+) no longer predicted an aversive scream (unconditioned stimulus, UCS). Considering differential valence and arousal ratings as well as skin conductance responses during the acquisition phase, we found a CS+ related increase in oxygenated haemoglobin concentration changes within the mPFC over the time course of extinction. Late CS+ trials further revealed higher activation than CS- trials in a cluster of probe set channels covering the mPFC. These results are in line with previous findings on extinction and further emphasize the mPFC as significant for associative learning processes. During extinction, the diminished fear association between a former CS+ and a UCS is inversely correlated with mPFC activity--a process presumably dysfunctional in anxiety disorders.
We investigate the stacking of dependency and phrase structure parsers, i.e. we define features from the output of a phrase structure parser for a dependency parser and vice versa. Our features are based on the original form of the external parses and we also compare this approach to converting phrase structures to dependencies then applying standard stacking on the converted output. The proposed method provides high accuracy gains for both phrase structure and dependency parsing. With the features derived from the phrase structures, we achieved a gain of 0.89 percentage points over a state-of-the-art parser and reach 93.95 UAS, which is the highest reported accuracy score on dependency parsing of the Penn Treebank. The phrase structure parser obtains 91.72 F-score with the features derived from the dependency trees, and this is also competitive with the best reported PARSEVAL scores for the Penn Treebank.
OBJECTIVE: To develop a Spanish version of the WHO-Composite International Diagnostic Interview (WHO-CIDI) applicable to Spain, through cultural adaptation of its most recent Latin American (LA v 20.0) version. METHODS: A 1-week training course on the WHO-CIDI was provided by certified trainers. An expert panel reviewed the LA version, identified words or expressions that needed to be adapted to the cultural or linguistic norms for Spain, and proposed alternative expressions that were agreed on through consensus. The entire process was supervised and approved by a member of the WHO-CIDI Editorial Committee. The changes were incorporated into a Computer Assisted Personal Interview (CAPI) format and the feasibility and administration time were pilot tested in a convenience sample of 32 volunteers. RESULTS: A total of 372 questions were slightly modified (almost 7% of approximately 5000 questions in the survey) and incorporated into the CAPI version of the WHO-CIDI. Most of the changes were minor - but important - linguistic adaptations, and others were related to specific Spanish institutions and currency. In the pilot study, the instrument's mean completion administration time was 2h and 10min, with an interquartile range from 1.5 to nearly 3h. All the changes made were tested and officially approved. CONCLUSIONS: The Latin American version of the WHO-CIDI was successfully adapted and pilot-tested in its computerized format and is now ready for use in Spain.
According to an influential dual-process model, a moral judgment is the outcome of a rapid, affect-laden process and a slower, deliberative process. If these outputs conflict, decision time is increased in order to resolve the conflict. Violations of deontological principles proscribing the use of personal force to inflict intentional harm are presumed to elicit negative affect which biases judgments early in the decision-making process. This model was tested in three experiments. Moral dilemmas were classified using (a) decision time and consensus as measures of system conflict and (b) the aforementioned deontological criteria. In Experiment 1, decision time was either unlimited or reduced. The dilemmas asked whether it was appropriate to take a morally questionable action to produce a "greater good" outcome. Limiting decision time reduced the proportion of utilitarian ("yes") decisions, but contrary to the model's predictions, (a) vignettes that involved more deontological violations logged faster decision times, and (b) violation of deontological principles was not predictive of decisional conflict profiles. Experiment 2 ruled out the possibility that time pressure simply makes people more like to say "no." Participants made a first decision under time constraints and a second decision under no time constraints. One group was asked whether it was appropriate to take the morally questionable action while a second group was asked whether it was appropriate to refuse to take the action. The results replicated that of Experiment 1 regardless of whether "yes" or "no" constituted a utilitarian decision. In Experiment 3, participants rated the pleasantness of positive visual stimuli prior to making a decision. Contrary to the model's predictions, the number of deontological decisions increased in the positive affect rating group compared to a group that engaged in a cognitive task or a control group that engaged in neither task. These results are consistent with the view that early moral judgments are influenced by affect. But they are inconsistent with the view that (a) violation of deontological principles are predictive of differences in early, affect-based judgment or that (b) engaging in tasks that are inconsistent with the negative emotional responses elicited by such violations diminishes their impact.
Linear Context-Free Rewriting System (LCFRS) is an extension of Context-Free Grammar (CFG) in which a non-terminal can dominate more than a single continu-ous span of terminals. Probabilistic LCFRS have recently successfully been used for the direct data-driven parsing of discontin-uous structures. In this paper we present a parser for binary PLCFRS of fan-out two, together with a novel monotonous estimate for A ∗ parsing, with which we conduct ex-periments on modified versions of the Ger-man NeGra treebank and the Discontinuous Penn Treebank in which all trees have block degree two. The experiments show that compared to previous work, our approach provides an enormous speed-up while de-livering an output of comparable richness. 1
Lexical substitutes have found use in areas such as paraphrasing, text simplification, machine translation, word sense disambiguation, and part of speech induction. However the computational complexity of accurately identifying the most likely substitutes for a word has made large scale experiments difficult. In this letter we introduce a new search algorithm, FASTSUBS, that is guaranteed to find the K most likely lexical substitutes for a given word in a sentence based on an n-gram language model. The computation is sublinear in both K and the vocabulary size V. An implementation of the algorithm and a dataset with the top 100 substitutes of each token in the WSJ section of the Penn Treebank are available at http://goo.gl/jzKH0.
Domain adaptation is an important topic for natural language processing. There has been extensive research on the topic and various methods have been explored, including training data selection, model combination, semi-supervised learning. In this study, we propose to use a goodness measure, namely, description length gain (DLG), for domain adaptation for Chinese word segmentation. We demonstrate that DLG can help domain adaptation in two ways: as additional features for supervised segmenters to improve system performance, and also as a similarity measure for selecting training data to better match a test set. We evaluated our systems on the Chinese Penn Treebank version 7.0, which has 1.2 million words from five different genres, and the Chinese Word Segmentation Bakeoff-3 data.
From the perspective of structural linguistics, we explore paradigmatic and syntagmatic lexical relations for Chinese POS tagging, an important and challenging task for Chinese language processing. Paradigmatic lexical relations are explicitly captured by word clustering on large-scale unlabeled data and are used to design new features to enhance a discriminative tagger. Syntagmatic lexical relations are implicitly captured by constituent parsing and are utilized via system combination. Experiments on the Penn Chinese Treebank demonstrate the importance of both paradigmatic and syntagmatic relations. Our linguistically motivated approaches yield a relative error reduction of 18 % in total over a stateof-the-art baseline. 1
Employing higher-order subtree structures in graph-based dependency parsing has shown substantial improvement over the accuracy, however suffers from the inefficiency increasing with the order of subtrees. We present a new reranking approach for dependency parsing that can utilize complex subtree representation by applying efficient subtree selection heuristics. We demonstrate the effective-ness of the approach in experiments conducted on the Penn Treebank and the Chinese Treebank. Our system improves the baseline accuracy from 91.88 % to 93.37 % for English, and in the case of Chinese from 87.39 % to 89.16%. 1.
This paper presents a new method to evaluate machine translation (MT) systems against a parallel treebank. This approach examines specific linguistic phenomena rather than the overall performance of the system. We show that the evaluation accuracy can be increased by using word alignments extracted from a parallel treebank. We compare the performance of our statistical MT system with two other competitive systems with respect to a set of problematic linguistic structures for translation between German and French.
Training data selection is a common method for domain adaptation, the goal of which is to choose a subset of training data that works well for a given test set. It has been shown to be effective for tasks such as machine translation and parsing. In this paper, we propose several entropy-based measures for training data selection and test their effectiveness on two tasks: Chinese word segmentation and part-of-speech tagging. The experimental results on the Chinese Penn Treebank indicate that some of the measures provide a statistically significant improvement over random selection for both tasks.
Parallel treebanks have received increasing attention in the past few years,\nprimarily due to their potential use in statistical machine translation. Creating\nparallel treebanks manually is a time-consuming and expensive task\nand for this reason there is considerable interest in creating treebanks automatically.\nThis task can be solved using standard tools such as parsers and\naligners. However, because parallel treebanks are based on parallel corpora,\nwe are in a special situation where the same meaning is represented\nin two different ways. This thesis is about how we can exploit this information\nto create better parallel treebanks than we can by using standard\ntools....
Development of interpersonal relationships is a fundamental human motivation, and behaviors facilitating social bonding are prized. Some individuals experience enhanced reward from alcohol in social contexts and may be at heightened risk for developing and maintaining problematic drinking. We employed a 3 (group beverage condition) ×2 (genotype) design (N = 422) to test the moderating influence of the dopamine D4 receptor gene (DRD4 VNTR) polymorphism on the effects of alcohol on social bonding. A significant gene x environment interaction showed that carriers of at least one copy of the 7-repeat allele reported higher social bonding in the alcohol, relative to placebo or control conditions, whereas alcohol did not affect ratings of 7-absent allele carriers. Carriers of the 7-repeat allele were especially sensitive to alcohol's effects on social bonding. These data converge with other recent gene-environment interaction findings implicating the DRD4 polymorphism in the development of alcohol use disorders, and results suggest a specific pathway by which social factors may increase risk for problematic drinking among 7-repeat carriers. More generally, our findings highlight the potential utility of employing transdisciplinary methods that integrate genetic methodologies, social psychology, and addiction theory to improve theories of alcohol use and abuse.