Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
In this thesis, we mainly investigate the influence of using unsupervised morphological segmentation as features on the dependency parsing of morphologically rich languages such as Finnish, Estonian, Hungarian, Turkish, Uyghur, and Kazakh. Studying the morphology of these languages is of great importance for the dependency parsing of morphologically rich languages since dependency relations in a sentence of these languages mostly rely on morphemes rather than word order. In order to investigate our research questions, we have conducted a large number of parsing experiments both on MaltParser and UDPipe. We have generated the supervised morphology and the predicted POS tags from UDPipe, and obtained the unsupervised morphological segmentation from Morfessor, and have converted the unsupervised morphological segmentation into features and added them to the UD treebanks of each language. We have also investigated the different ways of converting the unsupervised segmentation into features and studied the result of each method. We have reported the Labeled Attachment Score (LAS) for all of our experimental results. The main finding of this study is that dependency parsing of some languages can be improved simply by providing unsupervised morphology during parsing if there is no manually annotated or supervised morphology available for such languages. After adding unsupervised morphological information with predicted POS tags, we get improvement of 4.9%, 6.0%, 8.7%, 3.3%, 3.7%, and 12.0% on the test set of Turkish, Uyghur, Kazakh, Finnish, Estonian, and Hungarian respectively on MaltParser, and the parsing accuracies have been improved by 2.7%, 4.1%, 8.2%, 2.4%, 1.6%, and 2.6% on the test set of Turkish, Uyghur, Kazakh, Finnish, Estonian, and Hungarian respectively on UDPipe when comparing the results from the models which do not use any morphological information during parsing.
Please join the LAII and University Libraries for a presentation with Richard E. Greenleaf Visiting Library Scholar Jonathan Steuck, a doctoral candidate in Hispanic Linguistics and Language Science at Penn State University. Steuck’s main research interests include code-switching and bilingualism, the prosodic-syntactic interface, and dialectical variation of intonation. Much of his work focuses on Spanish and English in the US. Steuck will discuss his research findings regarding the tendency of speakers of Spanish and English in New Mexico to fluidly alternate between languages in the same conversation (i.e., code-switch). Research suggests that this is a skilled behavior, reflective of a high degree of language proficiency in both languages and the linguistic norms of a speech community. Recent studies have also found that bilinguals may utilize the phonetic features and quantitative patterns present in code-switching to anticipate an upcoming language switch (e.g. Fricke et al., 2016; Tamargo et al. 2016). Previous studies are constrained, however, by the exclusion of factors at the interface of prosody and syntax. This talk will help address this issue by comparing intra-sentential multi-word code-switches (MWCS) with i) English-only and Spanish-only prosodic sentences (see Chafe, 1994) from the same bilingual speakers and ii) Spanish prosodic sentences from more monolingual speakers of an older variety. A consideration of pause expression (i.e. no pause versus (un)filled pause(s)), length measures (e.g. words and seconds), and other prosodic-syntactic factors will illuminate the extent to which MWCS differ from speech where code-switches are absent. Data derive from the New Mexico Spanish-English Bilingual corpus (Torres Cacoullos & Travis, in prep) and the New Mexico-Colorado Spanish Survey (Bills & Vigil 2008). While MWCS are found to have prosodic-syntactic properties that are distinct from Spanish-only and English-only productions, MWCS are no less fluid than speech produced in only one language. Overall, this provides new insight into where code-switches tend to occur in natural discourse and indicates that the previously proposed constraints of code-switching will need to be reexamined in tandem with prosodic junctures. These findings also set the stage for experimental research exploring the prosodic-syntactic cues speakers may utilize to comprehend code-switches. Each year the LAII partners with University Libraries to offer Richard E. Greenleaf Visting Library Scholar awards to support scholars who work with UNM’s nationally-acclaimed Latin American library holdings. The award honors Dr. Richard E. Greenleaf, distinguished scholar of colonial Latin American history, and his extensive career in teaching, research, and service. Event sponsored by Latin American and Iberian Institute, University Libraries.
Listeners rate the speech of boys and girls as young as four years old as sounding gendered: boys are rated as sounding boy-like and girls as girl-like (Perry et al., J. Acoust. Soc. Am. [2001]). Recent research found that the extent to which boys’ speech sounds boy-like is correlated with measures of their gender identity and expression (Li et al., J. Phonetics [2016], Munson et al., J. Acoust. Soc. Am. [2015]). Munson et al. found that boys with a diagnosis of gender identity disorder [GID] were rated as sounding less boy-like than boys without GID. Munson et al.’s experiment used only a small number of girls’ productions as filler items. The current study examined listeners’ ratings of the gender typicality of speech of boys with GID and both boys and girls without GID. Significant differences in sex-typicality ratings were found between the two groups of boys. Boys with GID elicited ratings intermediate to those for boys and girls without GID. However, the differences between boys with and without GID were much smaller than those in Munson et al., suggesting that the sex distribution in the stimulus set can affect ratings of the sex typicality of children’s voices.
A number of firms in Northern Europe and especially in Denmark are owned by private foundations similarly to what would have been the case if the Ford Foundation had owned a majority of the shares in Ford Motor Company. Foundation-owned companies appear to perform surprisingly well in terms of profitability and growth despite lacking governance mechanisms like profit incentives or takeover threats. Given their non-profit ownership, they might be expected to behave more responsibly towards stakeholders such as employees or customers (Hansmann 1980), but so far there has been little empirical evidence to support this hypothesis. This paper presents new research on the reputation and responsibility of foundation-owned companies. In a panel of large Danish companies 2001-2011 we find that foundation-owned firms have better reputations and are regarded as more socially responsible in corporate image ratings. Secondary evidence on labour market behaviour is consistent with these findings. Using matched employer-employee data we show that foundation-owned companies are more stable employers, pay their employees better and keep them for longer. Altogether, the evidence indicates that foundation-ownership is associated with more responsible business behaviour towards employees.
Freedom-of-movement in thought (the degree to which thought is constrained in its variety as opposed to being free to change) can be empirically dissociated from other well-studied dimensions of thought such as its task-unrelatedness in everyday life setting, but has yet to be studied in a controlled experimental environment. While there are several proposed mechanisms by which thought can become constrained (and therefore less freely moving), none have bene explored empirically. The present study set out to uncover which constructs associated with thought’s task-relatedness were also related to its freedom-of-movement and to test potential mechanisms of constraint. Motivation was within-subjects through a variable-value time-sensitive task and administered experience sampling probes asking participants to self-report the level of freely-moving thought, task-unrelated thought, deliberate control, arousal, and valence they were experiencing. Electrodermal activity and pupillometry were used as an index of physiological arousal in addition to self-reports. When participants were more highly motivated they reported having more constrained and more task-related thoughts, and having greater control over their thoughts. Control fully mediated motivation’s impact on freedom-of-movement of thought, but only partially mediated motivation’s impact on task-unrelated thought. Neither self-report or physiological measures of arousal were impacted by the manipulation, but high levels of both task-unrelated and freely-moving thought were associated with high ratings or self-reported arousal, higher pupillary responses to stimuli and smaller average pupil size, with freely-moving thought being uniquely associated with reduced skin conductance. Additionally, high freely-moving thought ratings were uniquely associated with slower responses, while high task-unrelated thought ratings were uniquely associated with low valence ratings. Overall, these findings support and further extend the previously identified dissociation between task-unrelatedness and freedom-of-movement as two separable dimensions of thought. They indicate that a person’s degree of control over their own thoughts is a crucial determinant of the content and especially the dynamics of that thought, but further work needs to be done to explore what nonconscious factors constrain thought movement.
Customers’ opinions on social network platforms are known to influence peer behaviour (Bai, 2011; Eirinaki, Pisal, & Singh, 2012). Customers are also known to be more engaged in sharing their experiences by writing online reviews and recommendations that may be useful to others (Cantallops & Salvi, 2014; Tang & Guo, 2015; Xu & Li, 2016). Actually, user-generated content (UGC) on social network platforms has emerged as an important source for understanding and managing consumers’ expectations, particularly using automated and semi-automated knowledge extraction techniques from text such as text mining and sentiment analysis (Zhang, Zeng, Li, Wang, & Zuo, 2009). This research analyses dimensions of online customer engagement and associated concepts in customers’ reviews through (i) a global sentiment analysis using positive, neutral and negative sentiments and (ii) a topic-sentiment analysis to capture latent topics in online reviews. Furthermore, it examines what influences customers to contribute their online reviews, beyond the features of each focal company or brand. The research methodology is based on a text mining approach, using the MeaningCloud tool. The study focuses on Yelp.com reviews and includes a random sample of 15,000 unique reviews of restaurants, hotels and nightlife entertainment in eleven cities in the USA. An innovative customer engagement dictionary is created, based on previously validated scales using known dimensions of engagement, experience, emotions and brand advocacy, and extended using WordNet 2.1 lexical database. The research findings reveal a high impact of the engagement cognitive processing dimension and hedonic experience on customers’ review endeavour. The study results further indicate that customers seem to be more engaged in positively advocating a company/brand than the contrary. The findings will help social network managers to reinforce their platforms.
Demand-withdraw is an ineffective communication pattern frequently experienced by distressed couples. Therapists often attempt to address this pattern by helping partners understand and regulate the emotions that underlie these behaviors. To date, there is a lack of research focusing on the emotional experiences underlying the demand-withdraw pattern of interaction in couples. Related lines of research focus on emotional arousal and the expression of hard and soft emotions, but this research does not specifically investigate demand-withdraw interactions. The purpose of this study is to identify what emotions underlie demanding behavior in both men and women during marital demand-withdraw conflict interactions. Six couples were chosen from a five-year longitudinal randomized clinical trial that compared Integrative Behavioral Couple Therapy (IBCT) and Traditional Behavioral Couple Therapy (TBCT). Researchers viewed 10-minute pre-treatment problem-solving interactions to observe the demand-withdraw pattern in vivo among couples seeking therapy. The Behavioral Affective Rating Scale (BARS) was used to code the emotions observed during the interactions. The results indicated that the types of emotions varied not only depending on who initiated the problem-solving interaction (e.g., wife topic-husband topic) but also between the different couples, and when comparing gender. Anxiety (#2) and aggression (#4) were in the top four most commonly observed emotions for husbands, while they were two of the least observed emotions for wives. Moreover, frustration and hurt were the two most observed emotions for wives, while they were the least observed emotions for husbands.
This study aimed to address the following questions regarding the emotional experience of Dialectical Behavior Therapy clients with Borderline Personality Disorder: 1) How do positive and negative emotions change in therapy? 2) Does the severity of clients’ symptoms relate to affect? 3) Is affect related to clients’ perceptions of therapeutic alliance? 4) How are clients’ and therapists’ affect related? To test these questions, positive and negative affect ratings were collected from clients (N=77) and therapists (N=25) at the start and end of session. These ratings were tested in relation to alliance and severity ratings using Hierarchical Linear Modeling. Results indicated that clients’ positive affect increased while negative affect decreased from the start to the end of session. This pattern was mirrored over the course of treatment, but only the increase in positive affect was statistically significant over that time period. Severity was significantly related to affect, but in an unexpected direction (higher ratings of emotion dysregulation were associated with slight decreases in negative emotion and emotion lability). Additionally, clients’ positive emotion significantly predicted therapeutic alliance ratings, and therapist positive affect was significantly, positively related to clients’ positive emotion. These results indicate that client affect appears to change in treatment and may be related to severity, alliance, and therapist affect. Further exploration is needed to clarify these complex relationships given the differences between positive and negative affect and the surprising direction of the association between negative affect and emotion dysregulation.
Fibromyalgia syndrome (FMS), a common chronic pain condition, is often incompletely treated by conventional medical therapies. It can cause disability, psychological distress, work-related absenteeism, increased use of healthcare resources, and result in the inability to carry out the tasks of daily living. The purpose of this quantitative, correlational study was to investigate the potential influence of laughter on affect and pain in individuals with FMS. Laughter produces beneficial effects on acute pain and on chronic pain in general and has been found to improve temporary affective states, but there have been no studies testing the effects of laughter on the pain and affect of fibromyalgia patients. Informing this study were the gate control and neuromatrix theories of pain, as well as the dynamic model of affect theory. The research questions addressed whether laughter frequency is associated with affect and or with perceived chronic pain levels in these individuals. Forty-one adult fibromyalgia patients documented all laughter episodes daily and assessed their pain and affective states 3 times per day for 14 days. Hierarchical regressions revealed that increased overall laughter frequency was significantly associated with decreases in overall pain and increases in overall positive affect but was not associated with measures of negative affect. Also, morning laughter frequency was predictive of increased afternoon and evening positive affect ratings, as well as with decreased afternoon pain ratings, but was not significantly associated with evening pain ratings. The knowledge gained from these results may have positive social change implications at the individual level, within those individuals' larger social networks, and within the research and medical communities.
This study aims at exploring new norms as to the textual additions in parentheses (=TAiPs) in the translation of a Quranic text as writer-oriented devices of textuality. Coding for this sort of information could be useful in establishing an impact on any decision-making process on the TL version; such TAiPs can give a translated text of the Quran unity and purpose and distinguish it from a disconnected sequence of sentences. Six small-sized chapters of the Quran were selected as a research sample including a number of four handred forty two (442) TAiPs. Two writer-oriented kinds of textuality were found: cohesivity at the levels of grammar and lexis to be in form of recurrence, reference, substitution, ellipsis and conjunction; and relationality by coherence and intentionality to be in form of reiteration, collocation, connotation, evocation and interpretation. The study is a detailed analysis of such a severely criticized yet officially approved English interpretation of the Quran as the Hilali and Khan Translation (=HKT) against a predetermined set of text-linguistic norms. The strength or weakness of TAiPs as to how they might alleviate or aggravate the TL version is eventually identified for sake of improvement.
espanolLa interpretacion de espanol y castellano ha tenido casi siempre una tendencia sinonimica total a lo largo de su historia. Sin embargo, existen razones historicas y linguisticas para no considerarlo asi. Por un lado, la expansion del Imperio espanol parece ser la causa fundamental para el uso del primero de los terminos mientras que Norma linguistica sevillana serviria como prueba irrefutable de la validez del segundo, de una manera sistemica, y a pesar de que los propios sevillanistas normalmente se habrian manifestado en contra. Se presentan entonces una serie de argumentos contrapuestos que hacen especial hincapie en el marco normativo e historico del idioma y su descripcion geolectal en la actualidad. Con ello se pretende esclarecer diferencias entre los dos conceptos y, de paso, justificar tambien su significado como denominacion compuesta, espanol castellano. NB: Este articulo se basa en el texto “Del castellano al espanol y viceversa”, incluido en mi tesis doctoral (2017) y se redacta de acuerdo con “Opciones linguisticas avanzadas en clase ELE” (2016). EnglishThe interpretation of Spanish and Castilian has been almost meant to be synonymic in absolute terms in history. However, there exist linguistic and historic reasons in order not to state this. On the one hand, the expansion of the Spanish Empire seems to be the paramount cause for the use of the first term whereas Sevillian Linguistic Norm would prove successful in validating the second of them, in a systemic sense, and in spite of the fact that the selfsame sevillianists would have regularly claimed the opposite. Contrasted arguments are presented then, which put special emphasis on the historic and normative framework about this language and its geolectal description at present day. With this, it is intended to clarify some differences between the two concepts and, via that, to also justify their meaning as a compound definition, Castilian Spanish. NB: this article is based on the text “Del castellano al espanol y viceversa”, included in my doctoral thesis (2017) and it has been composed according to “Opciones linguisticas avanzadas en clase ELE” (2016).
Abstract People remember events and materials better when these are congruent with their mood at retrieval; this is known as the mood-congruent memory bias. This effect is largest when the materials are self-referential and this is known as the self-reference effect. We present two word rating studies, to create a list of self-referential valenced words that may be used as stimuli to investigate the influence of valence on cognitive processing in depressive ruminators. Words selected from the Affective Norms for English Words pool were rated by an unselected sample for self-referentiality (Study 1) and validated with ratings provided by depressive ruminators. As hypothesized, depressive ruminators rated negative words as more self-referential than an unselected sample. Using this list, valence differentiated performance between depressive ruminators and healthy controls in a working memory updating task. We thus created a list of self-referential valenced words matched on factors that influence word processing.
The focus of the current study was on idiom comprehension in younger and older adults. Due to inconsistent results in previous studies, it is unclear whether older adults may have problems understanding idioms. For the current study, I used a sentence-to-word matching task presented on an iPad with software that recorded participants’ response time and accuracy. Participants also completed a familiarity task where they rated idioms on how frequently these phrases were encountered. I predicted that older adults would have more difficulty comprehending idioms because of the context in which the idioms were embedded and the timed nature of the task. I also predicted that both age groups would rate the idioms as highly familiar because we purposefully selected these types of expressions. With respect to the sentence-to-word matching task, results showed that although older adults were slower overall, both younger and older adults showed faster response times and greater accuracy for idiomatic targets following idiomatically-biased contexts than for literal targets following literally-biased contexts. With respect to the familiarity ratings task, results showed that both age groups were very familiar with the idioms. These findings suggest that older adults are able to successfully use context to understand familiar ambiguous idioms and that they do not have difficulty comprehending idioms in a cognitively demanding timed task.
The statistical parsing of morphologically rich languages is hindered by the inability of parsers to collect solid statistics because of the large number of word types in such languages. There are however two separate but connected problems, reducing data sparsity of known words and handling rare and unknown words. Methods for tackling one problem may inadvertently negatively impact methods to handle the other. We perform a tightly controlled set of experiments to reduce data sparsity through class-based representations in combination with unknown word signatures with two PCFG-LA parsers that handle rare and unknown words differently on the German TiGer treebank. We demonstrate that methods that have improved results for other languages do not transfer directly to German, and that we can obtain better results using a simplistic model rather than a more generalized model for rare and unknown word handling.
It is no secret that people often use taboo words when speaking about persons and objects in their environment. Taboo words are charged with emotion and have observable impact on the listener as well as the speaker. The purpose of this study was to determine whether taboo words were quantitatively more offensive when used in combination with a proper name versus being used with a non-human object. We found that using taboo words to describe proper names does not cause a significant effect; however, we found that participants rated certain categories of taboo words as more offensive than other categories. In a second experiment, taboo words did affect ratings and memory for proper names and non-human objects.
Shi, Huang, and Lee (2017) obtained state-of-the-art results for English and Chinese dependency parsing by combining dynamic-programming implementations of transition-based dependency parsers with a minimal set of bidirectional LSTM features. However, their results were limited to projective parsing. In this paper, we extend their approach to support non-projectivity by providing the first practical implementation of the MH_4 algorithm, an $O(n^4)$ mildly nonprojective dynamic-programming parser with very high coverage on non-projective treebanks. To make MH_4 compatible with minimal transition-based feature sets, we introduce a transition-based interpretation of it in which parser items are mapped to sequences of transitions. We thus obtain the first implementation of global decoding for non-projective transition-based parsing, and demonstrate empirically that it is more effective than its projective counterpart in parsing a number of highly non-projective languages
Syntactic parsing plays a crucial role in improving the quality of natural language processing tasks. Although there have been several research projects on syntactic parsing in Vietnamese, the parsing quality has been far inferior than those reported in major languages, such as English and Chinese. In this work, we evaluated representative constituency parsing models on a Vietnamese Treebank to look for the most suitable parsing method for Vietnamese. We then combined the advantages of automatic and manual analysis to investigate errors produced by the experimented parsers and find the reasons for them. Our analysis focused on three possible sources of parsing errors, namely limited training data, part-of-speech (POS) tagging errors, and ambiguous constructions. As a result, we found that the last two sources, which frequently appear in Vietnamese text, significantly attributed to the poor performance of Vietnamese parsing.
Detecting lexical entailment plays a fundamental role in a variety of natural language processing tasks and is key to language understanding. Unsupervised methods still play an important role due to the lack of coverage of lexical databases in some domains and languages. Most of the previous approaches were either based on statistical hypothesis of specific entailment relations or tried to encode word relations in low-dimensional vector embeddings. This thesis builds upon one of the few approaches which intrinsically model entailment in a vector space. We then further generalize this model by introducing an alternative, distributional representations for words which harnesses tools from optimal transport to define distance or entailment measures between such representations. We evaluated the models on hypernymy detection where our distributional estimate significantly improves over the underlying model and even outperforms state-of-the-art on some datasets.
In this paper we present the linguistic databases developed during our 8-year lexicographic research on the Modern Greek Standard (MGS) verbal system. Apart from the intermediate databases presented, the main products are (a) a new conjugation system of 385 paradigmatic models, which allows for the automatic generation of all verbal lexical morphemes and monolexical forms (b) a statistically established database of 151,536 distinctive verb-final grapheme sequences which allow for the automatic tagging of all monolexical verbal tokens without the traditional intervention of any built-in lexicon, and (c) a linear Iemmatisation morphophonological rule system accessed on the basis of the distinctive grapheme sequences identified.
An important research field in the area of text mining is text categorization. Most of the real world documents are multi-label in nature. In this paper we have proposed a novel method for automated and effective categorization of multi-label text documents. The proposed method is based on lexical and semantics concepts. Tokens are identified in the text documents using standard IEEE taxonomy. To analyze the semantic relationships between tokens, standard lexical database WordNet is used. The proposed method is tested on a dataset of 150 research articles of computer science domain from IEEE Xplore digital library. It has shown a significantly good performance with an accuracy of 75%.
This paper presents initial experiments in data-driven morphological analysis for Finnish using deep learning methods. Our system uses a character based bidirectional LSTM and pretrained word embeddings to predict a set of morphological analyses for an input word form. We present experiments on morphological analysis for Finnish. We learn to mimic the output of the OMorFi analyzer on the Finnish portion of the Universal Dependency treebank collection. The results of the experiments are encouraging and show that the current approach has potential to serve as an extension to existing rule-based analyzers.
Despite all the impressive advances of recurrent neural networks, sequential data is still in need of better modelling. Truncated backpropagation through time (TBPTT), the learning algorithm most widely used in practice, suffers from the truncation bias, which drastically limits its ability to learn long-term dependencies.The Real Time Recurrent Learning algorithm (RTRL) addresses this issue, but its high computational requirements make it infeasible in practice. The Unbiased Online Recurrent Optimization algorithm (UORO) approximates RTRL with a smaller runtime and memory cost, but with the disadvantage of obtaining noisy gradients that also limit its practical applicability. In this paper we propose the Kronecker Factored RTRL (KF-RTRL) algorithm that uses a Kronecker product decomposition to approximate the gradients for a large class of RNNs. We show that KF-RTRL is an unbiased and memory efficient online learning algorithm. Our theoretical analysis shows that, under reasonable assumptions, the noise introduced by our algorithm is not only stable over time but also asymptotically much smaller than the one of the UORO algorithm. We also confirm these theoretical results experimentally. Further, we show empirically that the KF-RTRL algorithm captures long-term dependencies and almost matches the performance of TBPTT on real world tasks by training Recurrent Highway Networks on a synthetic string memorization task and on the Penn TreeBank task, respectively. These results indicate that RTRL based approaches might be a promising future alternative to TBPTT.
The contents and structure of semantic memory have been the focus of much recent research, with major advances in the development of distributional models, which use word co-occurrence information as a window into the semantics of language. In parallel, connectionist modeling has extended our knowledge of the processes engaged in semantic activation. However, these two lines of investigation have rarely been brought together. Here, we describe a processing model based on distributional semantics in which activation spreads throughout a semantic network, as dictated by the patterns of semantic similarity between words. We show that the activation profile of the network, measured at various time points, can successfully account for response times in lexical and semantic decision tasks, as well as for subjective concreteness and imageability ratings. We also show that the dynamics of the network is predictive of performance in relational semantic tasks, such as similarity/relatedness rating. Our results indicate that bringing together distributional semantic networks and spreading of activation provides a good fit to both automatic lexical processing (as indexed by lexical and semantic decisions) as well as more deliberate processing (as indexed by ratings), above and beyond what has been reported for previous models that take into account only similarity resulting from network structure.
The syntax of newspaper headlines in English displays features which, on a superficial level, set it apart from the norm of Standard English. This, however, is not to say that headlinese stands outside the notion of a linguistic norm: drawing from a corpus of various headlines, and looking at the specific fields of determiners and verb tenses, this article explores how the English of newspaper headlines constitutes a norm in itself which actually builds on the very potentialities of Standard English. Such an extension of the syntactic possibilities of English is designed to fit the pragmatic purpose of headlines, i.e. heighten the relevance of the newspaper article to the reader.
This paper describes our approach to developing the Turkish PropBank by adopting the semantic role-labeling guidelines of the original PropBank and using the translation of the English Penn-TreeBank as a resource. We discuss the semantic annotation process of the PropBank and language-specific cases for Turkish, the tools we have developed for annotation, and quality control for multiuser annotation. In the current phase of the project, more than 9500 sentences are semantically analyzed and predicate-argument information is extracted for 1330 verbs and 1914 verb senses. Our plan is to annotate 17,000 sentences by the end of 2017.
A dependency parser generates both a syntactic structure and a shallow semantic structure of a sentence. It is a fundamental component of natural language processing (NLP) based pipelines, which are critical to facilitate research using the Electronic Health Records (EHR). However, current works mainly apply parsers developed in the general English domain to clinical text. There are no formal evaluations and comparisons of deep learning based dependency parsers in the medical domain. No state-of-the-art dependency parsing performance has been established on clinical text, either. In this study, we investigated the performance of four state-ofthe-art deep learning based dependency parsers, Stanford parser, Bist-parser, dependency_tf parser and jPTDP parser, respectively. Experiments for evaluation are conducted on two datasets: (1) The MiPACQ Treebank and (2) A Treebank of progress notes. Our results showed that the original parsers achieved lower performance in clinical text compared to general English text. After retraining on the clinical Treebank, all parsers obtained better performance. Besides, using word embeddings from Gigaword and MIMICIII yielded comparable performance. Interestingly, the transition-based parsers demonstrated stronger generalizability on different treebanks than the graph-based parsers. Overall, Bist-parser achieved the best performance on MiPACQ (88.95% UAS, 92.69% LS, 86.10% LAS). Stanford parser achieved the best performance on progress notes (84.01% UAS, 89/97% LS, 80.72% LAS).
We propose a novel approach to Vietnamese word segmentation. Our approach is based on the Single Classification Ripple Down Rules methodology (Compton and Jansen, 1990), where rules are stored in an exception structure and new rules are only added to correct segmentation errors given by existing rules. Experimental results on the benchmark Vietnamese treebank show that our approach outperforms previous state-of-the-art approaches JVnSegmenter, vnTokenizer, DongDu and UETsegmenter in terms of both accuracy and performance speed. Our code is open-source and available at: https://github.com/datquocnguyen/RDRsegmenter.
We present ongoing work on data-driven parsing of German and French with Lexicalized Tree Adjoining Grammars. We use a supertagging approach combined with deep learning. We show the challenges of extracting LTAG supertags from the French Treebank, introduce the use of leftand right-sister-adjunction, present a neural architecture for the supertagger, and report experiments of n-best supertagging for French and German.
BACKGROUND: Neuroplastic underpinnings of meditation-induced changes in affective processing are largely unclear. METHODS: We included healthy older participants in an active-controlled experiment. They were involved a meditation training or a control relaxation training of eight weeks. Associations between behavioral and neural morphometric changes induced by the training were examined. RESULTS: The meditation group demonstrated a change in valence perception indexed by more neutral valence ratings of positive and negative affective images. These behavioral changes were associated with synchronous structural enlargements in a prefrontal network involving the ventromedial prefrontal cortex and the inferior frontal sulcus. In addition, these neuroplastic effects were modulated by the enlargement in the inferior frontal junction. In contrast, these prefrontal enlargements were absent in the active control group, which completed a relaxation training. Supported by a path analysis, we propose a model that describes how meditation may induce a series of prefrontal neuroplastic changes related to valence perception. These brain areas showing meditation-induced structural enlargements are reduced in older people with affective dysregulations. CONCLUSION: We demonstrated that a prefrontal network was enlarged after eight weeks of meditation training. Our findings yield translational insights in the endeavor to promote healthy aging by means of meditation.
Answering questions formulated in natural language is a long standing quest in Artificial Intelligence. However, even formulating the problem in precise terms has proven to be too challenging, which lead many researchers to focus on Multiple-Choice Question Answering problems. One particularly interesting type of the latter problem is solving standardized tests such as university entrance exams. The Exame Nacional do Ensino Médio (ENEM) is a High School level exam widely used by Brazilian universities as entrance exam, and the world's second biggest university entrance examination in number of registered candidates. In this work we tackle the problem of answering purely textual multiple-choice questions from the ENEM. We build on a previous solution that formulated the problem as a text information retrieval problem. In particular, we investigate how to enhance these methods by text augmentation using Word Embedding and WordNet, a structured lexical database where words are connected according to some relations like synonymy and hypernymy. We also investigate how to boost performance by building ensembles of weakly correlated solvers. Our approaches obtain accuracies ranging from 26% to 29.3%, outperforming the previous approach.
We propose a novel way to create categorized discourse lexicons for multiple languages. We combine information from the Penn Discourse Treebank with statistical machine translation techniques on the Europarl corpus. Using gender profiling as an application, we evaluate our approach by comparing it with an approach using features from a knowledge-based lexicon and with an Rhetorical structure theory (RST) discourse parser. Our experiments are performed on corpora for three languages (English, Dutch, and German) in two genres (news and blogs). We include a feature analysis in which we look for (in)consistencies of discourse features related to male and female authors between the different experimental settings.
Despite all the impressive advances of recurrent neural networks, sequential data is still in need of better modelling. Truncated backpropagation through time (TBPTT), the learning algorithm most widely used in practice, suffers from the truncation bias, which drastically limits its ability to learn long-term dependencies.The Real-Time Recurrent Learning algorithm (RTRL) addresses this issue, but its high computational requirements make it infeasible in practice. The Unbiased Online Recurrent Optimization algorithm (UORO) approximates RTRL with a smaller runtime and memory cost, but with the disadvantage of obtaining noisy gradients that also limit its practical applicability. In this paper we propose the Kronecker Factored RTRL (KF-RTRL) algorithm that uses a Kronecker product decomposition to approximate the gradients for a large class of RNNs. We show that KF-RTRL is an unbiased and memory efficient online learning algorithm. Our theoretical analysis shows that, under reasonable assumptions, the noise introduced by our algorithm is not only stable over time but also asymptotically much smaller than the one of the UORO algorithm. We also confirm these theoretical results experimentally. Further, we show empirically that the KF-RTRL algorithm captures long-term dependencies and almost matches the performance of TBPTT on real world tasks by training Recurrent Highway Networks on a synthetic string memorization task and on the Penn TreeBank task, respectively. These results indicate that RTRL based approaches might be a promising future alternative to TBPTT.
With the recent growth in attention to transgender people’s experiences, language has become a focused site of both anxiety and contestation. This chapter focuses on three challenges for trans-inclusive language: how to select gendered labels and pronouns, how to make language more gender-neutral when gender isn’t relevant, and how to talk about gender more precisely when it is relevant, such as in discussions of identity, social inequality, physiology, or socialization. In each case, concrete strategies are presented that reflect the way transgender people themselves reshape language to work in more inclusive and affirming ways. These include more direct negotiation regarding appropriate identifying terms, the selection of gender-neutral or gender-inclusive forms, and avoiding the conflation of different aspects of gender. Although transphobia will not be eliminated through linguistic reform alone, changes to linguistic norms that delegitimize trans identities and normalize cis (non-trans) identities are a critical part of trans liberation.
In this paper, we focus on parsing rare and non-trivial constructions, in particular ellipsis. We report on several experiments in enrichment of training data for this specific construction, evaluated on five languages: Czech, English, Finnish, Russian and Slovak. These data enrichment methods draw upon self-training and tri-training, combined with a stratified sampling method mimicking the structural complexity of the original treebank. In addition, using these same methods, we also demonstrate small improvements over the CoNLL-17 parsing shared task winning system for four of the five languages, not only restricted to the elliptical constructions.
This chapter reports on the comprehensive reform of an early childhood teacher education program aiming to substantively address the improvement of educational circumstances of culturally and linguistically diverse children. This project engages sociocultural theory, research, and praxis and expands a funds of knowledge perspective. This reform calls for new contexts for action, new forms of family and community engagement, and new circulation of literacies across households, neighborhoods, and U.S. schools. Overarching design goals for the program are focused on challenging the following: (a) narrow conceptions of language literacy and stories; (b) bounded or contained views of learning contexts; (c) hegemonic cultural and linguistic norms, matters regarding U.S. language and educational policies and power; (d) issues of voice and silencing; (e) quantitative and static views of “resources”; and (f) the limited attention to the agency, identities, and strategic actions of diverse students and their families as they traverse home, community, and institutional contexts.
This paper describes the system of our team Phoenix for participating CoNLL 2018 Shared Task: Multilingual Parsing from Raw Text to Universal Dependencies. Given the annotated gold standard data in CoNLL-U format, we train the tokenizer, tagger and parser separately for each treebank based on an open source pipeline tool UDPipe. Our system reads the plain texts for input, performs the preprocessing steps (tokenization, lemmas, morphology) and finally outputs the syntactic dependencies. For the low-resource languages with no training data, we use cross-lingual techniques to build models with some close languages instead. In the official evaluation, our system achieves the macro-averaged scores of 65.61%, 52.26%, 55.71% for LAS, MLAS and BLEX respectively.
Sentiment analysis involves classifying text into positive, negative and neutral classes according to the emotions expressed in the text. Extensive study has been carried out in performing sentiment analysis using the traditional ‘bag of words’ approach which involves feature selection, where the input is given to classifiers such as Naive Bayes and SVMs. A relatively new approach to sentiment analysis involves using a deep learning model. In this approach, a recently discovered technique called word embedding is used, following which the input is fed into a deep neural network architecture. As sentiment analysis using deep learning is a relatively unexplored domain, we plan to perform in-depth analysis into this field and implement a state of the art model which will achieve optimal accuracy. The proposed methodology will use a hybrid architecture, which consists of CNNs (Convolutional Neural Networks) and RNNs (Recurrent Neural Networks), to implement the deep learning model on the SAR14 and Stanford Sentiment Treebank data sets.
Web queries with question intent manifest a complex syntactic structure and the processing of this structure is important for their interpretation. However, their algorithms rely on resources typically not available outside of big web corporates. We propose a new BiLSTM query parser that: (1) Explicitly accounts for the unique grammar of web queries; and (2) Utilizes named entity (NE) information from a BiLSTM NE tagger, that can be jointly trained with the parser. In order to train our model we annotate the query treebank of When trained on 2500 annotated queries our parser achieves UAS of 83.5% and segmentation F1score of 84.5, substantially outperforming existing state-of-the-art parsers. 1
This paper proposes the model for searching similar collocations in English texts in order to determine semantically connected text fragments for social network data streams analysis. The logical-linguistic model uses semantic and grammatical features of words to obtain a sequence of semantically related to each other text fragments from different actors of a social network. In order to implement the model, we leverage Universal Dependencies parser and Natural Language Toolkit with the lexical database WordNet. Based on the Blog Authorship Corpus, the experiment achieves over 0.92 precision.
Dependency is dynamic manifestation of valency, and dependency distance (DD) is closely related to syntactic structure. By introducing the concept degree in graph theory, we analyze the relationship between dynamic valency (DV) and DD based on the Chinese and English treebanks. Our findings are: (1) the mean dependency distance (MDD) of the Chinese treebank is greater than that of English, while the variance of DV of Chinese is lower than that of English; (2) some values of the variance of DV exist in Chinese but not in English; (3) at a specific sentence length, there is a linear relationship between MDD and the variance of DV in the syntactic structures in English and Chinese. These findings suggest: (1) Chinese may have some unique syntactic dependency structures that are not found in English; (2) high DV may contribute to high MDD, but this effect on MDD may not be stronger than some grammatical factors.
The paper describes the conversion of an LFG treebank of Polish into enhanced Universal Dependencies, and—more generally—identifies the kinds of information lost in translation from LFG to UD. The paper also presents the resulting UD treebank of Polish and compares it to the previous UD treebank of Polish.
This paper demonstrates an end-to-end Chinese discourse parser. We propose a unified framework based on recursive neural network (RvNN) to jointly model the subtasks including elementary discourse unit (EDU) segmentation, tree structure construction, center labeling, and sense labeling. Experimental results show our parser achieves the state-of-the-art performance in the Chinese Discourse Treebank (CDTB) dataset. We release the source code with a pre-trained model for the NLP community. To the best of our knowledge, this is the first open source toolkit for Chinese discourse parsing. The standalone toolkit can be integrated into subsequent applications without the need of external resources such as syntactic parser.
Cross-domain sentiment classifiers aim to predict the polarity, namely the sentiment orientation of target text documents, by reusing a knowledge model learned from a different source domain. Distinct domains are typically heterogeneous in language, so that transfer learning techniques are advisable to support knowledge transfer from source to target. Distributed word representations are able to capture hidden word relationships without supervision, even across domains. Deep neural networks with memory (MemDNN) have recently achieved the state-of-the-art performance in several NLP tasks, including cross-domain sentiment classifica- tion of large-scale data. The contribution of this work is the massive experimentations of novel outstanding MemDNN architectures, such as Gated Recurrent Unit (GRU) and Differentiable Neural Computer (DNC) both in cross-domain and in-domain sentiment classification by using the GloVe word embeddings. As far as we know, only GRU neural networks have been applied in cross-domain sentiment classification. Senti- ment classifiers based on these deep learning architectures are also assessed from the viewpoint of scalability and accuracy by gradually increasing the training set size, and showing also the effect of fine-tuning, an ex- plicit transfer learning mechanism, on cross-domain tasks. This work shows that MemDNN based classifiers improve the state-of-the-art on Amazon Reviews corpus with reference to document-level cross-domain sen- timent classification. On the same corpus, DNC outperforms previous approaches in the analysis of a very large in-domain configuration in both binary and fine-grained document sentiment classification. Finally, DNC achieves accuracy comparable with the state-of-the-art approaches on the Stanford Sentiment Treebank dataset in both binary and fine-grained single-sentence sentiment classification.
We present an architecture based on neural networks to generate natural language from unordered dependency trees. The task is split into the two subproblems of word order prediction and morphology inflection. We test our model gold corpus (the Italian portion of the Universal Dependency treebanks) and an automatically parsed corpus from the Web.
We introduce a novel architecture for dependency parsing: \emph{stack-pointer networks} (\textbf{\textsc{StackPtr}}). Combining pointer networks~\citep{vinyals2015pointer} with an internal stack, the proposed model first reads and encodes the whole sentence, then builds the dependency tree top-down (from root-to-leaf) in a depth-first fashion. The stack tracks the status of the depth-first search and the pointer networks select one child for the word at the top of the stack at each step. The \textsc{StackPtr} parser benefits from the information of the whole sentence and all previously derived subtree structures, and removes the left-to-right restriction in classical transition-based parsers. Yet, the number of steps for building any (including non-projective) parse tree is linear in the length of the sentence just as other transition-based parsers, yielding an efficient decoding algorithm with $O(n^2)$ time complexity. We evaluate our model on 29 treebanks spanning 20 languages and different dependency annotation schemas, and achieve state-of-the-art performance on 21 of them.
The task of nuclearity recognition in Chinese discourse remains challenging due to the demand for more deep semantic information. In this paper, we propose a novel text matching network (TMN) that encodes the discourse units and the paragraphs by combining Bi-LSTM and CNN to capture both global dependency information and local n-gram information. Moreover, it introduces three components of text matching, the Cosine, Bilinear and Single Layer Network, to incorporate various similarities and interactions among the discourse units. Experimental results on the Chinese Discourse TreeBank show that our proposed TMN model significantly outperforms various strong baselines in both micro-F1 and macro-F1.
Increasingly, citizen scientists do work beyond the primary goal of the project (i.e., advanced work) such as writing articles. These activities often take place in discussion boards and have a set of linguistic norms for contributing. For newcomers, learning this language presents a challenge since there are no formal opportunities for them to learn the language and volunteers who join later need to learn more than volunteers who join earlier in a project life-cycle. In this poster, we examine how newcomers language use shifts over the course of two citizen science projects. We find that, although, newcomers joining later might face obstacles, newcomer language associated with advanced work increase over the project's life-cycle. The analysis can help the science team assess whether newcomers on the talk page have either adopted advanced terminologies or they need to have a more formal resource such as tutorial or blog posts.