Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Nous nous intéressons à la performance du langage et aux contraintes causées par la capacité cognitive. Miller (1956) a montré que la mémoire à court-terme est limitée à 7±2 éléments, mais aujourd’hui, cette limite est actualisée à 4 selon Cowan (2001). La première question que nous nous posons est: comment mesurer la complexité syntaxique à partir des données observables considérant ce contrainte de la mémoire à court-terme? Deux autres questions que nous voudrions résoudre sont: quelles sont les relations syntaxiques qui créent vraiment de la complexité? Lorsqu’il s’agit d’une structure syntaxique plus (ou moins) complexe, quels sont les phénomènes qu’on pourrait observer dans nos données?Pour répondre à la première question, je me concentre sur des différents mesures du flux de dépendances (Kahane et al., 2017; Yan & Kahane, 2018). Notre hypothèse est que la mesure du flux à partir de treebanks annotés en syntaxe de dépendance permettraient de vérifier les hypothèses sur la mémoire à court-terme. Les deux autres questions alimentent mon actuel travail de thèse sur les configurations du flux de dépendance, ainsi que les liens entre flux et structure prosodique.
Wikipedia has a strong norm of writing in a "neutral point of view" (NPOV). Articles that violate this norm are tagged, and editors are encouraged to make corrections. But the impact of this tagging system has not been quantitatively measured. Does NPOV tagging help articles to converge to the desired style? Do NPOV corrections encourage editors to adopt this style? We study these questions using a corpus of NPOV-tagged articles and a set of lexicons associated with biased language. An interrupted time series analysis shows that after an article is tagged for NPOV, there is a significant decrease in biased language in the article, as measured by several lexicons. However, for individual editors, NPOV corrections and talk page discussions yield no significant change in the usage of words in most of these lexicons, including Wikipedia's own list of "words to watch." This suggests that NPOV tagging and discussion does improve content, but has less success enculturating editors to the site's linguistic norms.
Facebook is the most popular social networking website. Every post on Facebook actually can imply the user's emotion or opinion. In this paper, we present our analysis on posts associated with left- and right-wing politics in the United States of America. Our dataset contains posts several related Facebook fan pages. We analyze sentiment of posts for the prediction of left- or right-wing politics. We build sentiment features for the prediction and evaluate prediction performance. The results show that F1-score can be as high as 0.95 when TF-IDF is used with a decision tree. Posts generally involve emotional words. We use the lexical databases for sentiment analysis. Our experiment results show that the sentiment analysis is sensitive to some classification algorithms.
PURPOSE: To assess postmatch perceived exertion, feeling, and wellness according to the match outcome (winning, drawing, or losing) in professional soccer players. METHODS: In total, 12 outfield players were followed during 52 official matches where the outcomes (win, draw, or lose) were noted. Following each match, players completed both a 10-point Borg scale modified by Foster and an 11-point Hardy and Rejeski scale rating of perceived feeling. Rating of perceived sleep quality, stress, fatigue, and muscle soreness was collected separately on a 7-point scale the day following each match. RESULTS: Player rating of perceived exertion was higher by a very large magnitude following a loss compared with a draw or a win and higher by a small magnitude after a draw compared with a win. Players felt more pleasure after a win compared with a draw or loss and more displeasure after a loss compared with draw. The players reported a largely and moderately better perceived sleep quality, less stress, and fatigue following a win compared with a draw or a loss and a moderately bad perceived sleep quality, higher stress, and fatigue following a draw compared with a loss. In contrast, only a trivial-small change was observed in perceived muscle soreness between all outcomes. CONCLUSION: Match outcomes moderately to largely affect rating of perceived exertion, feeling, sleep quality, stress, and fatigue, whereas perceived muscle soreness remains high regardless of the match outcome. However, winning a match decreases the strain and improves both pleasure and wellness in professional soccer players.
The paper presents the largest Polish Dependency Bank in Universal Dependencies format -PDBUD -with 22K trees and 352K tokens. PDBUD builds on its previous version, i.e. the Polish UD treebank (PL-SZ), and contains all 8K PL-SZ trees. The PL-SZ trees are checked and possibly corrected in the current edition of PDBUD. Further 14K trees are automatically converted from a new version of Polish Dependency Bank. The PDBUD trees are expanded with the enhanced edges encoding the shared dependents and the shared governors of the coordinated conjuncts and with the semantic roles of some dependents. The conducted evaluation experiments show that PDBUD is large enough for training a high-quality graph-based dependency parser for Polish.
Think aloud protocols are widely applied in user experience studies. In this paper, the effect of two different applications of the Retrospective Think Aloud (RTA) protocol on the number of user-reported usability issues is examined. To this end, 30 users were asked to use the National Cadastre and Mapping Agency web application and complete a set of tasks, such as measuring the land area of a square in their hometown. The order of tasks was randomized per participant. Next, participants were involved in RTA sessions. Each participant was involved in two different RTA modes: (a) the strict guidance, in which the facilitator stayed in the background and prompted participants to keep thinking aloud based on his judgement and experience, and (b) the physiology-supported interventions, in which the facilitator intervened based on real-time monitoring of user's physiological signals. During each session, three participant's physiological signals were recorded: skin conductance, skin temperature and blood volume pulse. Participants were also asked to provide valence-arousal ratings for each self-reported usability issue. Analysis of the collected data showed that participants in the physiology-supported RTA mode reported significantly more usability issues. No significant effect of the RTA mode was found on the va-lence-arousal ratings for the reported usability issues. Participants' physiological signals during the RTA sessions did not also differ significantly between the two modes.
Software industry is turning toward endorsing application (app, in short) development due to the ubiquitous use and interest in this computing pattern. Increasing trend and popularity of mobile apps reveal several issues for the developers to address. Absence of a scientific developmental approach adds further issues to the apps development. There are millions of daily downloads, use and views on mobile apps, which give rise to an interesting phenomenon of apps acceptance by the user community. It is significant to note that users tend to reject or dislike apps that present challenges, owing to the issues in the apps, to them. Therefore, it is imperative to know the different issues that affect mobile ratings. In this paper, we have identified 14 issues by reviewing the relevant literature and collected the data from numerous mobile app stores to find the influence of the identified issues on mobile app ratings. Further, an interpretive structure modeling (ISM) approach is used to categorize the identified issues into four groups — dependent, driving, linkage and autonomous — for better understanding and further analysis. There are two objectives of this research paper: (1) to identify issues in apps which affect ratings and (2) to find out mutual relationship between dominating issues in mobile apps.
Rhapsodie is a 33000-word treebank of spoken French that is annotated for syntax and prosody. It breaks down into 57 five-minute long samples produced by 89 male and female speakers. The discourse profile of each sample is captured by six variables: event structure (dialogue vs. monologue), social context (public vs. private), genre (argumentation, description, narrative, oratory, and procedural), interactivity (interactive, non-interactive, and semi-interactive), channel (broadcasting and face-to-face), and planning type (planned, semi-spontaneous, and spontaneous). The prosodic profile of each sample is captured by two sets of three variables. The first set consists of primary (i.e. structurally objective) variables, namely the mean number per second of pauses (fPauses), conversational overlaps (fOverlap), and gap fillers (fEuh). The second set is based on a model consisting of secondary variables determined a priori by the authors because they are likely to occur in certain discourse genres. They are the mean numbers per second of prosodic prominences (fProm), intonational periods (fIPE), intonation packages (fIPA). Our main research question is whether discourse types in French can be characterized and ultimately predicted by prosodic features. We also address two side questions. First, does the fact that the corpus is relatively small, heterogeneous, and not necessarily balanced affect the representativeness of our results? Second, are the secondary prosodic features representative of discourse genres? We compiled a data table that consists of 57 observations (the corpus samples) and the twelve above listed variables. We visualized the table with RhapVis, a tool we designed on purpose (http://ressources.modyco.fr/sm/RhapVis/), explored it with principal component analysis (http://ressources.modyco.fr/sm/RhapVis/PCA.html), and looked for confirmed tendencies with non-parametric one-way ANOVAs (Kruskal-Wallis H tests). Our exploration shows that argumentative and narrative sequences are prosodically marked, whereas descriptive and procedural sequences are not. A discourse genre is prosodically marked when it is characterized by a high frequency of prosodic features, namely the simultaneous occurrence of overlaps, prominences, and intonation packages. We also claim that a discourse genre is prosodically marked when it is atypical with respect to the other speech genres. This is the case with oratory speech, which is characterized by a high frequency of intonational periods and pauses and is consequently isolated from the other types. These results were partially confirmed by the ANOVAs. Focusing on primary variables, running an ANOVA on fPause showed a significant main effect of Genre (p < 0.05). Further inspection indicates that while the lowest fPause score was found in Narration (M = 0.32; SD = 0.04), the highest score was observed in Oratory (M = 0.42; SD = 0.01). For fOverlap, the main effect of Genre reached the level of significance (p < 0.001), indicating that fOverlap also varies according to Genre. The descriptive data showed that the fOverlap score was the highest for both Argumentation (M = 0.05, SD = 0.04) and Narration (M = 0.02, SD = 0.01). Conversely, no overlap was found in both Oratory and Procedural samples. References Lindqvist, Christina. Corpus transcrits de quelques journaux televises francais, Stockholm, Elanders Gotab, 2001, 289 pages Portele T, Heuft B, Widera C, Wagner P, Wolters M (2000) Perceptual Prominence In: Speech and Signals. Aspects of Speech Synthesis and Automatic Speech Recognition. Festschrift dedicated to Wolfgang Hess on his 60th birthday. Forum Phoneticum, 69. Hektor, Frankfurt a.M.: 97-116. Wagner, P. et al. (2015b), « Disentangling and connecting different perspectives on prosodic prominence », Communication a ICPL, International Conference Prominence in Language, 2015, Cologne, ICPH, 2015
There has been a vast development of personal informatics devices combining sleep monitoring with alarm systems, in order to find an optimal time to awaken a sleeping person in a pleasant way. Most of these systems implement auditory feedback, which is not always pleasant and may disturb other sleepers. We present an adaptive alarm system that detects sleeping cycles and triggers alarm signal during shallow sleep, to minimize sleep inertia. Since tactile sensation is associated with positive valence, vibrotactile stimulation is investigated as a silent alarm to enhance pleasant awakening. Three modulation techniques to render the tactile stimuli for pleasant awakening are considered, namely simultaneous, continuous, and successive stimulation. Two experimental studied are conducted. Experiment 1 studied exogenous attention towards tactile stimulation in a multimodal scenario (involving visual and haptic interactions) with fully awake individuals. Results from the attention task and the subjective valence rating suggest that the vibrotactile stimulation should be based on the continuous modulation, since this not only is very perceivable but also associated with positive attention. Experiment 2 evaluated the user experience with tactile stimulation patterns during sleep. Results confirmed the findings of experiment 1. Continuous modulation was rated highest for pleasant yet arousing sleep-awake transition.
Despite the significant improvement of datadriven dependency parsing systems in recent years, they still achieve a considerably lower performance in parsing spoken language data in comparison to written data. On the example of Spoken Slovenian Treebank, the first spoken data treebank using the UD annotation scheme, we investigate which speechspecific phenomena undermine parsing performance, through a series of training data and treebank modification experiments using two distinct state-of-the-art parsing systems. Our results show that utterance segmentation is the most prominent cause of low parsing performance, both in parsing raw and pre-segmented transcriptions. In addition to shorter utterances, both parsers perform better on normalized transcriptions including basic markers of prosody and excluding disfluencies, discourse markers and fillers. On the other hand, the effects of written training data addition and speech-specific dependency representations largely depend on the parsing system selected.
This paper presents SimpleNLG-NL, an adaptation of the SimpleNLG surface realisation engine for the Dutch language. It describes a novel method for determining and testing the grammatical constructions to be implemented, using target sentences sampled from a treebank.
The multiple state theory of working memory suggests that representations held in working memory are separated into two states: a currently-relevant active representation and accessory memory items held for future use. While the characteristics and consequences of active versus accessory states have been the subject of several investigations, the exact neurocognitive mechanisms that move representations between the two states remain unclear. Of the two competing hypotheses, one suggests that inhibition is applied to keep a representation in an accessory state, while the other suggests that accessory representations simply receive less top-down cortical amplification than active representations, but are not subjected to inhibition. Here we capitalize on the different affective consequences for stimuli whose memory representations are subjected to inhibition (negative ratings) or active enhancement (positive ratings) to test these competing hypotheses. On each trial participants memorized four items and then were cued to focus on a single item within memory. They then completed either a visual search or an affective evaluation task. Search times were slower when a search distractor matched the colour of the active item but not when it matched the colour of the accessory item, replicating findings of a division in working memory whereby only active items guide attention. Also, accessory items were affectively devalued compared to baseline and active memory items. This finding of devaluation supports the hypothesis that inhibition is used to keep representations in an accessory state, and adds to past findings that similar mechanisms of attention and emotion govern prioritization in working memory and the prioritization of external stimuli. Meeting abstract presented at VSS 2018
Recurrent neural networks (RNNs) provide excellent performance on applications with sequential data such as speech recognition. On-chip implementation of RNNs is difficult due to the significantly large number of parameters and computations. In this work, we first present a training method for LSTM model for language modeling on Penn Treebank dataset with binary weights and multi-bit activations and then map it onto a fully parallel RRAM array architecture ("XNOR-RRAM"). An energy-efficient XNOR-RRAM array based system for LSTM RNN is implemented and benchmarked on Penn Treebank dataset. Our results show that 4-bit activation precision can provide a near-optimal perplexity of 115.3 with an estimated energy-efficiency of ~27 TOPS/W.
We introduce a class of convolutional neural networks (CNNs) that utilize recurrent neural networks (RNNs) as convolution filters. A convolution filter is typically implemented as a linear affine transformation followed by a nonlinear function, which fails to account for language compositionality. As a result, it limits the use of high-order filters that are often warranted for natural language processing tasks. In this work, we model convolution filters with RNNs that naturally capture compositionality and long-term dependencies in language. We show that simple CNN architectures equipped with recurrent neural filters (RNFs) achieve results that are on par with the best published ones on the Stanford Sentiment Treebank and two answer sentence selection datasets. 1
Difficulties in the regulation of emotion are hypothesized to play a key role in the development and maintenance of posttraumatic stress disorder (PTSD). The current study used functional magnetic resonance imaging (fMRI) to assess neural activity during task preparation and image presentation during different emotion regulation strategies, cognitive reappraisal and expressive suppression, in PTSD. Patients with combat-related PTSD (n = 18) and combat-exposed controls (n = 27) were instructed to feel, reappraise or suppress their emotional response prior to viewing combat-related images during fMRI, while also providing arousal ratings. In the reappraise condition, patients showed lower medial prefrontal neural activity during task preparation and higher prefrontal neural activity during image presentation, compared with controls. No difference in neural activity was observed between the groups during the feel or suppress conditions, although patients rated images as more arousing than controls across all three conditions. By distinguishing between preparation and active regulation, and between reappraisal and suppression, the current findings reveal greater complexity regarding the dynamics of emotion regulation in PTSD and have implications for our understanding of the etiology and treatment of PTSD.
English in Singapore has always presented a balancing act for its founders. The colonial era saw a distinct role for English, i.e. to produce English-speaking officers for the British administration, while modern Singapore sees English being used as both a national and international lingua franca and as a major language that connects the island city-state to the world. ‘English-knowing bilingualism’ has gained ascendancy in Singapore and may become a core competency for the 21st-century world with the rise in status of English as a global language. However, the path to English-knowing bilingualism in the pluri-lingual and heterogeneous country was often marked by paradoxical debates surrounding the issues of language maintenance and shift, identity and the transmission of values, equity and meritocracy, as well as balancing between local versus global linguistic norms and standards. This paper focuses on the continuing debates, from the past to the present, as new challenges arise and argues how a new balance has to be achieved in the language strategy, policy and management for future-readiness in Singapore.
Recurrent neural networks (RNNs) are powerful models of sequential data. They have been successfully used in domains such as text and speech. However, RNNs are susceptible to overfitting; regularization is important. In this paper we develop Noisin, a new method for regularizing RNNs. Noisin injects random noise into the hidden states of the RNN and then maximizes the corresponding marginal likelihood of the data. We show how Noisin applies to any RNN and we study many different types of noise. Noisin is unbiased--it preserves the underlying RNN on average. We characterize how Noisin regularizes its RNN both theoretically and empirically. On language modeling benchmarks, Noisin improves over dropout by as much as 12.2% on the Penn Treebank and 9.4% on the Wikitext-2 dataset. We also compared the state-of-the-art language model of Yang et al. 2017, both with and without Noisin. On the Penn Treebank, the method with Noisin more quickly reaches state-of-the-art performance.
Anger is considered a unique high-arousal and approach-related negative emotion. The influence of individual differences in trait anger on the processing of visual stimuli is relevant to questions about emotional processing and remains to be explored. Using functional magnetic resonance imaging (fMRI), we explored the neural responses to standardized images, selected based on valence and arousal ratings in a group of men with high trait anger compared to those with normative to low anger scores (controls). Results show increased activation in the left-lateralized ventral fronto-parietal attention network to unpleasant images by individuals with high trait anger. There was also a group by arousal interaction in the left thalamus/pulvinar such that individuals with high trait anger had increased pulvinar activation to the high-arousal (versus low arousal) unpleasant images as compared to controls. Thus, individual differences in trait anger in men are associated with brain regions subserving executive attentional and sensory integration during the processing of unpleasant emotional stimuli, particularly to high arousal images.
Abstract. Affective science calls for methods to induce mood in an engaging and ecologically valid way. We present a method employing a naturally occurring scenario that fits these criteria: a job interview. Participants got positive or negative feedback from a fictive expert to induce positive or negative mood. After mood induction, we assessed participants’ decision making behavior in the so-called information sampling task (IST). Results show that our mood induction successfully changed valence, dominance, and state self-esteem ratings, while there were no differences in arousal ratings. Decision making in the IST was not influenced by the induced mood. Effect sizes of mood induction were equally high for positive and negative mood concerning valence ratings (d =.8) with participants scoring high on self-control showing smaller mood induction effects. We conclude that our mood induction technique is an effective and natural way to induce mood in the laboratory, meeting current criteria of affective science.
The present paper describes and illustrates the main naming strategies attested in a lexical database of 1233 Kakataibo names of plant and animals. Seven naming strategies are proposed for Kakataibo ethnobiological nomenclature: coining, morphological derivation, borrowing, ethnobiological polysemy, compounding and grammatical nominalization (the latter two being exclusively associated with lexically complex forms). Kakataibo ethnobiological terminology overally follows the general word-formation patterns available in the language, but it will be argued that some types of compounds and grammatical nominalizations found in the database are constraint to names of plants and animal. Indeed, one particular type of lexicalized grammatical nominalization seems to be cross-linguistically unusual.
The contrast between the contextual and general meaning of a word serves as an important clue for detecting its metaphoricity. In this paper, we present a deep neural architecture for metaphor detection which exploits this contrast. Additionally, we also use cost-sensitive learning by re-weighting examples, and baseline features like concreteness ratings, POS and WordNet-based features. The best performing system of ours achieves an overall F1 score of 0.570 on All POS category and 0.605 on the Verbs category at the Metaphor Shared Task 2018.
بنك المشجّرات محلّل حاسوبيّ للظّواهر التّركيبيّة في اللّغة العربيّة، استثمر مبادئ نظريّة التّحكّم والرّبط التّوليديّة، وحوسباتها وتصوّراتها للنّحو الكلّيّ، غايته في ذلك بناء نظام حوسبيّ آليّ، يحاكي في اشتغاله النّظام الحوسبيّ اللّغويّ الطّبيعيّ. وقد حقّق بنك المشجّرات نتائج مهمّة في هذا الشّأن، تتمثّل في بلوغه الانتظام والتّناسق في معالجة الأبنية الإعرابيّة، لكنّ العمل لم يخل من هنات، أهمّها عدم اتّسام السّيرورة الاشتقاقيّة بالخاصّيّة التّكراريّة المميّزة للّغة البشريّة، وخرق حوسبة النّقل للقيود الجزبريّة التي أقرّتها النّظريّة اللّسانيّة، وهو ما يجعلنا نشكّك في كفايته الوصفيّة لسانيّا.
Typing is a ubiquitous daily action for many individuals; yet, research on how these actions have changed our perception of language is limited. One such influence, deemed the QWERTY effect, is an increase in valence ratings for words typed more with the right hand on a traditional keyboard (Jasmin &amp; Casasanto, 2012). Although this finding is intuitively appealing given both right-handed dominance and the smaller number of letters typed with the right hand, extension and replication of the right-side advantage is warranted. The present paper reexamined the QWERTY effect expanding to other embodied cognition variables (Barsalou, 1999). First, we found that the right-side advantage is replicable to new valence stimuli. Further, when examining expertise, right-side advantage interacted with typing speed and typeability (i.e., alternating hand key presses or finger switches) portraying that both skill and our procedural actions play a role in judgment of valence on words.
We unify recent neural approaches to one-shot learning with older ideas of associative memory in a model for metalearning. Our model learns jointly to represent data and to bind class labels to representations in a single shot. It builds representations via slow weights, learned across tasks through SGD, while fast weights constructed by a Hebbian learning rule implement one-shot binding for each new task. On the Omniglot, Mini-ImageNet, and Penn Treebank one-shot learning benchmarks, our model achieves state-of-the-art results.
We propose Efficient Neural Architecture Search (ENAS), a faster and less expensive approach to automated model design than previous methods. In ENAS, a controller learns to discover neural network architectures by searching for an optimal path within a larger model. The controller is trained with policy gradient to select a path that maximizes the expected reward on the validation set. Meanwhile the model corresponding to the selected path is trained to minimize the cross entropy loss. On the Penn Treebank dataset, ENAS can discover a novel architecture thats achieves a test perplexity of 57.8, which is state-of-the-art among automatic model design methods on Penn Treebank. On the CIFAR-10 dataset, ENAS can design novel architectures that achieve a test error of 2.89%, close to the 2.65% achieved by standard NAS (Zoph et al., 2017). Most importantly, our experiments show that ENAS is more than 10x faster and 100x less resource-demanding than NAS.
We propose a novel neural network model for joint part-of-speech (POS) tagging and dependency parsing. Our model extends the well-known BIST graph-based dependency parser (Kiperwasser and Goldberg, 2016) by incorporating a BiLSTM-based tagging component to produce automatically predicted POS tags for the parser. On the benchmark English Penn treebank, our model obtains strong UAS and LAS scores at 94.51% and 92.87%, respectively, producing 1.5+% absolute improvements to the BIST graph-based parser, and also obtaining a state-of-the-art POS tagging accuracy at 97.97%. Furthermore, experimental results on parsing 61 "big" Universal Dependencies treebanks from raw texts show that our model outperforms the baseline UDPipe (Straka and Straková, 2017) with 0.8% higher average POS tagging score and 3.6% higher average LAS score. In addition, with our model, we also obtain state-of-the-art downstream task scores for biomedical event extraction and opinion analysis applications. Our code is available together with all pre-trained models at: https://github.com/datquocnguyen/jPTDP
Failing to recognize one's mirror image can signal an abnormality in one's sense of self. In dissociative identity disorder (DID), individuals often report that their mirror image can feel unfamiliar or distorted. They also experience some of their own thoughts, emotions, and bodily sensations as if they are nonautobiographical and sometimes as if instead, they belong to someone else. To assess these experiences, we designed a novel backwards masking paradigm in which participants were covertly shown their own face, masked by a stranger's face. Participants rated feelings of familiarity associated with the strangers' faces. 21 control participants without trauma-generated dissociation rated masks, which were covertly preceded by their own face, as more familiar compared to masks preceded by a stranger's face. In contrast, across two samples, 28 individuals with DID and similar clinical presentations (DSM-IV Dissociative Disorder Not Otherwise Specified type 1) did not show increased familiarity ratings to their own masked face. However, their familiarity ratings interacted with self-reported identity state integration. Individuals with higher levels of identity state integration had response patterns similar to control participants. These data provide empirical evidence of aberrant self-referential processing in DID/DDNOS and suggest this is restored with identity state integration.
Are gender differences in emotion culturally universal? To answer this question, the current study compared gender differences in emotional arousal (intensity) ratings for negative and positive pictures from the International Affective Picture System (IAPS) across cultures (Chinese vs. German culture) and age (younger vs. older adults). The raters were 53 younger Germans (24 women), 53 older Germans (28 women), 300 younger Chinese (176 women), and 126 older Chinese (86 women). The results showed that gender differences in arousal ratings were moderated by culture and age: Chinese women reported higher arousal for both negative and positive pictures compared with Chinese men; German women reported higher arousal for negative pictures, but lower arousal for positive pictures compared with German men. Moreover, the gender differences were larger for older than younger adults in the Chinese sample but smaller for older than younger adults in the German sample. The results indicated that gender differences in self-report emotional intensity induced by pictorial stimuli were more consistent with gender norms and stereotypes (i.e., women being more emotional than men) in the Chinese sample, compared with the German sample, and that gender differences were not constant across age groups. The study revealed that gender differences in emotion are neither constant nor universal, and it highlighted the importance of taking culture and age into account.
This paper describes a transduction language suitable for natural language treebank transformations and motivates its application to tasks that have been used and described in the literature. The language, which is the basis for a tree transduction tool allows for clean, precise and concise description of what has been very confusingly, ambiguously, and incompletely textually described in the literature also allowing easy non-hard-coded implementation. We also aim at getting feedback from the NLP community to eventually converge to a de facto standard for such transduction language.
Summary: Standard Catalan is based on the Central dialect and, specifically, on Barcelona speech. However, there are standard variants for all dialects, except for the Northern one. Furthermore, the sociolinguistic situation in Northern Catalan differs from that in other Catalan-speaking territories in that the language has almost disappeared. Some cultural activists are still trying to recover the Catalan language by using it in as many situations as possible. The objective of this article is to analyse the variety of Catalan – standard or dialectal forms – used in literature, the media, and education and what this usage demonstrates about Northern Catalans’ attitudes towards their own language. Keywords: Northern Catalan, standard Catalan, sociolinguistics, language attitudes
Imageability ratings for 3,000 monosyllabic words"<br>reference_to_cite: 'Cortese, M. J., & Fugett, A. (2004). Imageability ratings for 3,000 monosyllabic words. Behavior Research Methods, Instruments, & Computers (2004), Vol. 36, Issue 3, 384-387 [on-line]'<br>url_download: "http://psychonomic.org/archive"
We present the Uppsala system for the CoNLL 2018 Shared Task on universal\ndependency parsing. Our system is a pipeline consisting of three components:\nthe first performs joint word and sentence segmentation; the second predicts\npart-of- speech tags and morphological features; the third predicts dependency\ntrees from words and tags. Instead of training a single parsing model for each\ntreebank, we trained models with multiple treebanks for one language or closely\nrelated languages, greatly reducing the number of models. On the official test\nrun, we ranked 7th of 27 teams for the LAS and MLAS metrics. Our system\nobtained the best scores overall for word segmentation, universal POS tagging,\nand morphological features.\n
Au début de cette thèse, aucun corpus annoté syntaxiquement (treebank) n’était disponible pour le serbe. Or, les treebanks annotés manuellement sont une condition sine qua non du développement (entraînement et évaluation) d’outils statistiques dédiés à l’annotation syntaxique automatique (parsers). L’existence des parsers performants permet à son tour l’annotation syntaxique de corpus plus larges, qui peuvent ensuite alimenter des recherches en linguistique théorique. De fait, l’absence de ces ressources pour le serbe freine le développement des recherches sur cette langue dans ces deux directions, et plus généralement les efforts visant l’informatisation et la valorisation du serbe. Afin de combler cette lacune, nous avons constitué un ensemble de ressources pour le traitement automatique du serbe. Il s’agit en premier lieu du treebank ParCoTrain-Synt, qui contient 101 000 tokens annotés en morphosyntaxe, en lemmes et en syntaxe de dépendances. Nous avons également confectionné le lexique ParCoLex, doté de 7 millions d’entrées provenant de 157 000 lemmes différents. En exploitant ces deux ressources, nous avons développé des modèles pour le parsing, pour l’étiquetage et pour la lemmatisation.Toutes les ressources citées sont librement diffusées à l’adresse suivante: https://github.com/aleksandra-miletic/serbian-nlp-resources. Les ressources constituées ont également été exploitées dans le cadre de deux études linguistiques, montrant ainsi que le corpus ParCoTrain-Synt ouvre la porte aux études empiriques basées sur des analyses quantitatives dans le domaine de la linguistique serbe.
Deep neural networks (DNNs) have achieved impressive predictive performance due to their ability to learn complex, non-linear relationships between variables. However, the inability to effectively visualize these relationships has led to DNNs being characterized as black boxes and consequently limited their applications. To ameliorate this problem, we introduce the use of hierarchical interpretations to explain DNN predictions through our proposed method, agglomerative contextual decomposition (ACD). Given a prediction from a trained DNN, ACD produces a hierarchical clustering of the input features, along with the contribution of each cluster to the final prediction. This hierarchy is optimized to identify clusters of features that the DNN learned are predictive. Using examples from Stanford Sentiment Treebank and ImageNet, we show that ACD is effective at diagnosing incorrect predictions and identifying dataset bias. Through human experiments, we demonstrate that ACD enables users both to identify the more accurate of two DNNs and to better trust a DNN's outputs. We also find that ACD's hierarchy is largely robust to adversarial perturbations, implying that it captures fundamental aspects of the input and ignores spurious noise.
The main aim of the PhD study A computational syntactic analysis of Setswana(AS Berg, May 2018) is the computational syntactic analysis of the Setswana simple sentence, using Lexical Functional Grammar (LFG) as framework and XLE as the associated grammar development platform. The computational grammar is tested with a hand-crafted test suite constructed with 828 test items and consists of Setswana phrases and simple sentences. The analyses of these test items are stored in the the following available formats:.SExp for trees,.lfg for functional structures and.pl for trees and functional structures in prolog. The treebank consists of 2903 trees and functional structures for the 828 phrases and sentences.
Unsupervised learning of syntactic structure is typically performed using generative models with discrete latent variables and multinomial parameters. In most cases, these models have not leveraged continuous word representations. In this work, we propose a novel generative model that jointly learns discrete syntactic structure and continuous word representations in an unsupervised fashion by cascading an invertible neural network with a structured generative prior. We show that the invertibility condition allows for efficient exact inference and marginal likelihood computation in our model so long as the prior is well-behaved. In experiments we instantiate our approach with both Markov and tree-structured priors, evaluating on two tasks: part-of-speech (POS) induction, and unsupervised dependency parsing without gold POS annotation. On the Penn Treebank, our Markov-structured model surpasses state-of-the-art results on POS induction. Similarly, we find that our tree-structured model achieves state-of-the-art performance on unsupervised dependency parsing for the difficult training condition where neither gold POS annotation nor punctuation-based constraints are available. 1 e i N ( z i, z i
Cognitive variation due to language and culture has been shown in a range of domains, including visual perception,emotions, theory of mind, economic strategies, decision making, and categorization. While such patterns are robust,individuals within a given culture are affected by these cultural patterns differentially. One possible cause for theseindividual differences is personality (e.g. extroversion or agreeableness). The personality traits of individuals will affecthow they interact with and adopt cultural patterns. To explore this possibility, we perform analyses on online data fromindividuals with self-identified Myers-Briggs personality types (a popularized personality measure that is widely self-reported in social media). In particular, we examine how personality type predicts the rate at which individuals adopt novellexical items and conform to the linguistic norms of their surrounding community. The results make explicit predictionsabout which individuals will be more affect by cultural and linguistic patterns.
Film clips are proven to be one of the most efficient techniques in emotional induction. However, there is scant literature on the effect of this procedure in older adults and, specifically, the effect of using different positive stimuli. Thus, the aim of the present study was to examine emotional differences between young and older adults and to know how a set of film clips works as mood induction procedure in older adults, especially, when trying to elicit attachment-related emotions. To this end, we use this procedure to analyze differences in subjective emotional response between young and older adults. A sample of 57 older adults and 83 young adults watched a film set previously validated in young population. Their responses were studied in an individual laboratory session to elicit 6 target emotions (disgust, fear, sadness, anger, amusement and tenderness) and neutral state. Self-reported emotional experience was measured using the Self-Assessment Manikin (SAM). Our results show that film clips are capable of evoking positive and negative emotions in older adults. Furthermore, older adults experienced more intensely negative emotions than young adults, especially in response to disgust and fear clips. They also reported higher arousal than young adults, especially in the case of sadness, anger and tenderness clips. Nevertheless, the older adults recovered more easily from the effects of the emotion induction. The young adults reported higher arousal ratings than older adults in response to amusement film clips. On the other hand, this study reflects the importance of controlling the baseline state to study the real strength of mood induction. Overall, current data suggests significant differences occur in emotional response in adult age and that film clips are an effective tool for studying positive and negative emotions in aging research.
<h3>Introduction</h3><br> DEFT Spanish Treebank was developed by the Linguistic Data Consortium (LDC) and the <a href="http://clic.ub.edu/">Language and Computation Center (CLiC), University of Barcelona</a>. It contains treebank annotation of international Spanish newswire text and Latin American Spanish discussion forum data created for the DARPA Deep Exploration and Filtering of Text (DEFT) program. <br> DEFT aimed to improve state-of-the-art capabilities in automated deep natural language processing with a particular focus on technologies dealing with inference, casual relationships and anomaly detection across several languages. DEFT Spanish Treebank supported the program's goal of deep natural language understanding. <br> <h3>Data</h3><br> Newswire source files were selected from Spanish Gigaword Third Edition (<a href="../../../ldc2011t12">LDC2011T12</a>) and were manually sentence-segmented for DEFT. Discussion forum source files were selected from Spanish discussion forum source data collected by LDC, consisting of continuous multi-posts of 100-1000 words. <br> This release contains 114 files (54,394 tokens) of newswire data and 60 files (55,307 tokens) of discussion forum data all of which were annotated with constituents and syntactic functions. The annotation guidelines for DEFT Spanish Treebank are included in the documentation accompanying this release. <br> Source documents are presented as plain text files with one sentence unit per line. Treebank annotation files are in xml. <br> <h3>Samples</h3><br> Please view this <a href="desc/addenda/LDC2018T01.txt">source sample</a> and <a href="desc/addenda/LDC2018T01.xml">treebank sample</a>. <br> <h3>Updates</h3><br> None at this time. <br> <h3>Acknowledgement</h3><br> This material is based on research sponsored by Air Force Research Laboratory and Defense Advance Research Projects Agency under agreement number FA8750-13-2-0045. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of Air Force Research Laboratory and Defense Advanced Research Projects Agency or the U.S. Government. </br> Portions © 1994-2001, 2004-2009 The Associated Press, © 2002, 2005, 2007, 2009-2010 Xinhua News Agency, © 2006, 2009, 2011, 2018 Trustees of the University of Pennsylvania
Artificially created treebank of elliptical constructions (gapping), in the annotation style of Universal Dependencies. Data taken from UD 2.1 release, and from large web corpora parsed by two parsers. Input data are filtered, sentences are identified where gapping could be applied, then those sentences are transformed, one or more words are omitted, resulting in a sentence with gapping. Details in Droganova et al.: Parse Me if You Can: Artificial Treebanks for Parsing Experiments on Elliptical Constructions, LREC 2018, Miyazaki, Japan.
We demonstrate that replacing an LSTM encoder with a self-attentive architecture can lead to improvements to a state-of-the-art discriminative constituency parser. The use of attention makes explicit the manner in which information is propagated between different locations in the sentence, which we use to both analyze our model and propose potential improvements. For example, we find that separating positional and content information in the encoder can lead to improved parsing accuracy. Additionally, we evaluate different approaches for lexical representation. Our parser achieves new state-of-the-art results for single models trained on the Penn Treebank: 93.55 F1 without the use of any external data, and 95.13 F1 when using pre-trained word representations. Our parser also outperforms the previous best-published accuracy figures on 8 of the 9 languages in the SPMRL dataset.
BACKGROUND: Life events (LEs) are associated with future physical and mental health. They are crucial for understanding the pathways to mental disorders as well as the interactions with biological parameters. However, deeper insight is needed into the complex interplay between the type of LE, its subjective evaluation and accompanying factors such as social support. The "Stralsund Life Event List" (SEL) was developed to facilitate this research. METHODS: The SEL is a standardized interview that assesses the time of occurrence and frequency of 81 LEs, their subjective emotional valence, the perceived social support during the LE experience and the impact of past LEs on present life. Data from 2265 subjects from the general population-based cohort study "Study of Health in Pomerania" (SHIP) were analysed. Based on the mean emotional valence ratings of the whole sample, LEs were categorized as "positive" or "negative". For verification, the SEL was related to lifetime major depressive disorder (MDD; Munich Composite International Diagnostic Interview), childhood trauma (Childhood Trauma Questionnaire), resilience (Resilience Scale) and subjective health (SF-12 Health Survey). RESULTS: The report of lifetime MDD was associated with more negative emotional valence ratings of negative LEs (OR = 2.96, p < 0.0001). Negative LEs (b = 0.071, p < 0.0001, β = 0.25) and more negative emotional valence ratings of positive LEs (b = 3.74, p < 0.0001, β = 0.11) were positively associated with childhood trauma. In contrast, more positive emotional valence ratings of positive LEs were associated with higher resilience (b = - 7.05, p < 0.0001, β = 0.13), and a lower present impact of past negative LEs was associated with better subjective health (b = 2.79, p = 0.001, β = 0.05). The internal consistency of the generated scores varied considerably, but the mean value was acceptable (averaged Cronbach's alpha > 0.75). CONCLUSIONS: The SEL is a valid instrument that enables the analysis of the number and frequency of LEs, their emotional valence, perceived social support and current impact on life on a global score and on an individual item level. Thus, we can recommend its use in research settings that require the assessment and analysis of the relationship between the occurrence and subjective evaluation of LEs as well as the complex balance between distressing and stabilizing life experiences.