Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Roles are one of the most important concepts in understanding human sociocognitive behavior. During group interactions, members take on different roles within the discussion. Roles have distinct patterns of behavioral engagement (i.e., active or passive, leading or following), contribution characteristics (i.e., providing new information or echoing given material), and social orientation (i.e., individual or group). Different combinations of roles can produce characteristically different group outcomes, and thus can be either less or more productive with regard to collective goals. In online collaborative-learning environments, this can lead to better or worse learning outcomes for the individual participants. In this study, we propose and validate a novel approach for detecting emergent roles from participants’ contributions and patterns of interaction. Specifically, we developed a group communication analysis (GCA) by combining automated computational linguistic techniques with analyses of the sequential interactions of online group communication. GCA was applied to three large collaborative interaction datasets (participant N = 2,429, group N = 3,598). Cluster analyses and linear mixed-effects modeling were used to assess the validity of the GCA approach and the influence of learner roles on student and group performance. The results indicated that participants’ patterns of linguistic coordination and cohesion are representative of the roles that individuals play in collaborative discussions. More broadly, GCA provides a framework for researchers to explore the micro intra- and interpersonal patterns associated with participants’ roles and the sociocognitive processes related to successful collaboration.
When a shift in writing style is noticed in a document, doubts arise about its originality. Based on this clue to plagiarism, the intrinsic approach to plagiarism detection identifies the stolen passages by analysing the writing style of the suspicious document without comparing it to textual resources that may serve as sources for the plagiarist. Character n-grams are recognised as a successful approach to modelling text for writing style analysis. Although prior studies have investigated the best practice of using character n-grams in authorship attribution and other problems, there is still a need for such investigations in the context of intrinsic plagiarism detection. Moreover, it has been assumed in previous works that the ways of using character n-grams in authorship attribution remain the same for intrinsic plagiarism detection. In this paper, we study the effect of character n-grams frequency and length on the performance of intrinsic plagiarism detection. Our experiments utilise two state-of-the-art methods and five large document collections of PAN labs written in English and Arabic. We demonstrate empirically that the low- and the high-frequency n-grams are not equally relevant for intrinsic plagiarism detection, but their performance depends on the way they are exploited.
Combinatory Categorial Grammars provide a transparent interface between surface syntax and underlying semantic representation. Discourse Representation Theory allows the handling of meaning across sentence boundaries. Based on the foundations of these two theories along with the work of Johan Bos on the Boxer framework for English language, we propose an approach to the task of semantic parsing with Discourse Representation Structure for the French language. By giving an example of discourse analysis on French sentences and experimenting on 4,525 sentences taken from the French Treebank corpus, we demonstrate and evaluate the outcomes of our framework.
This paper investigates the historical (1850s–2000s) evolution of semantics in the English language using contemporaneous, decade-specific computational estimates of word concreteness. Study 1 describes the computational method of generating time-locked estimates of concreteness based on the Corpus of Historic American English, and makes available the computed scores for 25,000 English words over 15 decades. We also report several tests of reliability and validity, demonstrating that our historical concreteness scores have high levels of both. Study 2 uses concreteness scores to revisit findings of studies that use a static set of contemporary human concreteness norms to examine historical trends of semantic change. Specifically, we observed (contra Hills & Adelman, (Cognition, 143, 87–92 2015)) that distinct word types of the English language become increasingly more concrete over time and (in line with Hills & Adelman, (Cognition, 143, 87–92 2015) & Hills, Adelman & Noguchi, (The Quarterly Journal of Experimental Psychology, 70(8), 1603–1619 2016)) that relatively concrete words tend to be used more often than abstract ones. We discuss both contrastive and corroborative claims in light of recent work on semantic evolution and argue for the use of time-locked computed estimates over static human norms when examining diachronic linguistic phenomena.
espanolEn 1969, Philippe Jaccottet tradujo para la revista Aquila el poema Dunja de Giuseppe Ungaretti. El estudio tiene como objetivo destacar la singularidad de esta traduccion, arrojar luz sobre la genetica de su escritura: es decir, penetrar el lado invisible del proceso de traduccion en el que participan el autor y el traductor. A partir de la aclaracion terminologica de la nocion de transparencia, se aplica para dilucidar las estrategias del traductor, destacando la dialectica de la hermetica y de la hermeneutica. La yuxtaposicion de las versiones del poema, elaboradas por Jaccottet durante el trabajo en curso, proporciona informacion sobre los pasos de la traduccion transparente, comenzando por el trabajo sobre el lexico que se refiere a la busqueda de la equivalencia. Revela, ademas, las indecisiones y las dudas de un traductor advertido, consciente de la imposibilidad de comprender el enunciado poetico, descuidando sus aspectos emocional, imaginal y prosodico. Para llegar a la conclusion de que se trata de una traduccion “relevante”, resultado de un trabajo meticuloso sobre las potencialidades de la lengua meta, que no duda en transgredir sus normas gramaticales para preservar la extraneza del texto fuente. EnglishIn 1969 Philippe Jaccottet translated for the magazine Aquila the poem Dunja by Giuseppe Ungaretti. The study aims to highlight the singularity of this translation, shedding light on the genetics of its writing: that is to say to penetrate the invisible side of translating in which both the author and the translator participate. Starting from the terminological clarification of the notion of transparency, it applies itself to elucidate the translator’s strategies by highlighting the dialectic of hermetics and hermeneutics. The juxtaposition of the poem’s versions, elaborated by Jaccottet during the work in progress, provides information on the stages of transparent translation starting with the lexical work which concerns the search for equivalence. It reveals furthermore the hesitations and doubts of a wise translator, aware of the impossibility of understanding the poetic statement by neglecting its emotional, imaginal and prosodic aspects. In order to reach the conclusion that this is indeed a “relevant” translation, resulting from a meticulous work on the potentialities of the target language, which does not hesitate to transgress its grammatical norms with the purpose of preserving the foreignness of the source text. francaisEn 1969, Philippe Jaccottet traduit pour la revue Aquila le poeme Dunja de Giuseppe Ungaretti. L’etude tente de relever la singularite de cette traduction, en jetant de la lumiere sur la genetique de sa redaction: autrement dit, de penetrer le cote invisible du processus traductif auquel collaborent l’auteur et le traducteur. A partir de la clarification terminologique de la notion de transparence, elle s’applique a elucider les strategies du traducteur, en mettant en relief la dialectique de l’hermetique et de l’hermeneutique. La juxtaposition des versions jaccottiennes du poeme, elaborees au cours du « work in progress », renseigne sur les etapes de la traduction transparente, a commencer par le travail sur le lexique qui concerne la recherche d’equivalence. Elle revele en outre les hesitations et les doutes d’un traducteur averti, conscient de l’impossibilite de comprendre l’enonce poetique, en negligeant ses aspects affectif, imaginal et prosodique. Pour aboutir a la conclusion qu’il s’agit bien d’une traduction « relevante », resultat d’un travail minutieux sur les potentialites de la langue cible, qui n’hesite pas a transgresser les normes grammaticales de celle-ci afin de preserver l’etrangete du texte source.
Test publishers usually provide confidence intervals (CIs) for normed test scores that reflect the uncertainty due to the unreliability of the tests. The uncertainty due to sampling variability in the norming phase is ignored. To express uncertainty due to norming, we propose a flexible method that is applicable in continuous norming and allows for a variety of score distributions, using Generalized Additive Models for Location, Scale, and Shape (GAMLSS; Rigby & Stasinopoulos, 2005). We assessed the performance of this method in a simulation study, by examining the quality of the resulting CIs. We varied the population model, procedure of estimating the CI, confidence level, sample size, value of the predictor, extremity of the test score, and type of variance-covariance matrix. The results showed that good quality of the CIs could be achieved in most conditions. The method is illustrated using normative data of the SON-R 6-40 test. We recommend test developers to use this approach to arrive at CIs, and thus properly express the uncertainty due to norm sampling fluctuations, in the context of continuous norming. Adopting this approach will help (e.g., clinical) practitioners to obtain a fair picture of the person assessed.
The paper introduces the project of the Index Thomisticus Treebank (IT-TB). The IT-TB is a dependency-based treebank based on the corpus of the Index Thomisticus by father Roberto Busa (IT), which includes the opera omnia of Thomas Aquinas, for a total of approximately 11 million words. Currently, the IT-TB is the largest Latin treebank available, with more than 350,000 nodes in around 17,000 sentences. The annotation covers the entire books 1, 2 and 3 of Summa contra Gentiles, plus excerpts from Scriptum super Sententiis Magistri Petri Lombardi and Summa Theologiae. The paper details the multi-layer annotation style of the IT-TB and its background theoretical motivations. The conversion process to the now widely used Universal Dependencies style is described as well. Across more than a decade, the project has developed a number of linguistic resources and NLP tools for Latin connected to the IT-TB. As for the resources, the paper presents the syntaxbased subcategorization lexicon IT-VaLex and the valency lexicon Latin Vallex. As for the tools, the automatic dependency parsing process is described, highlighting the core issue of portability of NLP tools across the wide diachronic and diatopic span of Latin texts. A section is dedicated to automatic morphological analysis of Latin, introducing the analyzer Lemlat and its recent enhancement with information on derivational morphology and a new set of lexical entries covering a large Onomasticon (from Forcellini dictionary) and Medieval Latin (from Du Cange glossary).
Neural parsers obtain state-of-the-art results on benchmark treebanks for constituency parsing-but to what degree do they generalize to other domains? We present three results about the generalization of neural parsers in a zero-shot setting: training on trees from one corpus and evaluating on out-of-domain corpora. First, neural and non-neural parsers generalize comparably to new domains. Second, incorporating pre-trained encoder representations into neural parsers substantially improves their performance across all domains, but does not give a larger relative improvement for out-of-domain treebanks. Finally, despite the rich input representations they learn, neural parsers still benefit from structured output prediction of output trees, yielding higher exact match accuracy and stronger generalization both to larger text spans and to out-of-domain corpora. We analyze generalization on English and Chinese corpora, and in the process obtain state-of-the-art parsing results for the Brown, Genia, and English Web treebanks.
Recent studies have shown that architectural interior forms could impact the affective state of inhabitants. However, the direct relation of specific forms with specific affective states is difficult to determine. In addition, no systematic categorization of architectural forms and their relation to emotional states exists. The investigation of the impact of architectural features on inhabitants' emotions is further complicated by the use of two-dimensional images of forms in laboratory investigations, which cannot perceive real-world architecture. Furthermore, the interior form consists of a combination of different forms rather than only pure forms, which was considered in previous studies. This study aimed to fill these gaps by evaluating interior forms on the basis of clustering different images of built living rooms throughout history as well as their impact on emotions. This study used pleasure, arousal, and dominance ratings with an emphasis on individual differences in personality. Virtual sample rooms were created based on formal clusters of architectural forms. Results showed a relationship between forms and emotional states for different personality traits. This work provided a novel approach on the influence of architecture on emotion by considering systematic form categorization and combinations, personality differences, and a virtual reality setup.
Personality affects the way someone feels or acts. This paper examines the effect of personality traits, as operationalized by the Big-five questionnaire, on the number, type and severity of identified usability issues, physiological signals (skin conductance), and subjective emotional ratings (valence-arousal). Twenty-four users interacted with a web service and then participated in a retrospective thinking aloud session. Results revealed that the number of usability issues is significantly affected by the Openness trait. Emotional Stability significantly affects the type of reported usability issues. Problem severity is not affected by any trait. Valence ratings are significantly affected by Conscientiousness, whereas Agreeableness, Emotional Stability and Openness significantly affect arousal ratings. Finally, Openness has a significant effect on the number of detected peaks in user's skin conductance.
ЛИНГВИСТИЧЕСКАЯ БАЗА ДАННЫХ ОТРИЦАТЕЛЬНО-ОЦЕНОЧНОЙ ЛЕКСИКИ: КОНЦЕПЦИЯ, СТРУКТУРА, НАПОЛНЕНИЕИсследование проведено при поддержке Фонда содействия развитию малых форм предприятий в научно-технической сфере по программе «УМНИК» по теме «Разработка программного обеспечения для поддержки процедуры лингвистической экспертизы»
a certain experience may be just as important as the experience itself. The peak-and-end-rule (PE-rule) postulates that remembered experiences are best predicted by the peak emotional valence and the emotional valence at the end of an experience in the here and now. The PE-rule, however, has mostly been assessed in experimental paradigms that induce relatively simple, one-dimensional experiences (e.g., experienced pain in a clinical setting). This hampers generalizations of the PE-rule to the experiences in everyday life. This paper evaluates the generalizability of the PE-rule to more complex and heterogeneous experiences by examining the PE-rule in a virtual reality (VR) experience, as VR combines improved ecological validity with rigorous experimental control. Findings indicate that for more complex and heterogeneous experiences, peak and end emotional valence are inferior to other measures (such as averaged valence and arousal ratings over the entire experiential episode) in predicting remembered experience. These findings suggest that the PE-rule cannot be generalized to ecologically more valid experiential episodes.
Imagination is an internally-generated process, where one can make oneself or other people appear as protagonists of a scene. How does the brain tag the protagonist of an imagined scene as being oneself or someone else? Crucially, during imagination, neither external stimuli nor motor feedback are available to disentangle imagining oneself from imagining someone else. Here, we test the hypothesis that an internal mechanism based on the neural monitoring of heartbeats could distinguish between self and other. 23 participants imagined themselves (from a first-person perspective) or a friend (from a third-person perspective) in various scenarios, while their brain activity was recorded with magnetoencephalography and their cardiac activity was simultaneously monitored. We measured heartbeat-evoked responses, i.e. transients of neural activity occurring in response to each heartbeat, during imagination. The amplitude of heartbeat-evoked responses differed between imagining oneself and imagining a friend, in the precuneus and posterior cingulate regions bilaterally. Effect size was modulated by the daydreaming frequency scores of participants but not by their interoceptive abilities. These results could not be accounted for by other characteristics of imagination (e.g., the ability to adopt the perspective, valence or arousal), nor by cardiac parameters (e.g., heart rate) or arousal levels (e.g. arousal ratings, pupil diameter). Heartbeat-evoked responses thus appear as a neural marker distinguishing self from other during imagination.
Trustworthiness and dominance impressions summarize trait judgments from faces. Judgments on these key traits are negatively correlated to each other in impressions of female faces, implying less differentiated impressions of female faces. Here we test whether this is true across many trait judgments and whether less differentiated impressions of female faces originate in different facial information used for male and female impressions or different evaluation of the same information. Using multidimensional rating datasets and data-driven modeling, we show that (a) impressions of women are less differentiated and more valence-laden than impressions of men and find that (b) these impressions are based on similar visual information across face genders. Female face impressions were more highly intercorrelated and were better explained by valence (Study 1). These intercorrelations were higher when raters more strongly endorsed gender stereotypes. Despite the gender difference, male and female impression models-derived from separate trustworthiness and dominance ratings of male and female faces-were similar to each other (Study 2). Further, both male and female models could manipulate impressions of faces of both genders (Study 3). The results highlight the high-level, evaluative effect of face gender in impression formation-women are judged negatively to the extent their looks do not conform to expectations, not because people use different facial information across genders but because people evaluate the information differently across genders. (PsycINFO Database Record (c) 2020 APA, all rights reserved).
In this paper, we investigate the aspect of structured output modeling for the state-ofthe-art graph-based neural dependency parser (Dozat and Manning, 2017). With evaluations on 14 treebanks, we empirically show that global output-structured models can generally obtain better performance, especially on the metric of sentence-level Complete Match. However, probably because neural models already learn good global views of the inputs, the improvement brought by structured output modeling is modest.
In this paper we present the ConlluEditor annotation tool for manual annotation of files in CoNLL-U format, such as Universal Dependencies treebanks. Apart from providing a graphical editor for basic and enhanced dependencies, multi-token words, it also runs validation scripts to find potential errors. ConlluEditor uses a client-server architecture. It is freely-available under the 3-Clause BSD License.
Over the last 20 years, the development of a wide range of treebanks that track the evolution of languages’ syntactic patterns through time has revolutionized the field of historical syntax. The range of treebanks now available facilitates research into the long histories of many of the major Indo-European languages. Although the field's essentially corpus-based methodology has not changed, the quantity of data now available and the ease and precision with which those data can be extracted have created new opportunities. For example, with a treebank it is possible to extract all examples of surface strings associated only with abstract structures (e.g., relative clauses, extraposition), to investigate predictions made by syntactic analyses, to search for rare constructions, and to extract enough data to support sophisticated statistical analyses. Crucially, treebanks make verification and replicability of results possible.
This paper suggests one way to enhance the ability of Chinese learners to analyze sentences. It is building a treebank and visualizing it as a syntactic tree(or parsed tree) and providing it to learners. The process of building a treebank, which is the most important key in this method, is divided into three parts and described in detail.
This chapter is the first large-scale typological survey of the lexical means used in African languages to express color-related meanings. It is based on a very large sample, with data from 350 languages, most of which come from the RefLex online lexical database. It focuses on language-internal semantic sources, morphosyntactic strategies, and contact-induced terminology used for color naming. After a brief discussion of the issues raised by “basic” color terms, and “polychromatic” color terms, the chapter provides a review of the semantic sources of color terms, the origin of borrowings, colexifications and metaphorical uses of color terms, main patterns of lexicalization, and, briefly, color-related ideophones.
Callous-unemotional (CU) traits are associated with lower emotional reactivity in adolescents. However, since previous studies have focused mainly on reactivity to negative stimuli, it is unclear whether reactivity to positive stimuli is also affected. Further, few studies have addressed the link between CU traits and emotional reactivity in longitudinal community samples, which is important for determining its generalizability and developmental course. In the current study, pupil dilation and self-ratings of arousal and valence were assessed in 100 adolescents (15-17 years) from a community sample, while viewing images with negative and positive valence from the International Affective Pictures System (IAPS). Behavioral traits (CU) were assessed concurrently, as well as at ages 12-15, and 8-9 (subsample, n = 68, low levels of prosocial behavior were used as a proxy for CU traits). The results demonstrate that CU traits assessed at ages 12-15 and 8-9 predicted less pupil dilation to both positive and negative images at ages 15-17. Further, CU traits at ages 12-15 and concurrently were associated with less negative valence ratings for negative images and concurrently to less positive valence ratings for positive images. The current findings demonstrate that CU traits are related to lower emotional reactivity to both negative and positive stimuli in adolescents from a community sample.
Affective states underlie daily decision-making and pathological behaviours relevant to obsessive-compulsive disorders (OCD), mood disorders and addictions. Deep brain stimulation targeting the motor and associative-limbic subthalamic nucleus (STN) has been shown to be effective for Parkinson's disease (PD) and OCD, respectively. Cognitive and electrophysiological studies in PD showed responses of the motor STN to emotional stimuli, impairments in recognition of negative affective states and modulation of the intensity of subjective emotion. Here we studied whether the stimulation of the associative-limbic STN in OCD influences the subjective emotion to low-intensity positive and negative images and how this relates to clinical symptoms. We assessed 10 OCD patients with on and off STN DBS in a double-blind randomized manner by recording ratings of valence and arousal to low- and high-intensity positive and negative emotional images. STN stimulation increased positive ratings and decreased negative ratings to low-intensity positive and negative stimuli, respectively, relative to off stimulation. We also show that the change in severity of obsessive-compulsive symptoms pre- versus post-operatively interacts with both DBS and valence ratings. We show that stimulation of the associative-limbic STN might influence the negative cognitive bias in OCD and decreasing the negative appraisal of emotional stimuli with a possible relationship with clinical outcomes. That the effect is specific to low intensity might suggest a role of uncertainty or conflict related to competing interpretations of image intensity. These findings may have implications for the therapeutic efficacy of DBS.
Dependency distance minimization (DDm) is a word order principle favouring the placement of syntactically related words close to each other in sentences. Massive evidence of the principle has been reported for more than a decade with the help of syntactic dependency treebanks where long sentences abound. However, it has been predicted theoretically that the principle is more likely to be beaten in short sequences by the principle of surprisal minimization (predictability maximization). Here we introduce a simple binomial test to verify such a hypothesis. In short sentences, we find anti-DDm for some languages from different families. Our analysis of the syntactic dependency structures suggests that anti-DDm is produced by star trees.
Abstract In a usage-based framework, variation is part and parcel of our linguistic experiences, and therefore also of our mental representations of language. In this article, we bring attention to variation as a source of information. Instead of discarding variation as mere noise, we examine what it can reveal about the representation and use of linguistic knowledge. By means of metalinguistic judgment data, we demonstrate how to quantify and interpret four types of variation: variation across items, participants, time, and methods. The data concern familiarity ratings assigned by 91 native speakers of Dutch to 79 Dutch prepositional phrases such as in de tuin ‘in the garden’ and rond de ingang ‘around the entrance’. Participants performed the judgment task twice within a period of one to two weeks, using either a 7-point Likert scale or a Magnitude Estimation scale. We explicate the principles according to which the different types of variation can be considered information about mental representation, and we show how they can be used to test hypotheses regarding linguistic representations.
Abstract Lexicalized parsing models are based on the assumptions that (i) constituents are organized around a lexical head and (ii) bilexical statistics are crucial to solve ambiguities. In this paper, we introduce an unlexicalized transition-based parser for discontinuous constituency structures, based on a structure-label transition system and a bi-LSTM scoring system. We compare it with lexicalized parsing models in order to address the question of lexicalization in the context of discontinuous constituency parsing. Our experiments show that unlexicalized models systematically achieve higher results than lexicalized models, and provide additional empirical evidence that lexicalization is not necessary to achieve strong parsing results. Our best unlexicalized model sets a new state of the art on English and German discontinuous constituency treebanks. We further provide a per-phenomenon analysis of its errors on discontinuous constituents.
In this paper, we propose a novel data augmentation method with respect to the target context of the data via self-supervised learning. Instead of looking for the exact synonyms of masked words, the proposed method finds words that can replace the original words considering the context. For self-supervised learning, we can employ the masked language model (MLM), which masks a specific word within a sentence and obtains the original word. The MLM learns the context of a sentence through asymmetrical inputs and outputs. However, without using the existing MLM, we propose a label-masked language model (LMLM) that can include label information for the mask tokens used in the MLM to effectively use the MLM in data with label information. The augmentation method performs self-supervised learning using LMLM and then implements data augmentation through the trained model. We demonstrate that our proposed method improves the classification accuracy of recurrent neural networks and convolutional neural network-based classifiers through several experiments for text classification benchmark datasets, including the Stanford Sentiment Treebank-5 (SST5), the Stanford Sentiment Treebank-2 (SST2), the subjectivity (Subj), the Multi-Perspective Question Answering (MPQA), the Movie Reviews (MR), and the Text Retrieval Conference (TREC) datasets. In addition, since the proposed method does not use external data, it can eliminate the time spent collecting external data, or pre-training using external data.
Abstract Multiword expressions can have both idiomatic and literal occurrences. For instance pulling strings can be understood either as making use of one’s influence, or literally. Distinguishing these two cases has been addressed in linguistics and psycholinguistics studies, and is also considered one of the major challenges in MWE processing. We suggest that literal occurrences should be considered in both semantic and syntactic terms, which motivates their study in a treebank. We propose heuristics to automatically pre-identify candidate sentences that might contain literal occurrences of verbal VMWEs, and we apply them to existing treebanks in five typologically different languages: Basque, German, Greek, Polish and Portuguese. We also perform a linguistic study of the literal occurrences extracted by the different heuristics. The results suggest that literal occurrences constitute a rare phenomenon. We also identify some properties that may distinguish them from their idiomatic counterparts. This article is a largely extended version of Savary and Cordeiro (2018).
This paper presents work on the creation of a Universal Dependency (UD) treebank for Wolof as the first UD treebank within the Northern Atlantic branch of the Niger-Congo languages. The paper reports on various issues related to word segmentation for tokenization and the mapping of PoS tags, morphological features and dependency relations to existing conventions for annotating Wolof. It also outlines some specific constructions as a starting point for discussing several more general UD annotation guidelines, in particular for noun class marking, deixis encoding, and focus marking.
The current CoNLL version of the second part of the Late Latin Charter Treebank (LLCT2). Early medieval Latin documentary texts with morphological and syntactic annotation. Ancient Language Dependency Treebank (ALDT) compatible linguistic annotation with modifications concerning morphology (see Korkiakangas & Passarotti, 2011, “Challenges in Annotating Medieval Latin Charters”). LLCT2 expands the chronological span of LLCT1 up to AD 897. LLCT2 contains 521 charters and 257,918 tokens. See Korkiakangas, [in print], “Late Latin Charter Treebank: contents and annotation”.
Abstract A puzzling fact about linguistic norms is that they are mainly stable, but the conventional variant sometimes changes. These transitions seem to be mostly S-shaped and, therefore, directed. Previous models have suggested possible mechanisms to explain these directed changes, mainly based on a bias favoring the innovative variant. What is still debated is the origin of such a bias. In this paper, we propose a refined taxonomy of mechanisms of language change and identify a family of mechanisms explaining self-actuated language changes. We exemplify this type of mechanism with the preference-based selection mechanism that relies on agents having dynamic preferences for different variants of the linguistic norm. The key point is that if these preferences align through social interactions, then new changes can be actuated even in the absence of external triggers. We present results of a multi-agent model and demonstrate that the model produces trajectories that are typical of language change.
xml treebank Annotated by Toon Van Hal, with student contributions by Mathieu Cuijpers; Sanderijn Gijbels; Yoran Joosten; Yordi Lenaerts; Eva Uffing; Chiara Van der Hasselt; Lisa Vanhee and Jolien Volders (KU Leuven Bachelor 3, 2018-2019). Based on a preparsed text by Alek Keersmaekers. Controlled by Toon Van Hal, Sanderijn Gijbels and Yoran Joosten.
Despite the fact that there are a number of researches working on Khmer Language in the field of Natural Language Processing along with some resources regarding words segmentation and POS Tagging, we still lack of high-level resources regarding syntax, Treebanks and grammars, for example. This paper illustrates the semi-automatic framework of constructing Khmer Treebank and the extraction of the Khmer grammar rules from a set of sentences taken from the Khmer grammar books. Initially, these sentences will be manually annotated and processed to generate a number of grammar rules with their probabilities once the Treebank is obtained. In our experiments, the annotated trees and the extracted grammar rules are analyzed in both quantitative and qualitative way. Finally, the results will be evaluated in three evaluation processes including Self-Consistency, 5-Fold Cross-Validation, Leave-One-Out Cross-Validation along with the three validation methods such as Precision, Recall, F1-Measure. According to the result of the three validations, Self-Consistency has shown the best result with more than 92%, followed by the Leave-One-Out Cross-Validation and 5-Fold Cross Validation with the average of 88% and 75% respectively. On the other hand, the crossing bracket data shows that Leave-One-Out Cross Validation holds the highest average with 96% while the other two are 85% and 89%, respectively.
Imagining fictional creatures like zombies in survival situations boosts long-term memory for words encoded in these situations more than rating words for pleasantness (zombie effect). Study 1 required word-ratings in a zombie-survival scenario; participants were told they had to protect against either possible zombie attack or contamination. The zombie-survival situations yielded identical recall levels but higher recall rates than pleasantness. Study 2 matched a zombie-survival scenario on perceived fear with scenarios involving ghosts or predators. Perceived disgust in the zombie scenario was higher than in these other survival conditions. Words were remembered better when processed in survival scenarios than when rated for pleasantness, but there was no reliable difference in recall between the scenarios. In neither study did the number of death-related words produced in a word-fragment completion task fit the mortality salience account of the zombie memory effect. Overall findings suggest that this effect relates to the fear system.
Word vectors are at the core of many natural language processing tasks. Recently, there has been interest in post-processing word vectors to enrich their semantic information. In this paper, we introduce a novel word vector post-processing technique based on matrix conceptors (Jaeger 2014), a family of regularized identity maps. More concretely, we propose to use conceptors to suppress those latent features of word vectors having high variances. The proposed method is purely unsupervised: it does not rely on any corpus or external linguistic database. We evaluate the post-processed word vectors on a battery of intrinsic lexical evaluation tasks, showing that the proposed method consistently outperforms existing state-of-the-art alternatives. We also show that post-processed word vectors can be used for the downstream natural language processing task of dialogue state tracking, yielding improved results in different dialogue domains.
We introduce the first German treebank for Twitter microtext, annotated within the framework of Universal Dependencies. The new treebank includes over 12,000 tokens from over 500 tweets, independently annotated by two human coders. In the paper, we describe the data selection and annotation process and present baseline parsing results for the new testsuite.
We present UDify, a multilingual multi-task model capable of accurately predicting universal part-of-speech, morphological features, lemmas, and dependency trees simultaneously for all 124 Universal Dependencies treebanks across 75 languages. By leveraging a multilingual BERT self-attention model pretrained on 104 languages, we found that fine-tuning it on all datasets concatenated together with simple softmax classifiers for each UD task can result in state-of-the-art UPOS, UFeats, Lemmas, UAS, and LAS scores, without requiring any recurrent or language-specific components. We evaluate UDify for multilingual learning, showing that low-resource languages benefit the most from cross-linguistic annotations. We also evaluate for zero-shot learning, with results suggesting that multilingual training provides strong UD predictions even for languages that neither UDify nor BERT have ever been trained on. Code for UDify is available at https://github.com/hyperparticle/udify.
This is the first versioned collection of.xml files containing approximately 550,000 tokens of ancient Greek prose that have been hand-analyzed into dependency syntax using Perseids/Arethusa by Prof. Vanessa Gorman of the University of Nebraska-Lincoln. CC0 1.0 license.
We present a novel semantic framework for modeling linguistic expressions of\ngeneralization---generic, habitual, and episodic statements---as combinations\nof simple, real-valued referential properties of predicates and their\narguments. We use this framework to construct a dataset covering the entirety\nof the Universal Dependencies English Web Treebank. We use this dataset to\nprobe the efficacy of type-level and token-level information---including\nhand-engineered features and static (GloVe) and contextual (ELMo) word\nembeddings---for predicting expressions of generalization. Data and code are\navailable at decomp.io.\n
In the past decades, linguistic typology went through a renewing phase that involved a significant change in the research questions and methods of the discipline, which is now interested in fine-grained features underlying language diversity. In this paper, we propose a novel approach to address the newly defined needs of linguistic typology by extracting qualitative and quantitative information about a wide range of features from multilingual annotated corpora based on Natural Language Processing methods and techniques. We tested our method in a case study focusing on word order variation in two widely investigated constructions, VERB-SUBJ(ect) and NOUN-ADJ(ective), with a specific view to structural and functional factors underlying the preference for one or the other order, both intra- and cross-linguistically, and their interaction. Preliminary experiments have been carried out aimed at acquiring typological evidence from a selection of linguistically annotated treebanks for three different languages, namely Italian, Spanish and English. Our results show the effectiveness of the method in letting similarities and differences also emerge from typologically close languages.
Recently, it has been shown that various auditory stimuli modulate flavour perception. The present study attempts to understand the effects of environmental sounds (park, food court, fast food restaurant, cafe, and bar sounds) on the perception of chocolate gelato (specifically, sweet, bitter, milky, creamy, cocoa, roasted, and vanilla notes) using the Temporal Check-All-That-Apply (TCATA) method. Additionally, affective ratings of the auditory stimuli were obtained using the Self-Assessment Manikin (SAM) in terms of their valence, arousal, and dominance. In total, 58 panellists rated the sounds and chocolate gelato in a sensory laboratory. The results revealed that bitterness, roasted, and cocoa notes were more evident when the bar, fast food, and food court sounds were played. Meanwhile, sweetness was cited more in the early mastication period when listening to park and café sounds. The park sound was significantly higher in valence, while the bar sound was significantly higher in arousal. Dominance was significantly higher for the fast food restaurant, food court, and bar sound conditions. Intriguingly, the valence evoked by the pleasant park sound was positively correlated with the sweetness of the gelato. Meanwhile, the arousal associated with bar sounds was positively correlated with bitterness, roasted, and cocoa attributes. Taken together, these results clearly demonstrate that people's perception of the flavour of gelato varied with the different real-world sounds used in this study.
The affect associated with negative events fades faster than the affect associated with positive events (the fading affect bias). The fading affect bias is present in most participants and is thought to be evidence of a healthy coping mechanism operating in autobiographical memory. Prior research shows that the fading affect bias can be distorted by negative individual difference variables such as dysphoria and anxiety. The goal of this research is to link the fading affect bias to the positive individual difference variable of Grit. A total of 197 participants completed the short Grit Scale and were divided into four groups based on their Grit scores (i.e., low Grit to high Grit). Participants retrieved positive and negative event memories and then made affect ratings for the events. The results show that increased levels of Grit were associated with a stronger fading affect bias.
Several studies have attempted to investigate how the brain codes emotional value when processing music of contrasting levels of dissonance; however, the lack of control over specific musical structural characteristics (i.e., dynamics, rhythm, melodic contour or instrumental timbre), which are known to affect perceived dissonance, rendered results difficult to interpret. To account for this, we used functional imaging with an optimized control of the musical structure to obtain a finer characterization of brain activity in response to tonal dissonance. Behavioral findings supported previous evidence for an association between increased dissonance and negative emotion. Results further demonstrated that the manipulation of tonal dissonance through systematically controlled changes in interval content elicited contrasting valence ratings but no significant effects on either arousal or potency. Neuroscientific findings showed an engagement of the left medial prefrontal cortex (mPFC) and the left rostral anterior cingulate cortex (ACC) while participants listened to dissonant compared to consonant music, converging with studies that have proposed a core role of these regions during conflict monitoring (detection and resolution), and in the appraisal of negative emotion and fear-related information. Both the left and right primary auditory cortices showed stronger functional connectivity with the ACC during the dissonant portion of the task, implying a demand for greater information integration when processing negatively valenced musical stimuli. This study demonstrated that the systematic control of musical dissonance could be applied to isolate valence from the arousal dimension, facilitating a novel access to the neural representation of negative emotion.
This paper presents challenges and observations on creating a code-switching treebank based on ongoing annotation efforts of a Turkish-German spoken corpus following the Universal Dependencies annotation scheme. We present and discuss a number of issues that arise because of the need for consistent multilingual annotation within a single treebank, as well as the informal language which is where code-switching is observed most. Besides proposing solutions to these issues, our aim in this paper is to stimulate discussion and facilitate consistency over upcoming code-switching annotation projects.
Facial expressions are fundamental to interpersonal communication, including social interaction, and allow people of different ages, cultures, and languages to quickly and reliably convey emotional information. Historically, facial expression research has followed from discrete emotion theories, which posit a limited number of distinct affective states that are represented with specific patterns of facial action. Much less work has focused on dimensional features of emotion, particularly positive and negative affect intensity. This is likely, in part, because achieving inter-rater reliability for facial action and affect intensity ratings is painstaking and labor-intensive. We use computer-vision and machine learning (CVML) to identify patterns of facial actions in 4,648 video recordings of 125 human participants, which show strong correspondences to positive and negative affect intensity ratings obtained from highly trained coders. Our results show that CVML can both (1) determine the importance of different facial actions that human coders use to derive positive and negative affective ratings when combined with interpretable machine learning methods, and (2) efficiently automate positive and negative affect intensity coding on large facial expression databases. Further, we show that CVML can be applied to individual human judges to infer which facial actions they use to generate perceptual emotion ratings from facial expressions.
This article presents an analysis of experiments with statistical and neural parsing techniques for Urdu, a widely spoken South Asian language. We demonstrate state of the art constituency parsing results for an Urdu treebank. Urdu is a morphologically rich and is characterized by free word order. Language representation (e.g. input type, lemmatization, word clusters), part of speech tag set, phrase labels and the size of a training corpus are crucial for parsing such languages. In this article, probabilistic context-free grammars, data-oriented parsing, and recursive neural network based models have been experimented with several linguistic features which show improvements in the parsing results. Features include syntactic sub-categorization of POS tags, empirically learned horizontal and vertical markovizations and lexical head words. These features enable dependency information for case markers and add phrasal and lexical context to the parse trees. The data-oriented parsing and recursive neural network model give an f-score of 87.1 by considering gold POS tags in the test set, on textual input, they show a performance with f-scores of 83.4 and 84.2, respectively. To overcome the issue of data sparsity due to the morphological richness, lemmatization and unsupervised word clustering have been performed. A treebank should cover most probable word orders of the language so that models can learn various orders accurately. To analyze the order coverage of the treebank and learning capability of different parsers, a test set has been prepared conditioning different word orders. This test set is evaluated with the best performing parsing models and with gold POS tags, f-scores are above 90 and on textual input, the average f-score is 87.6.
Neural models have been investigated for sentiment classification over constituent trees. They learn phrase composition automatically by encoding tree structures but do not explicitly model sentiment composition, which requires to encode sentiment class labels. To this end, we investigate two formalisms with deep sentiment representations that capture sentiment subtype expressions by latent variables and Gaussian mixture vectors, respectively. Experiments on Stanford Sentiment Treebank (SST) show the effectiveness of sentiment grammar over vanilla neural encoders. Using ELMo embeddings, our method gives the best results on this benchmark.
Cross-lingual transfer is an effective way to build syntactic analysis tools in low-resource languages. However, transfer is difficult when transferring to typologically distant languages, especially when neither annotated target data nor parallel corpora are available. In this paper, we focus on methods for cross-lingual transfer to distant languages and propose to learn a generative model with a structured prior that utilizes labeled source data and unlabeled target data jointly. The parameters of source model and target model are softly shared through a regularized log likelihood objective. An invertible projection is employed to learn a new interlingual latent embedding space that compensates for imperfect crosslingual word embedding input. We evaluate our method on two syntactic tasks: part-ofspeech (POS) tagging and dependency parsing. On the Universal Dependency Treebanks, we use English as the only source corpus and transfer to a wide range of target languages. On the 10 languages in this dataset that are distant from English, our method yields an average of 5.2% absolute improvement on POS tagging and 8.3% absolute improvement on dependency parsing over a direct transfer method using state-of-the-art discriminative models. 1 3 Following Ahmad et al. ( fastText_multilingual, which contains alignment matrices for 78 languages, which also allows comparison with their numbers in Section 4.3.