Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The paper studies the effect of emotional states modulated by auditory stimuli on the cognitive control on decision making. Based on other previous neuroimaging studies, functional near-infrared spectroscopy provided reliable neuroimaging measurement in analyzing emotional states by studying the changes of hemodynamic response in prefrontal cortex (PFC). This experiment involved 16 nursing students. During the experiment, participants were given one minute to complete five nursing practice questions with five sequential repetitions in the presence of neutral and negative emotional auditory stimuli in two separated sessions under fNIRS measurement. The sound stimuli was selected from the International Affective Digitized Sound (IADS) System. The neutral auditory stimuli had neutral valence and medium arousal rating whereas negative auditory stimuli had negative valence and high arousal rating. The data collected was preprocessed by using wavelet transform to decompose the data into different frequency intervals. By selecting the frequency interval of interest, we analyzed the data based on functional connectivity within prefrontal cortex regions. We computed the regional wavelet coherence values between affective and neutral tasks. From the behavioral analysis, we found that subjects had significantly higher accuracy in affective task compared to neutral task. Based on the analysis, we found that left prefrontal cortex produced significantly lower wavelet coherence value but the highest coherence-accuracy correlation in affective task than in neutral task.
This paper is a linguistic as well as technical survey for the development of a shallow discourse parser for Czech. It focuses on long-distance discourse relations signalled by (mostly) anaphoric discourse connectives. Proceeding from the division of connectives on “structural” and “anaphoric” according to their (in)ability to accept distant (non-adjacent) text segments as their left-sided arguments, and taking into account results of related analyses on English data in the framework of the Penn Discourse Treebank, we analyze a large amount of language data in Czech. We benefit from the multilayer manual annotation of various language aspects from morphology to discourse, coreference and bridging relations in the Prague Dependency Treebank 3.0. We describe the linguistic parameters of long-distance discourse relations in Czechin connection with their anchoring connective, and suggest possible ways of their detection. Our empirical research also outlines some theoretical consequences for the underlying assumptions in discourse analysis and parsing, e.g. the risk of relying too much on different (language-specific?) part-of-speech categorizations of connectives or the different perspectives in shallow and global discourse analyses (the minimality principle vs. higher text structure).
OBJECTIVES: One-on-one structured Montessori-based activities conducted with people with dementia can improve agitation and enhance engagement. These activities may however not always be implemented by nursing home staff. Family members may present an untapped resource for enabling these activities. This study aimed to evaluate the impact of the Montessori activities implemented by family members on visitation experiences with people who have dementia. DESIGN: Cluster-randomized crossover design. SETTING: General and psychogeriatric nursing homes in the state of Victoria, Australia. PARTICIPANTS: Forty participants (20 residents and 20 carers) were recruited. INTERVENTION: During visits, family members interacted with their relative either through engaging in Montessori-based activities or reading a newspaper (the control condition) for four 30-minute sessions over 2 weeks. MEASUREMENTS: Residents' predominant affect and engagement were rated for each 30-second interval using the Philadelphia Geriatric Center Affect Rating Scale and the Menorah Park Engagement Scale. The Pearlin Mastery Scale was used to rate carers satisfaction with visits. The 15-item Mutuality Scale measured the carers quality of their relationship with the resident. Carers' mood and overall quality of life were measured using the Center for Epidemiological Studies Depression Scale and Carer-QoL questionnaires, respectively. RESULTS: Linear regressions within the generalized estimating equations approach assessed residents' and carers' outcomes. Relative to the control condition, the Montessori condition resulted in more positive engagement (b = 13.0, 95%CI 6.3-19.7, p < 0.001) and affect (b = 0.4, 95%CI 0.2-0.6, p < 0.001) for the residents and higher satisfaction with visits for carers (b = 1.7, 95%CI 0.45-3.00, p = 0.008). No correction was applied to p-values for multiple comparisons. CONCLUSION: This study strengthens the evidence base for the use of the Montessori programs in increasing well-being in nursing home residents. The findings also provide evidence that family members are an additional valuable resource in implementing structured activities such as the Montessori program with residents.
We present a comparative analysis of PP ordering in English and (Mandarin) Chinese, two languages with distinct typological word order characteristics. Previous work on PP orderings have mainly focused on English using data of relatively small size. Here we leverage corpora of much larger scale with straightforward annotations. We use the Penn Treebank for English, which includes three corpora that cover both written and spoken domains, and the Chinese Penn Treebank for Chinese. We explore the individual effect of dependency length, the argument status of the PP (argument or adjunct) and the traditional adverbial ordering rule, Manner before Place before Time. In addition, we evaluate the predictive power of dependency length and argument status with weights estimated from logistic regression models. We show that while dependency length plays a strong role across genre for English, it only exerts a mild effect in Chinese. On the other hand, the argument status of the PP has a pronounced role in both languages, that is, there exists a strong tendency for the argument-like PP to appear closer to the head verb than the adjunct-like PP. Our work contributes empirically to the long-standing proposal in linguistic typology that crosslinguistic word ordering preference is driven by cooperating and competing principles.
Previous research has shown that that evaluative verbal information (praise and criticism) conveys different affective values: criticism is perceived as unpleasant while praise is generally considered pleasant. Here, using praise and criticism in Chinese, we investigated how affective value is modulated in men and women, depending on the particular attribute (personality vs. appearance) targeted by social comments. Results showed that whereas praise was rated as pleasant and criticism as unpleasant overall, criticizing personality reduced pleasantness more than criticizing appearance. In men, moreover, criticism of personality was deemed more unpleasant than criticism of appearance while personality-targeted praise was rated more pleasant than appearance-targeted praise. This effect was absent in women and consistent with men's higher arousal ratings for personality- relative to appearance-targeted comments. Our findings suggest that men are more concerned about external perception of their personality than that of their appearance whereas women's affective judgment is more balanced. These gender-specific results may have implications for topic selection in evaluative social communication.
We describe a cross-lingual transfer method for dependency parsing that takes into account the problem of word order differences between source and target languages. Our model only relies on the Bible, a considerably smaller parallel data than the commonly used parallel data in transfer methods. We use the concatenation of projected trees from the Bible corpus, and the gold-standard treebanks in multiple source languages along with cross-lingual word representations. We demonstrate that reordering the source treebanks before training on them for a target language improves the accuracy of languages outside the European language family. Our experiments on 68 treebanks (38 languages) in the Universal Dependencies corpus achieve a high accuracy for all languages. Among them, our experiments on 16 treebanks of 12 non-European languages achieve an average UAS absolute improvement of 3.3% over a state-of-the-art method.
The paper aims to examine how the acoustic input (the surface form) and the abstract linguistic representation (the underlying representation) interact during spoken word recognition by investigating left-dominant tone sandhi, a tonal alternation in which the underlying tone of the first syllable spreads to the sandhi domain. We conducted two auditory-auditory priming lexical decision experiments on Shanghai left-dominant sandhi words with less-frequent and frequent Shanghai users, in which each disyllabic target was preceded by monosyllabic primes either sharing the same underlying tone, surface tone, or being unrelated to the tone of the first syllable of the sandhi targets. Results showed a surface priming effect but not an underlying priming effect in younger speakers who used Shanghai less frequently, but no surface or underlying priming effect in older speakers who used Shanghai more often. Moreover, the surface priming did not interact with speakers' familiarity ratings to the sandhi targets. These patterns suggest that left-dominant Shanghai sandhi words may be represented in the sandhi form in the mental lexicon. The results are discussed in the context of how phonological opacity, productivity, the non-structure-preserving nature of tone spreading, and speakers' semantic knowledge influence the representation and processing of tone sandhi words.
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple BERT-based LS approach that makes use of the pre-trained unsupervised deep bidirectional representations BERT. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement than these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 11 Accuracy points on three well-known benchmarks.
Although there is a wide consensus on how sleep processes declarative memories, how sleep affects emotional memories remains elusive. Moreover, studies assessing the long-term effect of sleep on emotional memory consolidation are scarce. Studies testing subclinical populations characterized by REM abnormalities are also lacking. Here we aimed to (i) investigate the fate of emotional memories and the potential unbinding (or preservation) between content and affective tone over time (i.e., 1 week), (ii) explore the role of seven nights of sleep (recorded via actigraphy) in emotional memory consolidation, and (iii) assess whether participants with self-reported mild-moderate depressive symptoms forget less emotional information compared to participants with low depression symptoms. We found that, although at the immediate recognition session emotional information was forgotten more than neutral information, a week later it was forgotten less than neutral information. This effect was observed both in participants with low and mild-moderate depressive symptoms. We also observed an increase in valence rating over time for negative pictures, whereas perceived arousal diminished a week later for both types of stimuli (unpleasant and neutral); an initial decrease was already observable at the immediate recognition session. Interestingly, we observed a negative association between sleep efficiency across the week and change in memory discrimination for unpleasant pictures over time, i.e., participants who slept worse were the ones who forgot less emotional information. Our results suggest that emotional memories are resistant to forgetting, particularly when sleep is disrupted, and they are not affected by non-clinical depression symptomatology.
In three experiments we investigated whether memory-independent evaluative conditioning (EC) and other memory-independent contingency learning (CL) effects occur in the valence contingency task (VCT). In the VCT, participants respond to the valence of a target word that is preceded by a nonword. Across trials, each nonword is mostly combined with either positive or negative targets. Schmidt and De Houwer (2012. Contingency learning with evaluative stimuli. Experimental Psychology, 59, 175–182. doi:10.1027/1618-3169/a000141) showed faster and more often correct responses on trials that conformed to this contingency. Additionally, the authors found EC on valence ratings assessed after the VCT. All effects occurred also in the absence of contingency memory. Our Experiments 1a and 1b replicated the CL effects on measures assessed during the VCT (RT, errors) and showed that they occurred in the absence of contingency memory, but they did not replicate the EC effect assessed after the VCT. In Experiment 2, we tested whether this dissociation between EC and other CL effects was due to the different phases (during vs. after VCT) with a CL measure that could be used in both phases. On this measure, the CL effect was memory-dependent after, but not during the VCT. Across measures and experiments, we thus find memory-independent CL during the VCT, but not afterwards.
Discourse markers are words and expressions (such as: firstly, then, for example, because, as a result, likewise, in comparison, in contrast) that explicitly state the relational structure of the information in the text, i.e. signalling a sequential relationship between the current message and the previous discourse. Using these markers improves the cohesion and coherence of texts, facilitating reading comprehension. Although often included in tools that support the rhetoric structuring of texts, discourse markers have hardly been explored in writing support tools for learners of a second language. However, learners of a second language, including those at advanced levels, have trouble producing these lexical items, frequently replacing them with items from their native language or with literal translations of items in their own language, which often do not result in proper lexical items in the second language. In addition, students learn a single marker per function and use it repeatedly, producing monotonous texts. With the aim of contributing to reducing these difficulties, this paper presents a lexicon that will be used to support the task of automatically detecting and correcting discourse marker errors. Several heuristics have been evaluated to generate different types of errors. Automatic translation methods were used to semi-automatically compile the lexicon used in these heuristics. Similarity measures were also combined with these heuristics to correct discourse marker errors. The evaluated methods proved to be suitable for the task of identifying some types of discourse marker errors and can potentially identify many others, as long as new lexical inputs are incorporated into them.
By building a part-of-speech (POS) tagger for Middle High German, we investigate strategies for dealing with a low resource, diverse and non-standard language in the domain of natural language processing. We highlight various aspects such as the data quantity needed for training and the influence of data quality on tagger performance. Since the lack of annotated resources poses a problem for training a tagger, we exemplify how existing resources can be adapted fruitfully to serve as additional training data. The resulting POS model achieves a tagging accuracy of about 91% on a diverse test set representing the different genres, time periods and varieties of MHG. In order to verify its general applicability, we evaluate the performance on different genres, authors and varieties of MHG, separately. We explore self-learning techniques which yield the advantage that unannotated data can be utilized to improve tagging performance on specific subcorpora.
In this paper we explain the difference between two aspects of semantic relatedness: taxonomic and thematic relations. We notice the lack of evaluation tools for measuring thematic relatedness, identify two datasets that can be recommended as thematic benchmarks, and verify them experimentally. In further experiments, we use these datasets to perform a comprehensive analysis of the performance of an extensive sample of computational models of semantic relatedness, classified according to the sources of information they exploit. We report models that are best at each of the two dimensions of semantic relatedness and those that achieve a good balance between the two.
The aim of this article is to analyse the attitude towards the linguistic norm of the first-year students of Romance Philology in Warsaw through the auto-narration. By this tool, belonging to the qualitative methodology, the learner performs a retrospective introspection and thus ‘evaluates’ the learning, in our context – his learning of grammar. The stories of students-future philologists will be analysed according to the following aspects: the importance given to the linguistic norm (understood here as the need for grammatical correction), the perception of its ‘utility’ in relation to other competences and language subsystems, the relationship between awareness of the norm and the effectiveness / quality of communication in a foreign language.
This paper presents experiments in part-of-speech tagging of low-resource languages. It addresses the case when no labeled data in the targeted language and no parallel corpus are available. We only rely on the proximity of the targeted language to a better-resourced language. We conduct experiments on three French regional languages. We try to exploit this proximity with two main strategies: delexicalization and transposition. The general idea is to learn a model on the (better-resourced) source language, which will then be applied to the (regional) target language. Delexicalization is used to deal with the difference in vocabulary, by creating abstract representations of the data. Transposition consists in modifying the target corpus to be able to use the source models. We compare several methods and propose different strategies to combine them and improve the state-of-the-art of part-of-speech tagging in this difficult scenario.
Perception of emotions and adequate responses are key factors of a successful conversational agent. However, determining emotions in a healthcare setting depends on multiple factors such as context and medical condition. Given the increase of interest in conversational agents integrated in mobile health applications, our objective in this work is to introduce a concept for analyzing emotions and sentiments expressed by a person in a mobile health application with a conversational user interface. The approach bases upon bot technology (Synthetic intelligence markup language) and deep learning for emotion analysis. More specifically, expressions referring to sentiments or emotions are classified along seven categories and three stages of strengths using treebank annotation and recursive neural networks. The classification result is used by the chatbot for selecting an appropriate response. In this way, the concerns of a user can be better addressed. We describe three use cases where the approach could be integrated to make the chatbot emotion-sensitive.
Like many other scientific disciplines, psychological science has felt the impact of the big-data revolution. This impact arises from the meeting of three forces: data availability, data heterogeneity, and data analyzability. In terms of data availability, consider that for decades, researchers relied on the Brown Corpus of about one million words (Kučera & Francis, 1969). Modern resources, in contrast, are larger by six orders of magnitude (e.g., Google’s 1T corpus) and are available in a growing number of languages. About 240 billion photos have been uploaded to Facebook,1 and Instagram receives over 100 million new photos each day.2 The large-scale digitization of these data has made it possible in principle to analyze and aggregate these resources on a previously unimagined scale. Heterogeneity refers to the availability of different types of data. For example, recent progress in automatic image recognition is owed not just to improvements in algorithms and hardware, but arguably more to the ability to merge large collections of images with linguistic labels (produced by crowdsourced human taggers) that serve as training data to the algorithms. Making use of heterogeneous data sources often depends on their standardization. For example, the ability to combine demographic and grammatical data about thousands of languages led to the finding that languages spoken by more people have simpler morphologies (Lupyan & Dale, 2010). The ability to combine these data types would have been substantially more difficult without the existence of standardized language and country codes that could be used to merge the different data sources. Finally, analyzability must be ensured, for without appropriate tools to process and analyze different types of data, the “data” are merely bytes.
The notion of spreading activation is a central theme in the cognitive sciences; however, the tools for implementing spreading activation computationally are not as readily available. This article introduces the spreadr R package, which can implement spreading activation within a specified network structure. The algorithmic method implemented in the spreadr subroutines follows the approach described in Vitevitch, Ercal, and Adagarla (Frontiers in Psychology, 2, 369, 2011), who viewed activation as a fixed cognitive resource that could “spread” among connected nodes in a network. Three sets of simulations were conducted using the package. The first set of simulations successfully reproduced the results reported in Vitevitch et al. (Frontiers in Psychology, 2, 369, 2011), who showed that a simple mechanism of spreading activation could account for the clustering coefficient effect in spoken word recognition. The second set of simulations showed that the same mechanism could be extended to account for higher false alarm rates for low clustering coefficient words in a false memory task. The final set of simulations demonstrated how spreading activation could be applied to a semantic network to account for semantic priming effects. It is hoped that this package will encourage cognitive and language scientists to explicitly consider how the structures of cognitive systems such as the mental lexicon and semantic memory interact with the process of spreading activation.
Purpose This paper aims to describe the structure of an aligned Serbian-German literary corpus (SrpNemKor) contained in a digital library Bibliša. The goal of the research was to create a benchmark Serbian-German annotated corpus searchable with various query expansions. Design/methodology/approach The presented research is particularly focused on the enhancement of bilingual search queries in a full-text search of aligned SrpNemKor collection. The enhancement is based on using existing lexical resources such as Serbian morphological electronic dictionaries and the bilingual lexical database Termi. Findings For the purpose of this research, the lexical database Termi is enriched with a bilingual list of German-Serbian translated pairs of lexical units. The list of correct translation pairs was extracted from SrpNemKor, evaluated and integrated into Termi. Also, Serbian morphological e-dictionaries are updated with new entries extracted from the Serbian part of the corpus. Originality/value A bilingual search of SrpNemKor in Bibliša is available within the user-friendly platform. The enriched database Termi enables semantic enhancement and refinement of user’s search query based on synonyms both in Serbian and German at a very high level. Serbian morphological e-dictionaries facilitate the morphological expansion of search queries in Serbian, thereby enabling the analysis of concepts and concept structures by identifying terms assigned to the concept, and by establishing relations between terms in Serbian and German which makes Bibliša a valuable Web tool that can support research and analysis of SrpNemKor.
For sequence models with large vocabularies, a majority of network parameters lie in the input and output layers. In this work, we describe a new method, DeFINE, for learning deep token representations efficiently. Our architecture uses a hierarchical structure with novel skip-connections which allows for the use of low dimensional input and output layers, reducing total parameters and training time while delivering similar or better performance versus existing methods. DeFINE can be incorporated easily in new or existing sequence models. Compared to state-of-the-art methods including adaptive input representations, this technique results in a 6% to 20% drop in perplexity. On WikiText-103, DeFINE reduces the total parameters of Transformer-XL by half with minimal impact on performance. On the Penn Treebank, DeFINE improves AWD-LSTM by 4 points with a 17% reduction in parameters, achieving comparable performance to state-of-the-art methods with fewer parameters. For machine translation, DeFINE improves the efficiency of the Transformer model by about 1.4 times while delivering similar performance.
Relation classification is a vital task in natural language processing, and it is screening for semantic relation between clauses in texts. This paper describes a study of relation classification on Chinese compound sentences without connectives. There exists an implicit relation in a compound sentence without connectives, which makes it difficult to realize the recognition of relation. The major challenges that relation classification modeling faces are how to obtain the contextual representation of sentence and relation dependence features between clauses. To solve this problem, we propose a novel Inatt-MCNN model to extract sentence features and classify relations by combining multi-channel CNN and Inner-attention mechanism. This network structure utilizes CNN to extract local features of sentences and Inner-attention to capture sentence-level feature representations for this relation classification task. Besides, since the Inner-attention is based on Bi-LSTM, the global and long-term dependence semantic information can be well obtained in Inatt-MCNN to promote the model performance. We conduct experiments on two public Chinese discourse datasets: the Chinese compound sentence corpus (CCCS) dataset and the Tsinghua Chinese Treebank(TCT) dataset. Compared with the previous public methods, Inatt-MCNN model has superior performance and achieves the highest accuracy, especially on the CCCS dataset.
We developed a method that can identify polarized public opinions by finding modules in a network of statistically related free word associations. Associations to the cue “migrant” were collected from two independent and comprehensive samples in Hungary (N1 = 505, N2 = 505). The co-occurrence-based relations of the free word associations reflected emotional similarity, and the modules of the association network were validated with well-established measures. The positive pole of the associations was gathered around the concept of “Refugees” who need help, whereas the negative pole associated asylum seekers with “Violence.” The results were relatively consistent in the two independent samples. We demonstrated that analyzing the modular organization of association networks can be a tool for identifying the most important dimensions of public opinion about a relevant social issue without using predefined constructs.
An important question that arises from autobiographical memory research is whether the variables that influence memory in the laboratory also drive memory for autobiographical episodes in real life. We explored this question within the context of e-mail communications and investigated the variables that influence recall for personally familiar names and temporal information in e-mails. We designed a Web-based program that analyzed each participant’s year-old sent e-mail archive and applied textual analysis algorithms to identify a set of sentences likely to be memorable. These sentences were then used as the stimuli in a cued recall task. Participants saw two sentences from their sent e-mail as a cue and attempted to recall the name of the e-mail recipient. Participants also rated the vividness of recall for the e-mail conversation and estimated the month in which they had written the e-mail. Linear mixed-effect analyses revealed that recipient name recall accuracy decreased with longer retention intervals and increased with greater frequency of contact with the recipient. Also, with longer retention intervals, participants dated e-mails as being more recent than their actual month. This telescoping error was moderately larger for e-mails with greater sentiment. These findings suggest that memory for personally familiar names and temporal information in e-mails closely follows the patterns for autobiographical memory and proper-name recall found in laboratory settings. This study introduces an innovative, Web-based experimental method for studying the cognitive processes related to autobiographical memories using ecologically valid, naturalistic communications.
Literary language is a style or form of language used in literary writing. The intent of this investigation is to disclose how and why literary writers foreground their texts and what meanings and effects are associated with foregrounding, deviation, creativity, Style and aesthetics on literature. This paper therefore appraised the characteristics of the language of literature, with a view to revealing the potency of creativity, style and aesthetics in some African and non-African poems and novels, which in turn portrays the skilfulness and dexterity of literary writers. Specifically, it examined the foregrounded parts of selected literary works; and to achieve this purpose, linguistic benchmarks were applied to these literary works. The descriptive system of data analysis, primary and secondary data collection methods and the foregrounding/deviation theory were employed. This survey therefore revealed that the literary genius contravenes the linguistic norms deliberately because he or she believes that the most proficient means of achieving distinction in writing is the use of distorted and strange forms. Thus, style heightens the language of literature to create a special effect and special meaning to the audience in order to arousing the interest and consciousness of the reader and society at large.Key Words: Deviation, Foregrounding, Literary Artist, Literature and Stylistics.
Mirror-sensory synesthetes mirror the pain or touch that they observe in other people on their own bodies. This type of synesthesia has been associated with enhanced empathy. We investigated whether the enhanced empathy of people with mirror-sensory synesthesia influences the experience of situations involving touch or pain and whether it affects their prosocial decision making. Mirror-sensory synesthetes (<i>N</i> = 18, all female), verified with a touch-interference paradigm, were compared with a similar number of age-matched control individuals (all female). Participants viewed arousing images depicting pain or touch; we recorded subjective valence and arousal ratings, and physiological responses, hypothesizing more extreme reactions in synesthetes. The subjective impact of positive and negative images was stronger in synesthetes than in control participants; the stronger the reported synesthesia, the more extreme the picture ratings. However, there was no evidence for differential physiological or hormonal responses to arousing pictures. Prosocial decision making was assessed with an economic game assessing altruism, in which participants had to divide money between themselves and a second player. Mirror-sensory synesthetes donated more money than non-synesthetes, showing enhanced prosocial behaviour, and also scored higher on the Interpersonal Reactivity Index as a measure of empathy. Our study demonstrates the subjective impact of mirror-sensory synesthesia and its stimulating influence on prosocial behaviour.This article is part of the discussion meeting issue ‘Bridging senses: new developments in synaesthesia’.
When using computer-aided translation systems in a typical, professional translation workflow, there are several stages at which there is room for improvement. The SCATE (Smart Computer-Aided Translation Environment) project investigated several of these aspects, both from a human-computer interaction point of view, as well as from a purely technological side. This paper describes the SCATE research with respect to improved fuzzy matching, parallel treebanks, the integration of translation memories with machine translation, quality estimation, terminology extraction from comparable texts, the use of speech recognition in the translation process, and human computer interaction and interface design for the professional translation environment. For each of these topics, we describe the experiments we performed and the conclusions drawn, providing an overview of the highlights of the entire SCATE project.
Classic natural language processing resources such as the Penn Treebank (Marcus et al. 1993) have long been used both as evaluation data for many linguistic tasks and as training data for a variety of off-the-shelf language processing tools. Recent work has highlighted a gender imbalance in the authors of this text data (Garimella et al. 2019) and hypothesized that tools created with such resources will privilege users from particular demographic groups (Hovy and Søgaard 2015). Domain adaptation is typically employed as a strategy in machine learning to adjust models trained and evaluated with data from different genres. However, the present work seeks to evaluate whether domain adaptation to demographic groups such as age or gender may be an effective strategy to ameliorate the effects of biased or outdated training corpora in linguistic preprocessing tasks. We find adaptation to demographic groups to be an effective strategy for improving preprocessing performance across all demographic groups.
Identity expression can be seen at either a personal or a social level. It can be shown in several ways, particularly through poetry. Therefore, this paper seeks to examine how identity is expressed modally in Darwish’s famous poem “Identity Card”. Leech’s (1974) theory of “seven types of meaning” is used as the theoretical framework for the study. As the study aims to figure out how modality is constructed when expressing identity in the poem, this linguistic norm has been used to track the mood and attitude of the poets in the composition of the stanzas. The findings of the study revealed that identity is expressed modally in the sensations that Mahmoud Darwish carries for the Palestinian, Arab, National, Cultural, Geographical and Historical identities. His language enacted the way he feels towards these six components of his belongings. Nationalism is seen as an important perspective in the affiliations of Darwish. The national identity is, however, the central concept of the poem. The modality provides a rough picture of what is going on in his mind as he experiencing the loss of land. It also comes in harmony with the general cultural context that contributes to his poetic experiences.
Machine translation (MT) is directly linked to its evaluation in order to both compare different MT system outputs and analyse system errors so that they can be addressed and corrected. As a consequence, MT evaluation has become increasingly important and popular in the last decade, leading to the development of MT evaluation metrics aiming at automatically assessing MT output. Most of these metrics use reference translations in order to compare system output, and the most well-known and widely spread work at lexical level. In this study we describe and present a linguistically-motivated metric, VERTa, which aims at using and combining a wide variety of linguistic features at lexical, morphological, syntactic and semantic level. Before designing and developing VERTa a qualitative linguistic analysis of data was performed so as to identify the linguistic phenomena that an MT metric must consider (Comelles et al. 2017). In the present study we introduce VERTa’s design and architecture and we report the experiments performed in order to develop the metric and to check the suitability and interaction of the linguistic information used. The experiments carried out go beyond traditional correlation scores and step towards a more qualitative approach based on linguistic analysis. Finally, in order to check the validity of the metric, an evaluation has been conducted comparing the metric’s performance to that of other well-known state-of-the-art MT metrics.
An important aspect of the perceived quality of vocal music is the degree to which the vocalist sings in tune. Although most listeners seem sensitive to vocal mistuning, little is known about the development of this perceptual ability or how it differs between listeners. Motivated by a lack of suitable preexisting measures, we introduce in this article an adaptive and ecologically valid test of mistuning perception ability. The stimulus material consisted of short excerpts (6 to 12 s in length) from pop music performances (obtained from MedleyDB; Bittner et al., 2014) for which the vocal track was pitch-shifted relative to the instrumental tracks. In a first experiment, 333 listeners were tested on a two-alternative forced choice task that tested discrimination between a pitch-shifted and an unaltered version of the same audio clip. Explanatory item response modeling was then used to calibrate an adaptive version of the test. A subsequent validation experiment applied this adaptive test to 66 participants with a broad range of musical expertise, producing evidence of the test’s reliability, convergent validity, and divergent validity. The test is ready to be deployed as an experimental tool and should make an important contribution to our understanding of the human ability to judge mistuning.
The recent rise in digitized historical text has made it possible to quantitatively study our psychological past. This involves understanding changes in what words meant, how words were used, and how these changes may have responded to changes in the environment, such as in healthcare, wealth disparity, and war. Here we make available a tool, the Macroscope, for studying historical changes in language over the last two centuries. The Macroscope uses over 155 billion words of historical text, which will grow as we include new historical corpora, and derives word properties from frequency-of-usage and co-occurrence patterns over time. Using co-occurrence patterns, the Macroscope can track changes in semantics, allowing researchers to identify semantically stable and unstable words in historical text and providing quantitative information about changes in a word’s valence, arousal, and concreteness, as well as information about new properties, such as semantic drift. The Macroscope provides information about both the local and global properties of words, as well as information about how these properties change over time, allowing researchers to visualize and download data in order to make inferences about historical psychology. Although quantitative historical psychology represents a largely new field of study, we see this work as complementing a wealth of other historical investigations, offering new insights and new approaches to understanding existing theory. The Macroscope is available online at http://www.macroscope.tech.
Neural parsers obtain state-of-the-art results on benchmark treebanks for constituency parsing -- but to what degree do they generalize to other domains? We present three results about the generalization of neural parsers in a zero-shot setting: training on trees from one corpus and evaluating on out-of-domain corpora. First, neural and non-neural parsers generalize comparably to new domains. Second, incorporating pre-trained encoder representations into neural parsers substantially improves their performance across all domains, but does not give a larger relative improvement for out-of-domain treebanks. Finally, despite the rich input representations they learn, neural parsers still benefit from structured output prediction of output trees, yielding higher exact match accuracy and stronger generalization both to larger text spans and to out-of-domain corpora. We analyze generalization on English and Chinese corpora, and in the process obtain state-of-the-art parsing results for the Brown, Genia, and English Web treebanks.
The history of legal lexis dates back to the ancient times of ancient peoples. The study of legal language enables the reconstruction of Indo-European ritual-legal ancients at verbal, linguistic levels. Archaic societies had no legal culture, instead, the norms of customary law of ancient societies were referred to as “pre-law”, which included syncretism of law, religion, myth, poetry, and morality. The syncretic ritual and legal consciousness of the ancient peoples in the pre-state period and in the early state formations has its specific reflection in a language that receives such a definition as “the language of law”. The system of “language of law” of Indo-European peoples is partly outlined in today’s scientific survey by describing changes in the semantics of legal lexis in the Indo-European languages, based on the analysis of the distinguished evolutionary models of semantics (EMS) in the Germanic, Slavonic and Iranian languages. The evolutionary model of semantics is a method of inquiry and a procedural scheme for explaining the history of legal meaning. 79 EMS were distinguished during the research, showing the genesis of the meaning 'power', 'lord', 'to rule', 'law', '(religious) law', 'pledge', '(blood) feud', 'court', 'judge'. Using data of the distinguished EMS, that clearly shows the change in the semantic volume of a word, a specific type of change in the meaning of legal lexis in the lexical and semantic system of the Indo-European languages was identified for each EMS, namely, expanding, narrowing (specializing), amelioration or pejoration of the meaning of the word. The study found that quantitatively the semantic derivation of the Indo-European legal terminology most experienced the type of narrowing of the meaning of the word, which, according to the researchers, belongs to the semantic universals. Metaphorical and metonymic changes in the meaning in the legal lexis of the Indo-European languages were also highlighted, that will need further study.
Identity expression can be seen at either a personal or a social level. It can be shown in several ways, particularly through poetry. Therefore, this paper seeks to examine how identity is expressed modally in Darwish’s famous poem “Identity Card”. Leech’s (1974) theory of “seven types of meaning” is used as the theoretical framework for the study. As the study aims to figure out how modality is constructed when expressing identity in the poem, this linguistic norm has been used to track the mood and attitude of the poets in the composition of the stanzas. The findings of the study revealed that identity is expressed modally in the sensations that Mahmoud Darwish carries for the Palestinian, Arab, National, Cultural, Geographical and Historical identities. His language enacted the way he feels towards these six components of his belongings. Nationalism is seen as an important perspective in the affiliations of Darwish. The national identity is, however, the central concept of the poem. The modality provides a rough picture of what is going on in his mind as he experiencing the loss of land. It also comes in harmony with the general cultural context that contributes to his poetic experiences.
SUD is an annotation scheme for syntactic dependency treebanks, near isomorphic to UD (Universal Dependencies). Contrary to UD, it is based on syntactic criteria (favoring functional heads) and the relations are defined on distributional and functional bases. In this paper, we will recall and specify the general principles underlying SUD, present the updated set of SUD relations, discuss the central question of MWEs, and introduce an orthogonal layer of deep-syntactic features converted from the deep-syntactic part of the UD scheme.
Eye tracking is a useful tool when studying the oscillatory eye movements associated with nystagmus. However, this oscillatory nature of nystagmus is problematic during calibration since it introduces uncertainty about where the person is actually looking. This renders comparisons between separate recordings unreliable. Still, the influence of the calibration protocol on eye movement data from people with nystagmus has not been thoroughly investigated. In this work, we propose a calibration method using Procrustes analysis in combination with an outlier correction algorithm, which is based on a model of the calibration data and on the geometry of the experimental setup. The proposed method is compared to previously used calibration polynomials in terms of accuracy, calibration plane distortion and waveform robustness. Six recordings of calibration data, validation data and optokinetic nystagmus data from people with nystagmus and seven recordings from a control group were included in the study. Fixation errors during the recording of calibration data from the healthy participants were introduced, simulating fixation errors caused by the oscillatory movements found in nystagmus data. The outlier correction algorithm improved the accuracy for all tested calibration methods. The accuracy and calibration plane distortion performance of the Procrustes analysis calibration method were similar to the top performing mapping functions for the simulated fixation errors. The performance in terms of waveform robustness was superior for the Procrustes analysis calibration compared to the other calibration methods. The overall performance of the Procrustes calibration methods was best for the datasets containing errors during the calibration.
Stack-augmented recurrent neural networks (RNNs) have been of interest to the deep learning community for some time. However, the difficulty of training memory models remains a problem obstructing the widespread use of such models. In this paper, we propose the Ordered Memory architecture. Inspired by Ordered Neurons (Shen et al., 2018), we introduce a new attention-based mechanism and use its cumulative probability to control the writing and erasing operation of the memory. We also introduce a new Gated Recursive Cell to compose lower-level representations into higher-level representation. We demonstrate that our model achieves strong performance on the logical inference task (Bowman et al., 2015) and the ListOps (Nangia and Bowman, 2018) task. We can also interpret the model to retrieve the induced tree structure, and find that these induced structures align with the ground truth. Finally, we evaluate our model on the Stanford Sentiment Treebank tasks (Socher et al., 2013), and find that it performs comparatively with the state-of-the-art methods in the literature.
We developed a method to automatically assess texts for features that help readers produce gist inferences. Following fuzzy-trace theory, we used a procedure in which participants recalled events under gist or verbatim instructions. Applying Coh-Metrix, we analyzed written responses in order to create gist inference scores (GISs), or seven variables converted to Z scores and averaged, which assess the potential for readers to form gist inferences from observable text characteristics. Coh-Metrix measures reflect referential cohesion and deep cohesion, which increase GIS because they facilitate coherent mental representations. Conversely, word concreteness, hypernymy for nouns and verbs (specificity), and imageability decrease GIS, because they promote verbatim representations. Also, the difference between abstract verb overlap among sentences (using latent semantic analysis) and more concrete verb overlap (using WordNet) should enhance coherent gist inferences, rather than verbatim memory for specific verbs. In the first study, gist condition responses scored nearly two standard deviations higher on GIS than did the verbatim condition responses. Predictions based on GIS were confirmed in two text analysis studies of 50 scientific journal article texts and 50 news articles and editorials. Texts from the Discussion sections of psychology journal articles scored significantly higher on GIS than did texts from the Method sections of the same journal articles. News reports also scored significantly lower than editorials on the same topics from the same news outlets. GIS proved better at discriminating among texts than did alternative formulae. In a behavioral experiment with closely matched text pairs, people randomly assigned to high-GIS versions scored significantly higher on knowledge and comprehension.
When a shift in writing style is noticed in a document, doubts arise about its originality. Based on this clue to plagiarism, the intrinsic approach to plagiarism detection identifies the stolen passages by analysing the writing style of the suspicious document without comparing it to textual resources that may serve as sources for the plagiarist. Character n-grams are recognised as a successful approach to modelling text for writing style analysis. Although prior studies have investigated the best practice of using character n-grams in authorship attribution and other problems, there is still a need for such investigations in the context of intrinsic plagiarism detection. Moreover, it has been assumed in previous works that the ways of using character n-grams in authorship attribution remain the same for intrinsic plagiarism detection. In this paper, we study the effect of character n-grams frequency and length on the performance of intrinsic plagiarism detection. Our experiments utilise two state-of-the-art methods and five large document collections of PAN labs written in English and Arabic. We demonstrate empirically that the low- and the high-frequency n-grams are not equally relevant for intrinsic plagiarism detection, but their performance depends on the way they are exploited.
In this paper we present a pipeline for the detection of spelling variants, i.e., different spellings that represent the same word, in non-standard texts. For example, in Middle Low German texts in and ihn (among others) are potential spellings of a single word, the personal pronoun ‘him’. Spelling variation is usually addressed by normalization, in which non-standard variants are mapped to a corresponding standard variant, e.g. the Modern German word ihn in the case of in. However, the approach to spelling variant detection presented here does not need such a reference to a standard variant and can therefore be applied to data for which a standard variant is missing. The pipeline we present first generates spelling variants for a given word using rewrite rules and surface similarity. Afterwards, the generated types are filtered. We present a new filter that works on the token level, i.e., taking the context of a word into account. Through this mechanism ambiguities on the type level can be resolved. For instance, the Middle Low German word in can not only be the personal pronoun ‘him’, but also the preposition ‘in’, and each of these has different variants. The detected spelling variants can be used in two settings for Digital Humanities research: On the one hand, they can be used to facilitate searching in non-standard texts. On the other hand, they can be used to improve the performance of natural language processing tools on the data by reducing the number of unknown words. To evaluate the utility of the pipeline in both applications, we present two evaluation settings and evaluate the pipeline on Middle Low German texts. We were able to improve the F1 score compared with previous work from \(0.39\) to \(0.52\) for the search setting and from \(0.23\) to \(0.30\) when detecting spelling variants of unknown words.
This paper investigates the historical (1850s–2000s) evolution of semantics in the English language using contemporaneous, decade-specific computational estimates of word concreteness. Study 1 describes the computational method of generating time-locked estimates of concreteness based on the Corpus of Historic American English, and makes available the computed scores for 25,000 English words over 15 decades. We also report several tests of reliability and validity, demonstrating that our historical concreteness scores have high levels of both. Study 2 uses concreteness scores to revisit findings of studies that use a static set of contemporary human concreteness norms to examine historical trends of semantic change. Specifically, we observed (contra Hills & Adelman, (Cognition, 143, 87–92 2015)) that distinct word types of the English language become increasingly more concrete over time and (in line with Hills & Adelman, (Cognition, 143, 87–92 2015) & Hills, Adelman & Noguchi, (The Quarterly Journal of Experimental Psychology, 70(8), 1603–1619 2016)) that relatively concrete words tend to be used more often than abstract ones. We discuss both contrastive and corroborative claims in light of recent work on semantic evolution and argue for the use of time-locked computed estimates over static human norms when examining diachronic linguistic phenomena.
Test publishers usually provide confidence intervals (CIs) for normed test scores that reflect the uncertainty due to the unreliability of the tests. The uncertainty due to sampling variability in the norming phase is ignored. To express uncertainty due to norming, we propose a flexible method that is applicable in continuous norming and allows for a variety of score distributions, using Generalized Additive Models for Location, Scale, and Shape (GAMLSS; Rigby & Stasinopoulos, 2005). We assessed the performance of this method in a simulation study, by examining the quality of the resulting CIs. We varied the population model, procedure of estimating the CI, confidence level, sample size, value of the predictor, extremity of the test score, and type of variance-covariance matrix. The results showed that good quality of the CIs could be achieved in most conditions. The method is illustrated using normative data of the SON-R 6-40 test. We recommend test developers to use this approach to arrive at CIs, and thus properly express the uncertainty due to norm sampling fluctuations, in the context of continuous norming. Adopting this approach will help (e.g., clinical) practitioners to obtain a fair picture of the person assessed.
The smart home brings together devices, the cloud, data, and people to make home living more comfortable and safer. Trigger-action programming enables users to connect smart devices using if-this-then-that (IFTTT)-style rules. With the increasing number of devices in smart home systems, multiple running rules that act on actuators in contradictory ways may cause unexpected and unpredictable interference problems, which can put residents and their belongings at risk. Previous studies have considered explicit interference problems related to multiple rules targeting a single actuator, whereas implicit interference (interference across different actuators) detection is still challenging and not yet well studied owing to the effort-intensive and time-consuming annotation work of obtaining device information. The lack of knowledge about devices is a critical reason that affects the accuracy and efficiency in implicit interference detection. In this article, we propose A3ID, an automatic detection method for implicit interference based on knowledge graphs. Using natural language processing (NLP) techniques and a lexical database, A3ID can extract knowledge of devices from a knowledge graph, including functionality, effect, and scope. Then, it analyzes and detects interferences among the different devices semantically in three steps, without human intervention. Furthermore, it provides user-friendly explanations in a well-designed structure to specify possible reasons for the implicit interference problems. Our experiment on 11 859 IFTTT-style rules shows that A3ID outperforms state-of-the-art methods by more than 33% in the F1-score for the detection of implicit interference. Moreover, evaluations on an extended data set for devices from ConceptNet (a knowledge graph) and five smart home systems suggest that A3ID also has favorable performance with other devices not limited to the smart home domain.
Recurrent Neural Networks (RNN), Long Short-Term Memory Networks (LSTM), and Memory Networks which contain memory are popularly used to learn patterns in sequential data. Sequential data has long sequences that hold relationships. RNN can handle long sequences but suffers from the vanishing and exploding gradient problems. While LSTM and other memory networks address this problem, they are not capable of handling long sequences (50 or more data points long sequence patterns). Language modelling requiring learning from longer sequences are affected by the need for more information in memory. This paper introduces Long Term Memory network (LTM), which can tackle the exploding and vanishing gradient problems and handles long sequences without forgetting. LTM is designed to scale data in the memory and gives a higher weight to the input in the sequence. LTM avoid overfitting by scaling the cell state after achieving the optimal results. The LTM is tested on Penn treebank dataset, and Text8 dataset and LTM achieves test perplexities of 83 and 82 respectively. 650 LTM cells achieved a test perplexity of 67 for Penn treebank, and 600 cells achieved a test perplexity of 77 for Text8. LTM achieves state of the art results by only using ten hidden LTM cells for both datasets.