Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Chapter 2 provides a sociohistorical analysis of the evolution of Yungueño Spanish, Chota Valley Spanish and Chincha Spanish. The chapter illustrates general aspects of the African Diaspora to the Americas and its specific linguistic consequences in Yungas (Bolivia), Chota Valley (Ecuador) and Chincha (Peru). Given the historical evidence available for these Afro-Hispanic Languages of the Americas, I propose that these contact varieties developed in isolated rural villages, not subject to the social pressures imposed by education, standardization and the linguistic norm. In such a context, advanced SLA processes could be nativized and conventionalized at the local level, thus crystallizing in the L1 varieties spoken by subsequent generations of these Afro-Andean communities.
The PapyGreek Treebanks dataset contains documentary texts written in Postclassical Greek (ca. 300 BCE–700 CE), morphosyntactically annotated according to Dependency Grammar. The source of the texts is the Duke Databank of Documentary Papyri (DDbDP), which preserves the modern editorial treatment of the documents in TEI Epidoc XML encoding. Aiming to expose linguistic variation in the DDbDP, we have annotated two versions of a selection of documents: the plain transcription and an editorially corrected version. The dataset also comprises metadata about the documents’ dating and provenance, text type, and the persons involved. Furthermore, it facilitates linguistic research on these texts.
This paper describes the construction and annotation of the Late Latin Charter Treebank, a set of three dependency treebanks (llct1, llct2 and llct3) which together contain 1,261 Early Medieval Latin documentary texts (i.e., original charters) written in Italy between ad 714 and 1000 (about 594,000 tokens). The paper focusses on matters which a linguistically or philologically inclined user of llct needs to know: the criteria on which the charters were selected, the special characteristics of the annotation types utilised, and the geographical and chronological distribution of the data. In addition to normal queries on forms, lemmas, morphology and syntax, complex philological research settings are enabled by the textual annotation layer of llct, which indicates abbreviated and damaged words, as well as the formulaic and non-formulaic passages of each charter.
Abstract The present study demonstrates that the process of linguistic Romanization, i.e. Latinization of the Roman Empire, is traceable by the data of the Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age (LLDB). A multi-level analysis of linguistic and non-linguistic data in the LLDB has shown that Latinization, i.e. the spread of spoken or vulgar Latin, became more and more intensive over time in all concerned provinces (i.e. Lusitania, Gallia Narbonensis, Venetia et Histria, Dalmatia, Moesia, Pannonia, and Britannia), although to a varying degree in each. What is more, in many aspects of the investigation, it was possible to find differences between the selected provinces of the Roman Empire corresponding mostly to the future Romance (both negative and positive) outcomes of the respective areas. All in all, the analysis of data of the LLDB database can contribute to solving the complex problem of Latinization, and is a lot more appropriate for this purpose than a simple comparative analysis of epigraphic corpora of the selected provinces.
This article introduces the working methods of the Parsed Historical Corpus of the Welsh Language (PARSHCWL). The corpus is designed to provide researchers with a tool for automatic exhaustive extraction of instances of grammatical structures from Middle and Modern Welsh texts in a way comparable to similar tools that already exist for various European languages. The major features of the corpus are outlined, along with the overall architecture of the workflow needed for a team of researchers to produce it. In this paper, the two first stages of the process, namely pre-processing of texts and automated part-of-speech (POS) tagging are discussed in some detail, focusing in particular on major issues involved in defining word boundaries and in defining a robust and useful tagset.
This paper describes the grammatical patterning of two parts of speech – nouns and adjectives – included in the corpus-driven “Lexical Database of Lithuanian” as a foreign language. The lexical database is a lexicographic application of the Lithuanian Pedagogic Corpus (approx. 620.000 tokens) which was used to develop headword lists and to collect word usage information in the form of corpus patterns. In this project, we adopted a partially automated inductive procedure of Corpus Pattern Analysis for 207 verbs, 386 nouns, 87 adjectives, and 41 adverbs. The detected corpus patterns reflect different meanings of the headword. Each pattern presents information on grammatical, semantic, and lexical levels. Manually selected examples illustrate all pattern components. In this paper, 673 patterns with nouns and 99 patterns with adjectives will be analysed discussing their syntactic behaviour in detail and providing some comments on lexis-grammar interface. The majority of patterns with nouns and adjectives are minimal patterns which include only the closest syntactical partners. This result is influenced by different procedures used to describe patterns with nouns, adjectives, and adverbs and patterns with verbs. Due to rich grammatical information, there are several similar patterns with one main (usually the most frequent) type and its variants. Pattern variants show that the grammatical characteristics of a specific word usage are rather individual.
Recent work on multilingual dependency parsing focused on developing highly multilingual parsers that can be applied to a wide range of low-resource languages. In this work, we substantially outperform such "one model to rule them all" approach with a heuristic selection of languages and treebanks on which to train the parser for a specific target language. Our approach, dubbed TOWER, first hierarchically clusters all Universal Dependencies languages based on their mutual syntactic similarity computed from human-coded URIEL vectors. For each low-resource target language, we then climb this language hierarchy starting from the leaf node of that language and heuristically choose the hierarchy level at which to collect training treebanks. This treebank selection heuristic is based on: (i) the aggregate size of all treebanks subsumed by the hierarchy level and (ii) the similarity of the languages in the training sample with the target language. For languages without development treebanks, we additionally use (ii) for model selection (i.e., early stopping) in order to prevent overfitting to development treebanks of closest languages. Our TOWER approach shows substantial gains for low-resource languages over two state-ofthe-art multilingual parsers, with more than 20 LAS point gains for some of those languages. Parsing models and code available at: https: //github.com/codogogo/towerparse.
Treebanks are valuable linguistic resources that include the syntactic structure of a language sentence in addition to part-of-speech tags and morphological features. They are mainly utilized in modeling statistical parsers. Although the statistical natural language parser has recently become more accurate for languages such as English, those for the Arabic language still have low accuracy. The purpose of this article is to construct a new Arabic dependency treebank based on the traditional Arabic grammatical theory and the characteristics of the Arabic language, to investigate their effects on the accuracy of statistical parsers. The proposed Arabic dependency treebank, called I3rab, contrasts with existing Arabic dependency treebanks in two main concepts. The first concept is the approach of determining the main word of the sentence, and the second concept is the representation of the joined and covert pronouns. To evaluate I3rab, we compared its performance against a subset of Prague Arabic Dependency Treebank that shares a comparable level of details. The conducted experiments show that the percentage improvement reached up to 10.24% in UAS and 18.42% in LAS.
Arabic dependency parsers have a poor performance compared to parsers of other languages. Recently the impact of annotation at lexical level of dependency treebank on the overall performance of the dependency parses has been extensively investigated. This paper focuses on the impact of coarse-grained and fine-grained dependency relations on the performance of Arabic dependency parsers. Moreover, this paper introduces the annotation rules for I3rab dependency treebank. Experimentally, the obtained results showed that having an appropriate set of dependency relations improves the performance of an Arabic dependency parser up to 27.55%.
Towards explainable affective computing (XAC), researchers have invested considerable effort into post hoc approaches and reverse engineering to seek explanations for deep learning models. However, alternative, intrinsic approaches that aim to build inherently interpretable models by restricting their complexity are yet to be widely explored. In this study, we integrate an explanatory polytomous item response model that provides a well-established psychological interpretation for ordinal scales with deep neural networks to realize high prediction performance and good result interpretability. We conducted an experiment on a growing task (i.e., predicting the idiosyncratic perception of emotional faces of an individual); as expected theoretically, the topmost parameters of our model demonstrated strong correlations with those of the corresponding ordinal item response model: r = 0.928 to 1.00. Our proposed intrinsic approach can used as a complementary framework for post-hoc methods in XAC to coach and support human social interactions.
Recursive Deep Models have been used as powerful models to learn \ncompositional representations of text for many natural language processing tasks. \nHowever, they require structured input (i.e. sentiment treebank) to encode sentences \nbased on their tree-based structure to enable them to learn latent semantics \nof words using recursive composition functions. In this paper, we present our \ncontributions and efforts for the Turkish Sentiment Treebank construction. We \nintroduce MS-TR, a Morphologically Enriched Sentiment Treebank, which was \nimplemented for training Recursive Deep Models to address compositional sentiment \nanalysis for Turkish, which is one of the well-known Morphologically Rich \nLanguage (MRL). We propose a semi-supervised automatic annotation, as a distantsupervision \napproach, using morphological features of words to infer the polarity of \nthe inner nodes of MS-TR as positive and negative. The proposed annotation model \nhas four different annotation levels: morph-level, stem-level, token-level, and \nreview-level. Each annotation level’s contribution was tested using three different \ndomain datasets, including product reviews, movie reviews, and the Turkish Natural \nCorpus essays. Comparative results were obtained with the Recursive Neural Tensor Networks (RNTN) model which is operated over MS-TR, and conventional machine learning methods. Experiments proved that RNTN outperformed the baseline methods and achieved much better accuracy results compared to the baseline methods, which cannot accurately capture the aggregated sentiment information.
The ventromedial and dorsolateral prefrontal cortex are two major prefrontal regions that usually interact in serving different cognitive functions. On the other hand, these regions are also involved in cognitive processing of emotions but their contribution to emotional processing is not well-studied. In the present study, we investigated the role of these regions in three dimensions (valence, arousal and dominance) of emotional processing of stimuli via ratings of visual stimuli performed by the study participants on these dimensions. Twenty- two healthy adult participants (mean age 25.21 ± 3.84 years) were recruited and received anodal and sham transcranial direct current stimulation (tDCS) (1.5 mA, 15 min) over the dorsolateral prefrontal cortex (dlPFC) and and ventromedial prefrontal cortex (vmPFC) in three separate sessions with an at least 72-h interval. During stimulation, participants underwent an emotional task in each stimulation condition. The task included 100 visual stimuli and participants were asked to rate them with respect to valence, arousal, and dominance. Results show a significant effect of stimulation condition on different aspects of emotional processing. Specifically, anodal tDCS over the dlPFC significantly reduced valence attribution for positive pictures. In contrast, anodal tDCS over the vmPFC significantly reduced arousal ratings. Dominance ratings were not affected by the intervention. Our results suggest that the dlPFC is involved in control and regulation of valence of emotional experiences, while the vmPFC might be involved in the extinction of arousal caused by emotional stimuli. Our findings implicate dimension-specific processing of emotions by different prefrontal areas which has implications for disorders characterized by emotional disturbances such as anxiety or mood disorders.
International audience
Recent advances in deep learning techniques have enabled machines to generate cohesive open-ended text when prompted with a sequence of words as context. While these models now empower many downstream applications from conversation bots to automatic storytelling, they have been shown to generate texts that exhibit social biases. To systematically study and benchmark social biases in open-ended language generation, we introduce the Bias in Open-Ended Language Generation Dataset (BOLD), a large-scale dataset that consists of 23,679 English text generation prompts for bias benchmarking across five domains: profession, gender, race, religion, and political ideology. We also propose new automated metrics for toxicity, psycholinguistic norms, and text gender polarity to measure social biases in open-ended text generation from multiple angles. An examination of text generated from three popular language models reveals that the majority of these models exhibit a larger social bias than human-written Wikipedia text across all domains. With these results we highlight the need to benchmark biases in open-ended language generation and caution users of language generation models on downstream tasks to be cognizant of these embedded prejudices.
The COVID-19 pandemic has dramatically changed the nature of our social interactions. In order to understand how protective equipment and distancing measures influence the ability to comprehend others' emotions and, thus, to effectively interact with others, we carried out an online study across the Italian population during the first pandemic peak. Participants were shown static facial expressions (Angry, Happy and Neutral) covered by a sanitary mask or by a scarf. They were asked to evaluate the expressed emotions as well as to assess the degree to which one would adopt physical and social distancing measures for each stimulus. Results demonstrate that, despite the covering of the lower-face, participants correctly recognized the facial expressions of emotions with a polarizing effect on emotional valence ratings found in females. Noticeably, while females' ratings for physical and social distancing were driven by the emotional content of the stimuli, males were influenced by the "covered" condition. The results also show the impact of the pandemic on anxiety and fear experienced by participants. Taken together, our results offer novel insights on the impact of the COVID-19 pandemic on social interactions, providing a deeper understanding of the way people react to different kinds of protective face covering.
OBJECTIVE: The study sought to develop and evaluate neural natural language processing (NLP) packages for the syntactic analysis and named entity recognition of biomedical and clinical English text. MATERIALS AND METHODS: We implement and train biomedical and clinical English NLP pipelines by extending the widely used Stanza library originally designed for general NLP tasks. Our models are trained with a mix of public datasets such as the CRAFT treebank as well as with a private corpus of radiology reports annotated with 5 radiology-domain entities. The resulting pipelines are fully based on neural networks, and are able to perform tokenization, part-of-speech tagging, lemmatization, dependency parsing, and named entity recognition for both biomedical and clinical text. We compare our systems against popular open-source NLP libraries such as CoreNLP and scispaCy, state-of-the-art models such as the BioBERT models, and winning systems from the BioNLP CRAFT shared task. RESULTS: For syntactic analysis, our systems achieve much better performance compared with the released scispaCy models and CoreNLP models retrained on the same treebanks, and are on par with the winning system from the CRAFT shared task. For NER, our systems substantially outperform scispaCy, and are better or on par with the state-of-the-art performance from BioBERT, while being much more computationally efficient. CONCLUSIONS: We introduce biomedical and clinical NLP packages built for the Stanza library. These packages offer performance that is similar to the state of the art, and are also optimized for ease of use. To facilitate research, we make all our models publicly available. We also provide an online demonstration (http://stanza.run/bio).
Abstract This article contributes to a dialogue between childhood studies and the sociolinguistic subfield ‘Family Language Policy’ (‘FLP’). The article argues that the two fields provide complementary vantage points for exploring child agency. It explains a revised version of a model I developed to conceptualise child agency in FLP, consisting of four intersecting dimensions: compliance regimes; linguistic norms; linguistic competence and generational positioning (Smith‐Christmas, Handbook of home language maintenance and development. De Gruyter Mouton, pp. 218–235, 2020a). The article examines two conversational excerpts as a means to illustrating the dynamic and relational nature of child agency and how it is both shaped by as well as shapes interactional practices over time and space.
The therapeutic effect of antidepressants has been demonstrated for anhedonia in patients with depression. However, antidepressants may cause side-effects, such as cardiovascular dysfunction. Although physical activity has minor side-effects, it may serve as an alternative for improving anhedonia and depression. We sought to investigate whether physical activity reduces the level of anhedonia in individuals with depression. Fifty-six university students with moderate depressive symptoms (Beck Depression Inventory total score > 16) were divided into three training groups: the Running Group (RG, n = 19), the Stretching Group (SG, n = 19), and the Control Group (n = 18). We employed the Monetary Incentive Delay (MID) task and the Temporal Experience of Pleasure Scale (TEPS) to evaluate hedonic capacity. All participants in the RG and SG received 8 weeks of jogging and stretching training, respectively. The RG experienced an increase in the level of arousal during anticipation of a future reward and recalled less negativity towards the loss condition. The SG exhibited enhanced scores on the Anticipatory and Consummatory Pleasure subscales of the TEPS after training. Moreover, in the RG, greater improvements in anticipatory arousal ratings for pleasure and remembered valence ratings for negative affect were associated with longer training duration, lower maximum heart rate, and higher consumed calories during training. To conclude, physical activity is effective in improving anticipatory anhedonia in individuals with depressive symptoms.
Semantics is a research field that has gained an extensive interest recently. This survey describes recent works in the field of semantics, a part of the broader area of computational linguistics. One of the important aspects of computational linguistics is using proper methods to distribute semantics for obtaining representations of the meaning of words. This survey summarizes the latest state of the art approaches in semantics that use deep learning methods, datasets, and lexical databases, specifying semantics under two categories such as semantic similarity and sentence modeling.
Popular Geopolitics and the Conceptualization of Linguistic Norm CentresInger Schoonderbeek Hansen & Yonatan Goldshtein Abstract This article challenges the prevalent idea of Copenhagen as the only linguistic norm centre in Denmark. It is based on an experimental study of language attitudes conducted in Salling, a peninsula in North-Western Jutland. The experiment consisted of three tasks […]
Many philosophers working today on the normativity of language have concluded that linguistic activity is not a matter of rule-following. These conversations have been framed by a conception of linguistic normativity with roots in Wittgenstein and Kripke, however, and in this chapter, I use conceptual resources developed by the classical American pragmatists and their descendants to argue that punctate linguistic acts are governed by rules in a sense that has been neglected in the literature. In doing so, I show that this work draws on themes from German idealism discussed in the Introduction. I also argue that, in order to account for the development and propagation of (proto-) linguistic norms within hominid communities, our ancestors’ capacity for shared practical picturing – discussed in chapter 2 as a physiological basis for shared intentionality – gave rise to a capacity for deontic picturing, which involves the exercise of affective and reactive evaluative attitudes across different points of view within a community. In closing out part 1 of the book, chapter 3 thereby both clears the ground for using the ability to speak a rule-governed language as a basis for constructing an account of discursive cognition – beginning in chapter 4 ’s formal semantics for the deontic modalities – and reinforces my claim that this project has its roots in American appropriations of German idealism directed at making naturalistic sense of our existence as norm-governed rational animals.
The design of widespread vision-and-language datasets and pre-trained encoders directly adopts, or draws inspiration from, the concepts and images of ImageNet. While one can hardly overestimate how much this benchmark contributed to progress in computer vision, it is mostly derived from lexical databases and image queries in English, resulting in source material with a North American or Western European bias. Therefore, we devise a new protocol to construct an ImageNet-style hierarchy representative of more languages and cultures. In particular, we let the selection of both concepts and images be entirely driven by native speakers, rather than scraping them automatically. Specifically, we focus on a typologically diverse set of languages, namely, Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish. On top of the concepts and images obtained through this new protocol, we create a multilingual dataset for Multicultural Reasoning over Vision and Language (MaRVL) by eliciting statements from native speaker annotators about pairs of images. The task consists of discriminating whether each grounded statement is true or false. We establish a series of baselines using state-of-the-art models and find that their cross-lingual transfer performance lags dramatically behind supervised performance in English. These results invite us to reassess the robustness and accuracy of current state-of-the-art models beyond a narrow domain, but also open up new exciting challenges for the development of truly multilingual and multicultural systems.
The DiGreC (DIachrony of GREek Case) treebank is a corpus of selected sentences from Greek texts, ranging from Homer to Modern Greek, which have been annotated morphosyntactically and semantically. The corpus comprises excerpts from 655 texts, for a total of 3385 sentences and 56,440 word tokens; automated tagging and lemmatisation has been supplemented with manual review to ensure accuracy. The data exist in xml and csv formats, which can be manipulated and converted automatically to other schemata. A web site has also been created to allow users to interact with the data more easily, and to provide specialised functionality for searching and visualisation. This corpus was created to inform theoretical debates regarding the role of case in grammar, and may be of use to researchers searching for specific attestations of a range of different constructions in Greek.
Although social media appear to be welcoming spaces that enable easy access to target-language communities, second language (L2) participation is not necessarily full and equitable. Drawing on computer-mediated discourse analysis (Herring, 2007) and critical discourse analysis (Wodak & Meyer, 2009), this analysis of a discussion forum on the social media platform Reddit uses social positioning theory (Harré, 2012; see also Debray & Spencer-Oatey, 2019) to show how L2 errors are construed as obstacles to full participation. I argue that linguistic gatekeeping is linked to community norms that reproduce language ideologies, affirm the authority of the idealized native speaker, and position L2 participants as L2 learners rather than L2 users. When L2 users cannot participate fully in what seem to be welcoming spaces they may exclude themselves. At the same time, the data also provide compelling evidence that Reddit offers a new mode of inclusion for L2 users.
Study 1 Data - self-and other-affect ratings of younger and older adult dyads<br>
How do thoughts arise, unfold, and change over time? Are the contents and dynamics of everyday thought rooted in conceptual associations within one's semantic networks? To address these questions, we developed the Free Association Semantic task (FAST), whereby participants generate dynamic chains of conceptual associations in response to seed words that vary in valence. Ninety-four adults from a community sample completed the FAST task and additionally described and rated six of their most frequently occurring everyday thoughts. Text analysis and valence ratings revealed similarities in thematic and affective content between FAST concept chains and recurrent autobiographical thoughts. Dynamic analyses revealed that individuals higher in rumination were more strongly attracted to negative conceptual spaces and more likely to remain there longer. Overall, these findings provide quantitative evidence that conceptual associations may act as a semantic scaffold for more complex everyday thoughts, and that more negative and less dynamic conceptual associations in ruminative individuals mirror maladaptive repetitive thoughts in daily life. (PsycInfo Database Record (c) 2022 APA, all rights reserved).
The model of lexicographic description of English, Italian and Kazakh language figurative means explicating metaphorical rethinking of the gastronomic sphere phenomena in the digital “Multilingual Dictionary of Metaphors” is presented.
Anxiety patients over-generalize fear, possibly because of an incapacity to discriminate threat and safety signals. Discrimination trainings are promising approaches for reducing such fear over-generalization. Here we investigated the efficacy of a fear-relevant vs. a fear-irrelevant discrimination training on fear generalization and whether the effects are increased with feedback during training. Eighty participants underwent two fear acquisition blocks, during which one face (conditioned stimulus, CS+), but not another face (CS-), was associated with a female scream (unconditioned stimulus, US). During two generalization blocks, both CSs plus four morphs (generalization stimuli, GS1-GS4) were presented. Between these generalization blocks, half of the participants underwent a fear-relevant discrimination training (discrimination between CS+ and the other faces) with or without feedback and the other half a fear-irrelevant discrimination training (discrimination between the width of lines) with or without feedback. US expectancy, arousal, valence ratings, and skin conductance responses (SCR) indicated successful fear acquisition. Importantly, fear-relevant vs. fear-irrelevant discrimination trainings and feedback vs. no feedback reduced generalization as reflected in US expectancy ratings independently from one another. No effects of training condition were found for arousal and valence ratings or SCR. In summary, this is a first indication that fear-relevant discrimination training and feedback can improve the discrimination between threat and safety signals in healthy individuals, at least for learning-related evaluations, but not evaluations of valence or (physiological) arousal.
Human beings have a fundamental need to belong. Evaluating and dealing with social exclusion and social inclusion events, which represent negative and positive social interactions, respectively, are closely linked to our physical and mental health. In addition to traditional paradigms that simulate scenarios of social interaction, images are utilized as effective visual stimuli for research on socio-emotional processing and regulation. Since the current mainstream emotional image database lacks social stimuli based on a specific social context, we introduced an open-access image database of social inclusion/exclusion in young Asian adults (ISIEA). This database contains a set of 164 images depicting social interaction scenarios under three categories of social contexts (social exclusion, social neutral, and social inclusion). All images were normatively rated on valence, arousal, inclusion score, and vicarious feeling by 150 participants in Study 1. We additionally examined the relationships between image ratings and the potential factors influencing ratings. The importance of facial expression and social context in the image rating of ISIEA was examined in Study 2. We believe that this database allows researchers to select appropriate materials for socially related studies and to flexibly conduct experimental control.
Intranasal oxytocin exerts wide-ranging effects on socioemotional behavior and is proposed as a potential therapeutic intervention in psychiatric disorders. However, following intranasal administration, oxytocin could penetrate directly into the brain or influence its activity via increased peripheral concentrations crossing the blood-brain barrier or influencing vagal projections. In the current randomized, placebo-controlled, pharmaco-imaging clinical trial we investigated effects of 24IU oral (lingual) oxytocin spray, restricting it to peripherally mediated blood-borne and vagal effects, on responses to face emotions in 80 male subjects and compared them with 138 subjects treated intranasally with 24IU. Oral, but not intranasal oxytocin administration increased both arousal ratings for faces and associated brain reward responses, the latter being partially mediated by blood concentration changes. Furthermore, while oral oxytocin increased amygdala and arousal responses to face emotions, after intranasal administration they were decreased. Thus, oxytocin can produce markedly contrasting motivational effects in relation to socioemotional cues when it influences brain function via different routes. These findings have important implications for future therapeutic use since administering oxytocin orally may be both easier and have potentially stronger beneficial effects by enhancing responses to emotional cues and increasing their associated reward.
Abstract The annotation scheme of dependency treebanks might have an impact on the results of linguistic analysis, thus leading to different interpretations of linguistic phenomena. This study compares the results of two widely used dependency measures, i.e., dependency direction and dependency distance, based on 18 parallel Universal Dependencies (UD) annotated treebanks and 18 corresponding Surface‐Syntactic Universal Dependencies (SUD) annotated treebanks. The results show that (1) Based on the semantic UD and syntactic SUD, dependency relations between function words and content words share the opposite dependency directions but similar dependency distances; (2) Annotation scheme has a significant impact on dependency direction, though the effect size is small. We find that the proportions of head‐final dependencies based on the syntactic SUD can better group language families than those based on semantic UD; (3) Annotation scheme also affects dependency distance significantly, though its effect size is small. Mean dependency distances (MDDs) based on UD are always higher than those based on SUD. However, the MDDs based on both annotation schemes are within a certain threshold, which shows that the linguistic universal of dependency distance minimization is independent of annotation schemes.
The prescriptive approach has been prevalent in discussions about the linguistic norm for many decades. Many linguists question the primacy of social custom and make many arbitrary changes to establish the subjective form of the norm. In connection with the planned The Dictionary of Proper Uses of Languagethe author of the article presents the best structuralist traditions and calls for research on the linguistic norm which is based on descriptive methods. It is necessary to completely break away from all manifestations of arbitrariness and subjectivity in contemporary prescriptive linguistics. The fundamental premise that the linguistic norm is a fact based on usus must be reflected in relevant procedures aimed at analyzing corpora consisting of millions of words. Such an approach will make it possible to establish a model that comprises more than just individual language uses. As far as dictionary definitions are concerned, the most frequent, widespread and thus typical linguistic units should be primarily considered to be normative. Typicality, determined by frequency, as well as textual, social and territorial conditions, is the most important category.
Using a very large lexical database and generalized additive modeling, this article reveals that labial-velar (LV) stops are marginal phonemes in many of the languages of Northern Sub-Saharan Africa that have them, and that the languages in which they are not marginal are grouped into three compact zones of high lexical LV frequency. The resulting picture allows us to formulate precise hypotheses about the spread of the Niger-Congo and Central Sudanic languages and about the origins of the linguistic area known as the Sudanic zone or Macro-Sudan belt. It shows that LV stops are a substrate feature that should not be reconstructed into the early stages of the languages that currently have them. We illustrate the implications of our findings for linguistic prehistory with a short discussion of the Bantu expansion. Our data also indirectly confirm the hypothesis that LV stops are more recurrent in expressive parts of the vocabulary, and we argue that this has a common explanation with the well-known fact that they tend to be restricted to stem-initial position in what we call C-emphasis prosody.
This paper develops the concept of word order universals based on a data analysis of the Universal Dependencies project, which proposes treebanks of more than 90 languages encoded with the same annotation scheme. The nature of the data we work on allows us to extract rich details for testing well-known typological implicational universals and, further, explore new kinds of universals that we call quantitative universals. We show how such quantitative universals are in essence different from implicational universals, including statistical universals, by the fact that they no longer lay down any claims on categorical statements, but rather on continuous parameters, opening a new field of research we propose to call typometrics.
The article explores the use of contextual slang as linguistic performance by three all-female friendship groups in Calabar metropolis, Cross River State, south-eastern Nigeria. I argue that slang constitutes critical components of the discursive practices of young urban Nigerian women in maintaining friendship and deviating from stereotyped cultural and linguistic norms. Drawing insights from the analytical tools of African feminism and linguistic ideology, the article discusses recurrent themes in young women’s contextual slanguage and the motivations for the use of these creative linguistic and cultural resources in defining participants’ authentic social selves and in enacting their different modes of belonging. Qualitative ethnographic data for the study were sourced from participant observations, semi-structured interviews and informal conversations with 30 participants. The study concludes that young urban women utilise contextual slang as indexical tools in their everyday narratives to negotiate meaning in relation to the experience of their social lives, to acculturate to male linguistic norms and to affiliate with ideologies that represent gendered identity.
Manually annotated corpus is a perquisite for several natural language processing applications including parsing. Nevertheless, annotated corpus is not always available for resource-poor languages, especially when domain under consideration is noisy user-generated data found on social media platforms such as Twitter. To overcome this deficiency of hand-annotated corpus, researchers have focused their attention on semi-automatic corpus annotation methods. This paper describes the experiments carried out using semi-automatic methods like self-training and co-training in an attempt for creating silver-standard dependency treebank of Urdu tweets. Six iterations of each approach were performed using same experimental conditions using MaltParser and Parsito parser, both statistical data driven parsers. For self-training experiments, the best performing MaltParser model was trained on 1250 Urdu tweets, with an accuracy of 70.2% LA, 74.4% UAS, 63% LAS. Whereas the best performing Parsito model was also trained on 1250 Urdu tweets with an accuracy of 70.8% LA, 74.8% UAS, 63.4% LAS. For co-training experiments, best performing MaltParser model was trained on 1500 Urdu tweets, with an accuracy of 70.5% LA, 74.4% UAS, 63.2% LAS. The best performing Parsito model was also trained on 1500 Urdu tweets with an accuracy of 70.5% LA, 74.3% UAS, 63% LAS. Although, there was not much difference between the results of both approaches, co-training results were slightly better for both parsers and is used for generating a silver-standard dependency treebank of 4500 Urdu tweets.
We present and evaluate the concept of FeelMusic and evaluate an implementation of it. It is an augmentation of music through the haptic translation of core musical elements. Music and touch are intrinsic modes of affective communication that are physically sensed. By projecting musical features such as rhythm and melody into the haptic domain, we can explore and enrich this embodied sensation; hence, we investigated audio-tactile mappings that successfully render emotive qualities. We began by investigating the affective qualities of vibrotactile stimuli through a psychophysical study with 20 participants using the circumplex model of affect. We found positive correlations between vibration frequency and arousal across participants, but correlations with valence were specific to the individual. We then developed novel FeelMusic mappings by translating key features of music samples and implementing them with “Pump-and-Vibe”, a wearable interface utilising fluidic actuation and vibration to generate dynamic haptic sensations. We conducted a preliminary investigation to evaluate the FeelMusic mappings by gathering 20 participants’ responses to the musical, tactile and combined stimuli, using valence ratings and descriptive words from Hevner’s adjective circle to measure affect. These mappings, and new tactile compositions, validated that FeelMusic interfaces have the potential to enrich musical experiences and be a means of affective communication in their own right. FeelMusic is a tangible realisation of the expression “feel the music”, enriching our musical experiences.
Inferring emotions from Head Movement (HM) and Eye Movement (EM) data in 360° Virtual Reality (VR) can enable a low-cost means of improving users’ Quality of Experience. Correlations have been shown between retrospective emotions and HM, as well as EM when tested with static 360° images. In this early work, we investigate the relationship between momentary emotion self-reports and HM/EM in HMD-based 360° VR video watching. We draw on HM/EM data from a controlled study (N=32) where participants watched eight 1-minute 360° emotion-inducing video clips, and annotated their valence and arousal levels continuously in real-time. We analyzed HM/EM features across fine-grained emotion labels from video segments with varying lengths (5-60s), and found significant correlations between HM rotation data, as well as some EM features, with valence and arousal ratings. We show that fine-grained emotion labels provide greater insight into how HM/EM relate to emotions during HMD-based 360° VR video watching.
Abstract The paper investigates formal language in persuasive discourse on the r /C hange M y V iew subreddit. We collected a corpus of 100 million messages, split into subcorpora based on the user-awarded marker delta, which rewards changing an original poster’s view. Assuming that formality/informality is potentially an important factor in the persuasiveness of a message, we examine the two subcorpora with respect to formality markers. The results indicate no systematic variation along the formality/informality continuum between persuasive and non-persuasive posts on r /C hange M y V iew. The posters use personal pronouns, suasive verbs, emphatics, imperatives, elaborate connectors and WH-questions with similar frequency, and express themselves using vocabulary and syntax of similar complexity. Moreover, keyword lists and n-gram rankings indicate no register difference. A qualitative analysis of concordance lines for persuade and change PRONOUN view paints a picture of a community that values factual, evidence-based discourse and openness to logical persuasion, with a linguistic norm of relatively formal, sophisticated register.
Recent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection. Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups. We harness this genre metadata as a weak supervision signal for targeted data selection in zeroshot dependency parsing. Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations. We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios. For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection. Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages.
Introduction<br><br> Penn Discourse Treebank Version 2.0 - German Translation was developed at the University of Potsdam's Applied Computational Linguistics group and consists of approximately one million tokens derived from Penn Discourse Treebank Version 2.0 (LDC2008T05). This data was translated into German and annotated for shallow discourse relations in the financial news domain.<br><br> The aim of the University of Pennsylvania's Penn Discourse Treebank (PDTB) project is to annotate the Wall Street Journal text in Treebank-2 with discourse relations. PDTB 2.0 contains 40,600 tokens of annotation relations. PDTB2-German is based on a subset of PDTB2.0 used in the 2016 CoNLL Shared Task on Multilingual Shallow Discourse Parsing. Data<br><br> Data is in CoNLL format. Text was automatically translated into German with deepL, and projections of the annotations using word alignments were produced with GIZA++. See the included documentation for more information on the relation annotations.<br><br> Source text and CoNLL format annotations are each presented in their own tab separated plain text file, encoded in UTF-8. Samples<br><br> Please view this source sample and annotation sample. Updates<br><br> None at this time. Copyright Portions © 1987-1989 Dow Jones & Company, Inc., © 2008, 2012, 2021 The Penn Discourse Treebank Group, © 2021 Manfred Stede, © 1993-1995, 2008, 2012, 2021 Trustees of the University of Pennsylvania
This article deals with a vision for the achievement of the project of a linguistic atlas of the dialects of Algeria supported by colored digital maps, showing the dialectical diversity of the selected region. And the circulation, the dialects used and current on the tongue of the inhabitants of Algeria We will focus in our research on the aspect of processing linguistic data in the manufacture of a digital linguistic atlas, in an attempt to invest computer data in describing local dialects and access to digital content, with the possibility of this content audio recordings of dialect variations, by providing our data bank through this linguistic research Also to pay tribute to the importance of the linguistic atlas in meeting the need of dialect workers for linguistic maps that identify the locations, nature and types of dialectal diversity in Algeria, The research also comes mainly to embody the digital principle and support the Arabic language by storing it digitally with the possibility of managing and printing linguistic maps according to the changes in them.
The Mongolian written language and the traditional Mongolian script were a “pre-modern” language and script that transcended ethnicity and dialect. The Mongolian script can be read in any dialect. However, modern languages demand pronunciation norms, making the Mongolian script unsuitable as the official script of a modern nation-state. The countries and regions that used the Mongolian script changed the script in the first half of the 20th century, except for the Mongolian ethnic areas of China which continued to use the original Mongolian script. This was possible as Mongolian was a minority language in China, and the scope of its use was limited. However, Inner Mongolia, too, faced the issue of the written language not conforming with the spoken language. Therefore, in the 1930s, Mongolian literary figures and others in Manchukuo attempted to unify the Mongolian written and spoken language based on the genbun itchi movement in Japan. Genbun itchi sought to unify the Japanese written and spoken language, and was translated literally as “üge üsüg-i nigen bolcaqui” in Mongolian. This effort was only partially successful in changing the Mongolian script to match the spoken pronunciation, and systematic genbun itchi could not be achieved. Since 1945, Inner Mongolia learned from the example of the Mongolian People’s Republic and transformed the style of Mongolian script to that of modern Mongolian while still using the Mongolian script. However, the discord between the Mongolian script and the spoken pronunciation remains unresolved, and to address this, changes are being made to the Mongolian script to this day as they have been for the last eight decades. It is important to know that the Mongolian script is not only the modern script used in Chinese territory, but also the traditional script of the written language common to areas using the Mongolian written language. For this reason, the Mongolian script must be passed on to future generations. A means for unifying the Mongolian language while using the Mongolian script is to follow the Hanyu Pinyin system, which has succeeded in standardizing pronunciation while using Chinese characters—namely, enhance the Inner Mongolian “standard-sounding” writing system so that it can also function at the sentence level. The orthography of the Cyrillic alphabet of Mongolia offers an objectively good example for the development of a standard-sounding writing system in Inner Mongolia. After all, the problem of the Mongolian script reform comes down to that of genbun itchi.
Para un lingüista o computólogo orientado al análisis de textos en español, analizar oraciones sintácticamente puede ser un esfuerzo de mucho tiempo cuando la cantidad de oraciones es elevada. Esta labor es más delicada si se toma en cuenta la variante del español que se emplee. Más aún, seleccionar cómo se etiquetan las palabras según la función puede complicar el proceso y el su impacto al compartir el conocimiento adquirido. Esta investigación propone realizar parte del esfuerzo en forma automática, utilizando reglas gramaticales, con el fin de analizar las oraciones sintácticamente y etiquetarlas con dependencias universales; un etiquetado estándar y nemónico, capaz de ser aplicado a diferentes idiomas.
Eesti veebipuudepanga tekstid (Muischnek et al., 2019), mis on annoteeritud käsitsi nii ortograafiliste kui süntaktiliste lausepiiridega, samuti on kontrollitud ja parandatud sõnestust. Lausete annoteerimisprotsessi kirjeldavad Sirts ja Peekman (2020), sõnestuse kontrolli kirjeldab Kairit Peekmani (2020) bakalaureusetöö. Andmete kasutamisel palume viidata Sirts ja Peekman (2020) artiklile. Muischnek, K., Müürisep, K., & Särg, D. D. (2019). CG Roots of UD Treebank of Estonian Web Language. In Proceedings of the NoDaLiDa 2019 Workshop on Constraint Grammar-Methods, Tools and Applications, 30 September 2019, Turku, Finland (No. 168, pp. 23-26). Linköping University Electronic Press. Peekman, K. (2020). Automaatse lausestamise ja sõnestamise hindamine uue meedia keele korpusel (bakalaureusetöö). Tartu Ülikool. Kättesaadav https://comserv.cs.ut.ee/ati_thesis/datasheet.php?id=69690&year=2020. Sirts, K., & Peekman, K. (2020). Evaluating Sentence Segmentation and Word Tokenization Systems on Estonian Web Texts. In Volume 328: Human Language Technologies – The Baltic Perspective, Frontiers in Artificial Intelligence and Applications, pages 174-181.
Abstract Despite the importance of mastering different types of formulaic sequences in a second language, little is known about the relative effect of different input modes on their acquisition. This study explores the learning of a particular type of formulaic language (binomials) in three input modes (reading-only, listening-only, and reading-while-listening) at different frequencies of exposure (2, 4, 5 and 6 occurrences). Arabic learners of English were presented with three stories, each in a different mode, that contained novel binomials (e.g., wires and pipes ) and existing binomials (e.g., brother and sister ). Two post-tests (multiple-choice and familiarity ratings) assessed learners’ knowledge of the binomials. Results showed that reading-only and reading-while-listening led to better performance on the tasks than listening-only. Frequency of exposure had an effect on the perceived familiarity of binomials.
While high performance have been obtained for high-resource languages, performance on low-resource languages lags behind. In this paper we focus on the parsing of the low-resource language Frisian. We use a sample of code-switched, spontaneously spoken data, which proves to be a challenging setup. We propose to train a parser specifically tailored towards the target domain, by selecting instances from multiple treebanks. Specifically, we use Latent Dirichlet Allocation (LDA), with word and character N-grams. We use a deep biaffine parser initialized with mBERT. The best single source treebank (nl_alpino) resulted in an LAS of 54.7 whereas our data selection outperformed the single best transfer treebank and led to 55.6 LAS on the test data. Additional experiments consisted of removing diacritics from our Frisian data, creating more similar training data by cropping sentences and running our best model using XLM-R. These experiments did not lead to a better performance.
Motivated by collective emotions theories that propose emotions shared between individuals predict group-level qualities, we hypothesized that co-experienced affect during interactions is associated with relationship quality, above and beyond the effects of individually experienced affect. Consistent with positivity resonance theory, we also hypothesized that co-experienced positive affect would have a stronger association with relationship quality than would co-experienced negative affect. We tested these hypotheses in 150 married couples across 3 conversational interactions: a conflict, a neutral topic, and a pleasant topic. Spouses continuously rated their individual affective experience during each conversation while watching video-recordings of their interactions. These individual affect ratings were used to determine, for positive and negative affect separately, the number of seconds of co-experienced affect and individually experienced affect during each conversation. In line with hypotheses, results from all 3 conversational topics suggest that more co-experienced positive affect is associated with greater marital quality, whereas more co-experienced negative affect is associated with worse marital quality. Individual level affect factors added little explanatory value beyond co-experienced affect. Comparing co-experienced positive affect and co-experienced negative affect, we found that co-experienced positive affect generally outperformed co-experienced negative affect, although co-experienced negative affect was especially diagnostic during the pleasant conversational topic. Findings suggest that co-experienced positive affect may be an integral component of high-quality relationships and highlight the power of co-experienced affect for individual perceptions of relationship quality. (PsycInfo Database Record (c) 2022 APA, all rights reserved).