Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Recent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection. Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups. We harness this genre metadata as a weak supervision signal for targeted data selection in zeroshot dependency parsing. Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations. We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios. For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection. Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages.
Abstract The paper investigates formal language in persuasive discourse on the r /C hange M y V iew subreddit. We collected a corpus of 100 million messages, split into subcorpora based on the user-awarded marker delta, which rewards changing an original poster’s view. Assuming that formality/informality is potentially an important factor in the persuasiveness of a message, we examine the two subcorpora with respect to formality markers. The results indicate no systematic variation along the formality/informality continuum between persuasive and non-persuasive posts on r /C hange M y V iew. The posters use personal pronouns, suasive verbs, emphatics, imperatives, elaborate connectors and WH-questions with similar frequency, and express themselves using vocabulary and syntax of similar complexity. Moreover, keyword lists and n-gram rankings indicate no register difference. A qualitative analysis of concordance lines for persuade and change PRONOUN view paints a picture of a community that values factual, evidence-based discourse and openness to logical persuasion, with a linguistic norm of relatively formal, sophisticated register.
Inferring emotions from Head Movement (HM) and Eye Movement (EM) data in 360° Virtual Reality (VR) can enable a low-cost means of improving users’ Quality of Experience. Correlations have been shown between retrospective emotions and HM, as well as EM when tested with static 360° images. In this early work, we investigate the relationship between momentary emotion self-reports and HM/EM in HMD-based 360° VR video watching. We draw on HM/EM data from a controlled study (N=32) where participants watched eight 1-minute 360° emotion-inducing video clips, and annotated their valence and arousal levels continuously in real-time. We analyzed HM/EM features across fine-grained emotion labels from video segments with varying lengths (5-60s), and found significant correlations between HM rotation data, as well as some EM features, with valence and arousal ratings. We show that fine-grained emotion labels provide greater insight into how HM/EM relate to emotions during HMD-based 360° VR video watching.
We present and evaluate the concept of FeelMusic and evaluate an implementation of it. It is an augmentation of music through the haptic translation of core musical elements. Music and touch are intrinsic modes of affective communication that are physically sensed. By projecting musical features such as rhythm and melody into the haptic domain, we can explore and enrich this embodied sensation; hence, we investigated audio-tactile mappings that successfully render emotive qualities. We began by investigating the affective qualities of vibrotactile stimuli through a psychophysical study with 20 participants using the circumplex model of affect. We found positive correlations between vibration frequency and arousal across participants, but correlations with valence were specific to the individual. We then developed novel FeelMusic mappings by translating key features of music samples and implementing them with “Pump-and-Vibe”, a wearable interface utilising fluidic actuation and vibration to generate dynamic haptic sensations. We conducted a preliminary investigation to evaluate the FeelMusic mappings by gathering 20 participants’ responses to the musical, tactile and combined stimuli, using valence ratings and descriptive words from Hevner’s adjective circle to measure affect. These mappings, and new tactile compositions, validated that FeelMusic interfaces have the potential to enrich musical experiences and be a means of affective communication in their own right. FeelMusic is a tangible realisation of the expression “feel the music”, enriching our musical experiences.
AbstractData were checked for univariate outliers using standardised scores (<i>z</i> > ± 3.29) and for multivariate outliers using the Mahalanobis distance test (<i>p</i> <.001; Tabachnick & Fidell, 2019). Data were also examined for the parametric assumptions that underlie within-subjects ANOVA (Tabachnick & Fidell, 2019), other than the behavioural data from video analysis, which were derived from frequency counts. Where the assumption of sphericity was violated, Greenhouse–Geisser-adjusted <i>F</i> tests were used. Initial analyses employed repeated-measures (RM) 2 (Load) × 3 (Tempo) (M)ANOVAs for the three psychological measures (i.e., RSME, NASA-TLX, and Affect Grid) and cardiac measures (HR and HRV indices). Additionally, exploratory analyses were conducted using a mixed-model approach, adopting the between-subject factors of personality (introvert vs. extrovert), sex (women vs. men), and age group (young adults vs. middle-aged adults). Significant <i>F</i> tests were followed up with pairwise/multiple comparisons, or in the case of interaction effects, examination of 95% confidence intervals (95% CIs) to identify where differences lay. Behavioural data were collated for the urban environment (high load) simulation under the following categories: (a) video data pertaining to four triggers (pedestrian, garbage truck, traffic lights, and vehicle cutting); (b) and simulator-derived data from the accelerator and brake pedal positions (i.e., 0 = no pressure applied, 1 = maximum braking); (c) mean speed (mph), and (d) course completion time (min). For the highway environment (low load), simulator data were collated for: (a) accelerator and brake pedal positions; (b) mean speed, and (c) completion time. From among these data, where parametric assumptions were not met and transformations would not serve to normalise the distribution, nonparametric analyses was adopted using rank-based, nonparametric tests. Specifically, the Wald-type statistic (WTS) and the ANOVA-type statistic (ATS) was computed within the nparLD package (Noguchi et al., 2012) of data analysis software R. In the absence of the load factor for the trigger and pedal data, a within-subjects, one-way ANOVA for the effect of music tempo was computed. In the exploratory analyses, a factorial approach was used, with a series of mixed-model ANOVAs 3 ([Tempo] × 2 [Personality], 3 [Tempo] × 2 [Sex], and 3 [Tempo] × 2 [Age Group]). Note that the main effect of tempo for trigger and pedal data is relevant to the main analysis but in the interest of parsimony is incorporated within exploratory factorial analyses.<br>Detailed Description of Data FileThis SPSS data file contains the demographic data (i.e. sex, age, age group [1 = young adult, 2 = middle-aged adult], personality [1 = introvert, 2 = extrovert]) for each of the 46 participants (presented with one participant per row). Behavioural measures relating to the driving simulation are included. These include the elapsed time (mins) for each trial. Also, mean speed (mph), brake pedal use (i.e. 0 = no pressure applied, 1 = maximal braking), accelerator pedal use (i.e., 0 = no pressure applied, 1 = maximal acceleration) and risk ratings (on a scale from 1 [<i>safe driving</i>] to 4 [<i>reckless driving</i>]). Note that these performance-related measures appear 12 times in total; that is for each simulator trigger (i.e. a pedestrian who walked at 5 km/h across a zebra crossing, a garbage truck that moved slowly in the left-hand lane and prompted an overtaking manoeuvre, traffic lights that changed to red, a slow vehicle on a stretch of road on which overtaking was prohibited and a vehicle that cut across unexpectedly at a four-way intersection) across all three high-load (urban) conditions. Additionally, the scores across all conditions for the measures of the NASA Task Load Index (NASA-TLX), Affect Grid (affective valence and affective arousal), Rating Scale Mental Effort (RSME) and wordsearch task are included. The psychophysiological measures of heart rate variability (HRV) and mean heart rate (HR) are also included. For HRV and HR, specifically, we present mean HR, minimum HR, maximum HR, standard deviation of normal RR intervals (SDNN), HR standard deviation and root mean square of successive differences (RMSSD). <i>z</i>-scores (i.e. standardised scores) for each variable are also included. Note that each participant was exposed to six experimental conditions (high load/fast tempo, high load/slow tempo, high load/no music, low load/fast music, low load/slow music and low load/no music). Accordingly, the measures that pertain to each trial (i.e. NASA-TLX, RSME, Affect Grid, wordsearch task, HRV indices, risk ratings, mean speed, brake pedal use and accelerator pedal use) appear six times in the data file.
Manually annotated corpus is a perquisite for several natural language processing applications including parsing. Nevertheless, annotated corpus is not always available for resource-poor languages, especially when domain under consideration is noisy user-generated data found on social media platforms such as Twitter. To overcome this deficiency of hand-annotated corpus, researchers have focused their attention on semi-automatic corpus annotation methods. This paper describes the experiments carried out using semi-automatic methods like self-training and co-training in an attempt for creating silver-standard dependency treebank of Urdu tweets. Six iterations of each approach were performed using same experimental conditions using MaltParser and Parsito parser, both statistical data driven parsers. For self-training experiments, the best performing MaltParser model was trained on 1250 Urdu tweets, with an accuracy of 70.2% LA, 74.4% UAS, 63% LAS. Whereas the best performing Parsito model was also trained on 1250 Urdu tweets with an accuracy of 70.8% LA, 74.8% UAS, 63.4% LAS. For co-training experiments, best performing MaltParser model was trained on 1500 Urdu tweets, with an accuracy of 70.5% LA, 74.4% UAS, 63.2% LAS. The best performing Parsito model was also trained on 1500 Urdu tweets with an accuracy of 70.5% LA, 74.3% UAS, 63% LAS. Although, there was not much difference between the results of both approaches, co-training results were slightly better for both parsers and is used for generating a silver-standard dependency treebank of 4500 Urdu tweets.
The article explores the use of contextual slang as linguistic performance by three all-female friendship groups in Calabar metropolis, Cross River State, south-eastern Nigeria. I argue that slang constitutes critical components of the discursive practices of young urban Nigerian women in maintaining friendship and deviating from stereotyped cultural and linguistic norms. Drawing insights from the analytical tools of African feminism and linguistic ideology, the article discusses recurrent themes in young women’s contextual slanguage and the motivations for the use of these creative linguistic and cultural resources in defining participants’ authentic social selves and in enacting their different modes of belonging. Qualitative ethnographic data for the study were sourced from participant observations, semi-structured interviews and informal conversations with 30 participants. The study concludes that young urban women utilise contextual slang as indexical tools in their everyday narratives to negotiate meaning in relation to the experience of their social lives, to acculturate to male linguistic norms and to affiliate with ideologies that represent gendered identity.
This paper develops the concept of word order universals based on a data analysis of the Universal Dependencies project, which proposes treebanks of more than 90 languages encoded with the same annotation scheme. The nature of the data we work on allows us to extract rich details for testing well-known typological implicational universals and, further, explore new kinds of universals that we call quantitative universals. We show how such quantitative universals are in essence different from implicational universals, including statistical universals, by the fact that they no longer lay down any claims on categorical statements, but rather on continuous parameters, opening a new field of research we propose to call typometrics.
Using a very large lexical database and generalized additive modeling, this article reveals that labial-velar (LV) stops are marginal phonemes in many of the languages of Northern Sub-Saharan Africa that have them, and that the languages in which they are not marginal are grouped into three compact zones of high lexical LV frequency. The resulting picture allows us to formulate precise hypotheses about the spread of the Niger-Congo and Central Sudanic languages and about the origins of the linguistic area known as the Sudanic zone or Macro-Sudan belt. It shows that LV stops are a substrate feature that should not be reconstructed into the early stages of the languages that currently have them. We illustrate the implications of our findings for linguistic prehistory with a short discussion of the Bantu expansion. Our data also indirectly confirm the hypothesis that LV stops are more recurrent in expressive parts of the vocabulary, and we argue that this has a common explanation with the well-known fact that they tend to be restricted to stem-initial position in what we call C-emphasis prosody.
The prescriptive approach has been prevalent in discussions about the linguistic norm for many decades. Many linguists question the primacy of social custom and make many arbitrary changes to establish the subjective form of the norm. In connection with the planned The Dictionary of Proper Uses of Languagethe author of the article presents the best structuralist traditions and calls for research on the linguistic norm which is based on descriptive methods. It is necessary to completely break away from all manifestations of arbitrariness and subjectivity in contemporary prescriptive linguistics. The fundamental premise that the linguistic norm is a fact based on usus must be reflected in relevant procedures aimed at analyzing corpora consisting of millions of words. Such an approach will make it possible to establish a model that comprises more than just individual language uses. As far as dictionary definitions are concerned, the most frequent, widespread and thus typical linguistic units should be primarily considered to be normative. Typicality, determined by frequency, as well as textual, social and territorial conditions, is the most important category.
Abstract The annotation scheme of dependency treebanks might have an impact on the results of linguistic analysis, thus leading to different interpretations of linguistic phenomena. This study compares the results of two widely used dependency measures, i.e., dependency direction and dependency distance, based on 18 parallel Universal Dependencies (UD) annotated treebanks and 18 corresponding Surface‐Syntactic Universal Dependencies (SUD) annotated treebanks. The results show that (1) Based on the semantic UD and syntactic SUD, dependency relations between function words and content words share the opposite dependency directions but similar dependency distances; (2) Annotation scheme has a significant impact on dependency direction, though the effect size is small. We find that the proportions of head‐final dependencies based on the syntactic SUD can better group language families than those based on semantic UD; (3) Annotation scheme also affects dependency distance significantly, though its effect size is small. Mean dependency distances (MDDs) based on UD are always higher than those based on SUD. However, the MDDs based on both annotation schemes are within a certain threshold, which shows that the linguistic universal of dependency distance minimization is independent of annotation schemes.
Semantics is a research field that has gained an extensive interest recently. This survey describes recent works in the field of semantics, a part of the broader area of computational linguistics. One of the important aspects of computational linguistics is using proper methods to distribute semantics for obtaining representations of the meaning of words. This survey summarizes the latest state of the art approaches in semantics that use deep learning methods, datasets, and lexical databases, specifying semantics under two categories such as semantic similarity and sentence modeling.
The therapeutic effect of antidepressants has been demonstrated for anhedonia in patients with depression. However, antidepressants may cause side-effects, such as cardiovascular dysfunction. Although physical activity has minor side-effects, it may serve as an alternative for improving anhedonia and depression. We sought to investigate whether physical activity reduces the level of anhedonia in individuals with depression. Fifty-six university students with moderate depressive symptoms (Beck Depression Inventory total score > 16) were divided into three training groups: the Running Group (RG, n = 19), the Stretching Group (SG, n = 19), and the Control Group (n = 18). We employed the Monetary Incentive Delay (MID) task and the Temporal Experience of Pleasure Scale (TEPS) to evaluate hedonic capacity. All participants in the RG and SG received 8 weeks of jogging and stretching training, respectively. The RG experienced an increase in the level of arousal during anticipation of a future reward and recalled less negativity towards the loss condition. The SG exhibited enhanced scores on the Anticipatory and Consummatory Pleasure subscales of the TEPS after training. Moreover, in the RG, greater improvements in anticipatory arousal ratings for pleasure and remembered valence ratings for negative affect were associated with longer training duration, lower maximum heart rate, and higher consumed calories during training. To conclude, physical activity is effective in improving anticipatory anhedonia in individuals with depressive symptoms.
Abstract This article contributes to a dialogue between childhood studies and the sociolinguistic subfield ‘Family Language Policy’ (‘FLP’). The article argues that the two fields provide complementary vantage points for exploring child agency. It explains a revised version of a model I developed to conceptualise child agency in FLP, consisting of four intersecting dimensions: compliance regimes; linguistic norms; linguistic competence and generational positioning (Smith‐Christmas, Handbook of home language maintenance and development. De Gruyter Mouton, pp. 218–235, 2020a). The article examines two conversational excerpts as a means to illustrating the dynamic and relational nature of child agency and how it is both shaped by as well as shapes interactional practices over time and space.
OBJECTIVE: The study sought to develop and evaluate neural natural language processing (NLP) packages for the syntactic analysis and named entity recognition of biomedical and clinical English text. MATERIALS AND METHODS: We implement and train biomedical and clinical English NLP pipelines by extending the widely used Stanza library originally designed for general NLP tasks. Our models are trained with a mix of public datasets such as the CRAFT treebank as well as with a private corpus of radiology reports annotated with 5 radiology-domain entities. The resulting pipelines are fully based on neural networks, and are able to perform tokenization, part-of-speech tagging, lemmatization, dependency parsing, and named entity recognition for both biomedical and clinical text. We compare our systems against popular open-source NLP libraries such as CoreNLP and scispaCy, state-of-the-art models such as the BioBERT models, and winning systems from the BioNLP CRAFT shared task. RESULTS: For syntactic analysis, our systems achieve much better performance compared with the released scispaCy models and CoreNLP models retrained on the same treebanks, and are on par with the winning system from the CRAFT shared task. For NER, our systems substantially outperform scispaCy, and are better or on par with the state-of-the-art performance from BioBERT, while being much more computationally efficient. CONCLUSIONS: We introduce biomedical and clinical NLP packages built for the Stanza library. These packages offer performance that is similar to the state of the art, and are also optimized for ease of use. To facilitate research, we make all our models publicly available. We also provide an online demonstration (http://stanza.run/bio).
The COVID-19 pandemic has dramatically changed the nature of our social interactions. In order to understand how protective equipment and distancing measures influence the ability to comprehend others' emotions and, thus, to effectively interact with others, we carried out an online study across the Italian population during the first pandemic peak. Participants were shown static facial expressions (Angry, Happy and Neutral) covered by a sanitary mask or by a scarf. They were asked to evaluate the expressed emotions as well as to assess the degree to which one would adopt physical and social distancing measures for each stimulus. Results demonstrate that, despite the covering of the lower-face, participants correctly recognized the facial expressions of emotions with a polarizing effect on emotional valence ratings found in females. Noticeably, while females' ratings for physical and social distancing were driven by the emotional content of the stimuli, males were influenced by the "covered" condition. The results also show the impact of the pandemic on anxiety and fear experienced by participants. Taken together, our results offer novel insights on the impact of the COVID-19 pandemic on social interactions, providing a deeper understanding of the way people react to different kinds of protective face covering.
Recent advances in deep learning techniques have enabled machines to generate cohesive open-ended text when prompted with a sequence of words as context. While these models now empower many downstream applications from conversation bots to automatic storytelling, they have been shown to generate texts that exhibit social biases. To systematically study and benchmark social biases in open-ended language generation, we introduce the Bias in Open-Ended Language Generation Dataset (BOLD), a large-scale dataset that consists of 23,679 English text generation prompts for bias benchmarking across five domains: profession, gender, race, religion, and political ideology. We also propose new automated metrics for toxicity, psycholinguistic norms, and text gender polarity to measure social biases in open-ended text generation from multiple angles. An examination of text generated from three popular language models reveals that the majority of these models exhibit a larger social bias than human-written Wikipedia text across all domains. With these results we highlight the need to benchmark biases in open-ended language generation and caution users of language generation models on downstream tasks to be cognizant of these embedded prejudices.
International audience
The ventromedial and dorsolateral prefrontal cortex are two major prefrontal regions that usually interact in serving different cognitive functions. On the other hand, these regions are also involved in cognitive processing of emotions but their contribution to emotional processing is not well-studied. In the present study, we investigated the role of these regions in three dimensions (valence, arousal and dominance) of emotional processing of stimuli via ratings of visual stimuli performed by the study participants on these dimensions. Twenty- two healthy adult participants (mean age 25.21 ± 3.84 years) were recruited and received anodal and sham transcranial direct current stimulation (tDCS) (1.5 mA, 15 min) over the dorsolateral prefrontal cortex (dlPFC) and and ventromedial prefrontal cortex (vmPFC) in three separate sessions with an at least 72-h interval. During stimulation, participants underwent an emotional task in each stimulation condition. The task included 100 visual stimuli and participants were asked to rate them with respect to valence, arousal, and dominance. Results show a significant effect of stimulation condition on different aspects of emotional processing. Specifically, anodal tDCS over the dlPFC significantly reduced valence attribution for positive pictures. In contrast, anodal tDCS over the vmPFC significantly reduced arousal ratings. Dominance ratings were not affected by the intervention. Our results suggest that the dlPFC is involved in control and regulation of valence of emotional experiences, while the vmPFC might be involved in the extinction of arousal caused by emotional stimuli. Our findings implicate dimension-specific processing of emotions by different prefrontal areas which has implications for disorders characterized by emotional disturbances such as anxiety or mood disorders.
Recursive Deep Models have been used as powerful models to learn \ncompositional representations of text for many natural language processing tasks. \nHowever, they require structured input (i.e. sentiment treebank) to encode sentences \nbased on their tree-based structure to enable them to learn latent semantics \nof words using recursive composition functions. In this paper, we present our \ncontributions and efforts for the Turkish Sentiment Treebank construction. We \nintroduce MS-TR, a Morphologically Enriched Sentiment Treebank, which was \nimplemented for training Recursive Deep Models to address compositional sentiment \nanalysis for Turkish, which is one of the well-known Morphologically Rich \nLanguage (MRL). We propose a semi-supervised automatic annotation, as a distantsupervision \napproach, using morphological features of words to infer the polarity of \nthe inner nodes of MS-TR as positive and negative. The proposed annotation model \nhas four different annotation levels: morph-level, stem-level, token-level, and \nreview-level. Each annotation level’s contribution was tested using three different \ndomain datasets, including product reviews, movie reviews, and the Turkish Natural \nCorpus essays. Comparative results were obtained with the Recursive Neural Tensor Networks (RNTN) model which is operated over MS-TR, and conventional machine learning methods. Experiments proved that RNTN outperformed the baseline methods and achieved much better accuracy results compared to the baseline methods, which cannot accurately capture the aggregated sentiment information.
Towards explainable affective computing (XAC), researchers have invested considerable effort into post hoc approaches and reverse engineering to seek explanations for deep learning models. However, alternative, intrinsic approaches that aim to build inherently interpretable models by restricting their complexity are yet to be widely explored. In this study, we integrate an explanatory polytomous item response model that provides a well-established psychological interpretation for ordinal scales with deep neural networks to realize high prediction performance and good result interpretability. We conducted an experiment on a growing task (i.e., predicting the idiosyncratic perception of emotional faces of an individual); as expected theoretically, the topmost parameters of our model demonstrated strong correlations with those of the corresponding ordinal item response model: r = 0.928 to 1.00. Our proposed intrinsic approach can used as a complementary framework for post-hoc methods in XAC to coach and support human social interactions.
Arabic dependency parsers have a poor performance compared to parsers of other languages. Recently the impact of annotation at lexical level of dependency treebank on the overall performance of the dependency parses has been extensively investigated. This paper focuses on the impact of coarse-grained and fine-grained dependency relations on the performance of Arabic dependency parsers. Moreover, this paper introduces the annotation rules for I3rab dependency treebank. Experimentally, the obtained results showed that having an appropriate set of dependency relations improves the performance of an Arabic dependency parser up to 27.55%.
Treebanks are valuable linguistic resources that include the syntactic structure of a language sentence in addition to part-of-speech tags and morphological features. They are mainly utilized in modeling statistical parsers. Although the statistical natural language parser has recently become more accurate for languages such as English, those for the Arabic language still have low accuracy. The purpose of this article is to construct a new Arabic dependency treebank based on the traditional Arabic grammatical theory and the characteristics of the Arabic language, to investigate their effects on the accuracy of statistical parsers. The proposed Arabic dependency treebank, called I3rab, contrasts with existing Arabic dependency treebanks in two main concepts. The first concept is the approach of determining the main word of the sentence, and the second concept is the representation of the joined and covert pronouns. To evaluate I3rab, we compared its performance against a subset of Prague Arabic Dependency Treebank that shares a comparable level of details. The conducted experiments show that the percentage improvement reached up to 10.24% in UAS and 18.42% in LAS.
Recent work on multilingual dependency parsing focused on developing highly multilingual parsers that can be applied to a wide range of low-resource languages. In this work, we substantially outperform such "one model to rule them all" approach with a heuristic selection of languages and treebanks on which to train the parser for a specific target language. Our approach, dubbed TOWER, first hierarchically clusters all Universal Dependencies languages based on their mutual syntactic similarity computed from human-coded URIEL vectors. For each low-resource target language, we then climb this language hierarchy starting from the leaf node of that language and heuristically choose the hierarchy level at which to collect training treebanks. This treebank selection heuristic is based on: (i) the aggregate size of all treebanks subsumed by the hierarchy level and (ii) the similarity of the languages in the training sample with the target language. For languages without development treebanks, we additionally use (ii) for model selection (i.e., early stopping) in order to prevent overfitting to development treebanks of closest languages. Our TOWER approach shows substantial gains for low-resource languages over two state-ofthe-art multilingual parsers, with more than 20 LAS point gains for some of those languages. Parsing models and code available at: https: //github.com/codogogo/towerparse.
This paper describes the grammatical patterning of two parts of speech – nouns and adjectives – included in the corpus-driven “Lexical Database of Lithuanian” as a foreign language. The lexical database is a lexicographic application of the Lithuanian Pedagogic Corpus (approx. 620.000 tokens) which was used to develop headword lists and to collect word usage information in the form of corpus patterns. In this project, we adopted a partially automated inductive procedure of Corpus Pattern Analysis for 207 verbs, 386 nouns, 87 adjectives, and 41 adverbs. The detected corpus patterns reflect different meanings of the headword. Each pattern presents information on grammatical, semantic, and lexical levels. Manually selected examples illustrate all pattern components. In this paper, 673 patterns with nouns and 99 patterns with adjectives will be analysed discussing their syntactic behaviour in detail and providing some comments on lexis-grammar interface. The majority of patterns with nouns and adjectives are minimal patterns which include only the closest syntactical partners. This result is influenced by different procedures used to describe patterns with nouns, adjectives, and adverbs and patterns with verbs. Due to rich grammatical information, there are several similar patterns with one main (usually the most frequent) type and its variants. Pattern variants show that the grammatical characteristics of a specific word usage are rather individual.
This article introduces the working methods of the Parsed Historical Corpus of the Welsh Language (PARSHCWL). The corpus is designed to provide researchers with a tool for automatic exhaustive extraction of instances of grammatical structures from Middle and Modern Welsh texts in a way comparable to similar tools that already exist for various European languages. The major features of the corpus are outlined, along with the overall architecture of the workflow needed for a team of researchers to produce it. In this paper, the two first stages of the process, namely pre-processing of texts and automated part-of-speech (POS) tagging are discussed in some detail, focusing in particular on major issues involved in defining word boundaries and in defining a robust and useful tagset.
This paper describes the construction and annotation of the Late Latin Charter Treebank, a set of three dependency treebanks (llct1, llct2 and llct3) which together contain 1,261 Early Medieval Latin documentary texts (i.e., original charters) written in Italy between ad 714 and 1000 (about 594,000 tokens). The paper focusses on matters which a linguistically or philologically inclined user of llct needs to know: the criteria on which the charters were selected, the special characteristics of the annotation types utilised, and the geographical and chronological distribution of the data. In addition to normal queries on forms, lemmas, morphology and syntax, complex philological research settings are enabled by the textual annotation layer of llct, which indicates abbreviated and damaged words, as well as the formulaic and non-formulaic passages of each charter.
The PapyGreek Treebanks dataset contains documentary texts written in Postclassical Greek (ca. 300 BCE–700 CE), morphosyntactically annotated according to Dependency Grammar. The source of the texts is the Duke Databank of Documentary Papyri (DDbDP), which preserves the modern editorial treatment of the documents in TEI Epidoc XML encoding. Aiming to expose linguistic variation in the DDbDP, we have annotated two versions of a selection of documents: the plain transcription and an editorially corrected version. The dataset also comprises metadata about the documents’ dating and provenance, text type, and the persons involved. Furthermore, it facilitates linguistic research on these texts.
Chapter 2 provides a sociohistorical analysis of the evolution of Yungueño Spanish, Chota Valley Spanish and Chincha Spanish. The chapter illustrates general aspects of the African Diaspora to the Americas and its specific linguistic consequences in Yungas (Bolivia), Chota Valley (Ecuador) and Chincha (Peru). Given the historical evidence available for these Afro-Hispanic Languages of the Americas, I propose that these contact varieties developed in isolated rural villages, not subject to the social pressures imposed by education, standardization and the linguistic norm. In such a context, advanced SLA processes could be nativized and conventionalized at the local level, thus crystallizing in the L1 varieties spoken by subsequent generations of these Afro-Andean communities.
We present an approach for automatic punctuation restoration with BERT models for English and Hungarian. For English, we conduct our experiments on Ted Talks, a commonly used benchmark for punctuation restoration, while for Hungarian we evaluate our models on the Szeged Treebank dataset. Our best models achieve a macro-averaged $F_1$-score of 79.8 in English and 82.2 in Hungarian. Our code is publicly available.
The purpose of this study is to examine the orthographic and phonological characteristics of the Yeongsan Sillok(the biography of Yeongsan), published in Jeollabuk-do in the early 20th century. The author of this book is considered to be Jang Bong-seon, an educator from Jeongeup city in Jeollabuk-do. Accordingly, it is expected that this book contains the orthographic characteristics and attitudes toward the language of young intellectuals in Jeollabuk-do in the early 20th century. In Chapter 3, we looked at the orthographic characteristics of this book. The writing characteristics of this book largely follow the characteristics of the 19th century Jeollabuk-do dialect based on the tradition of modern Korean. However, a transitional characteristic of the language transforming into present-day Korean was also present. Although only a few examples have been confirmed, the writing of double consonant letters for tense consonant are gradually similar to the notation method of modern Korean. This can be understood as a dissolution process. At the same time, with the exception of some circumstances of verbs, the tendency to split consonants is widely confirmed, and the modern Korean notation for the /ㄹㄹ/ chain (ㄹㄴ, ​​ㄹㅇ) is gradually changing to ㄹㄹ. Above all, the fact that the notation of ․ or diphthong after sibilants no longer appears in this book is a characteristic feature that differs from data from the Jeollabuk-do region of the same period. This writing trend seems to be related to a set of linguistic norms compiled in the first half of the 20th century. Recalling that the author of this book established a private school in the 1920s and 1930s and devoted himself to educational activities, this assumption is somewhat probable. In Chapter 4, we looked at the phonological characteristics of the Yeongsan Sillok(the biography of Yeongsan). Front-vowelization was very active inside the morpheme, but at the morpheme boundary, it appeared only in the environment behind c. The simple vowelization of jə>e is confirmed throughout the interior and boundary of the morpheme, and it must have been a productive phonological phenomenon in the Jeollabuk-do dialect in the early 20th century, as hypercorrection types also appeared. Regarding the alternation of the ending ‘-a/ə’, when the stem vowel is ‘ø’, there is a high tendency to combine these to ‘-ə’. This is different from the 19th century and modern Jeollabuk-do dialects. In the case of umlauts, only very limited examples were shown. And although t-palatalization is quite actively realized, only a few examples of k-palatalization were shown. Through this realization of phonological phenomena, we were able to confirm whether the young intellectuals in the Jeollabuk-do region in the early 20th century had linguistic attitudes toward the Jeollabuk-do dialect. In this book, the typical phonological phenomenon of the Jeollabuk-do dialect was confirmed only to a very limited extent due to its negative evaluation by the author.
Abstract This paper proposes to study the contrastive syntax of French and Chinese through the lens of syntactic mismatches, and by making use of parallel treebanks. A syntactic mismatch is the non-similarity between the syntactic structures of one linguistic unit and its translation. Syntactic mismatches are formalized using the notion of paraphrase from the Meaning-Text Theory, which allows for capturing mismatches at different levels of the linguistic description (e.g. Semantic, Deep-Syntactic, and Surface-Syntactic). In this paper, we report in details on the types of paraphrases found in the seed corpus used, demonstrating that the Deep-Syntactic paraphrases constitute the best starting point for our study. Then, we show how, starting from the seed corpus, we semi-automatically constructed a multi-layer parallel treebank with the alignment and annotation of paraphrases.
The main motivation behind this exam document is to look at the extent to which EWOM among customers can affect the brand image and the intent of buying the consumer in the clothing industry. A key condition display process is linked to the E-WOM impacts survey on brand image and buyer's purchase target. The exploration program was tested using an example of 385 respondents who included information within online purchasing groups and examined buyers of Pakistan's textile industry at the time of the investigation. The document recalls the methodologies to help a brand profitably through client-based social networking on the web, as well as typical suggestions for delegated websites and dialogues to enhance this note on a major path with people in their online dating. This explorative document extends the winning image rating to another set, in particular e-WOM. This document provides profitable knowledge on e-WOM estimation, brand image and purchasing expectations of the purchaser in the clothing industry and provides a facility for future search for tagging items.
In this paper, we propose a method for learning representations in the space of Gaussian-like distribution defined on a novel geometrical space called Kinematic space. The utility of non-Euclidean geometry for deep representation learning has recently been in vogue, specifically models of hyperbolic geometry such as Poincaré and Lorentz models have proven useful for learning hierarchical representations. Going beyond manifolds with constant curvature, albeit has better representation capacity might lead to unhanding of computationally tractable tools like Riemannian optimization methods. Here, we explore a pseudo-Riemannian auxiliary Lorentzian space called Kinematic space and provide a principled approach for constructing a Gaussian-like distribution, which is compatible with gradient-based learning methods, to formulate a probabilistic word embedding framework. Contrary to, mapping lexically distributed representations to a single point vector in Euclidean space, we advocate for mapping entities to density-based representations, as it provides explicit control over the uncertainty in representations. We test our framework by embedding WordNet-Noun hierarchy, a large lexical database, our experiments report strong consistent improvements in Mean Rank and Mean Average Precision (MAP) values compared to probabilistic word embedding frameworks defined on Euclidean and hyperbolic spaces. We show an average improvement of 72.68% in MAP and 82.60% in Rank compared to the hyperbolic version. Our work serves as evidence for the utility of novel geometrical spaces for learning hierarchical representations.
Abstract While databases of taboo language word norms exist, none focus specifically on slurs as a category of taboo language. Furthermore, no existing databases include measures of linguistic reclamation, a phenomenon which may specifically affect the processing of slurs. I produced a database in which 155 native or near-native speakers of British English rated 41 LGBTQ+ slurs for a number of word properties and measures of linguistic reclamation. I then ran correlation and demographic group comparison analyses on the resulting database. I found a clear correlation pattern between properties and reclamation behaviours. I also found that there were age-related differences in age of acquisition and familiarity ratings; that gender identity and sexual identity differences were affected by being the target of slurs; and that sexual identity particularly affected differences in reclamation ratings.
This paper compares two influential theories of processing difficulty: Gibson (2000)'s Dependency Locality Theory (DLT) and Hale (2001)'s Surprisal Theory. While prior work has aimed to compare DLT and Surprisal Theory (see I compare estimated surprisal values from two models, an RNN and a Transformer neural network, as well as DLT integration cost from a hand-parsed treebank, to reading times from the Dundee Corpus. The results for integration cost corroborate those of Ultimately, I conclude that a broad-coverage model must integrate both theories in order to most accurately predict processing difficulty.
Cloud-based enterprise search services (e.g., AWS Kendra) have been\nentrancing big data owners by offering convenient and real-time search\nsolutions to them. However, the problem is that individuals and organizations\npossessing confidential big data are hesitant to embrace such services due to\nvalid data privacy concerns. In addition, to offer an intelligent search, these\nservices access the user search history that further jeopardizes his/her\nprivacy. To overcome the privacy problem, the main idea of this research is to\nseparate the intelligence aspect of the search from its pattern matching\naspect. According to this idea, the search intelligence is provided by an\non-premises edge tier and the shared cloud tier only serves as an exhaustive\npattern matching search utility. We propose Smartness At Edge (SAED mechanism\nthat offers intelligence in the form of semantic and personalized search at the\nedge tier while maintaining privacy of the search on the cloud tier. At the\nedge tier, SAED uses a knowledge-based lexical database to expand the query and\ncover its semantics. SAED personalizes the search via an RNN model that can\nlearn the user interest. A word embedding model is used to retrieve documents\nbased on their semantic relevance to the search query. SAED is generic and can\nbe plugged into existing enterprise search systems and enable them to offer\nintelligent and privacy-preserving search without enforcing any change on them.\nEvaluation results on two enterprise search systems under real settings and\nverified by human users demonstrate that SAED can improve the relevancy of the\nretrieved results by on average 24% for plain-text and 75% for encrypted\ngeneric datasets.\n
Literary Works byAkaki Tsereteli are considered as versatile and diverse. In his works he touches upon almost everything by his poetry, prose, journalism or public work. It is obvious that he established "a type of versatile writer who is equally engaged in prose, poetry, journalism, dramaturgy, translations, children's literature and fables”. He was an extremely optimistic person who deeply believed in the future. The following words from one of his works seem amazingly and expressive: “Even if you kill a swallow, Spring will definitely come”. Connection between the old and the new forms, that is clearly shown within this emotionally colored expression, has become the goal of the research. We tried to find an answer to the question- what is the role of using old Georgian forms in Akaki's work?! Given paper analyses the samples such as: 1. Using proper name by its stem form in nominative case; 2. Ending words by - მან [-man] in the ergative form; 3. Full stems of demonstrative pronouns - ‘ამ’ [am], ‘ეგ’ [eg] (=this, that); 4. Using postposition – ‘ზე’ [ze] (=on), along with the forms - ზედ [-zed] and -ზედა [-zeda] (=on, over); 5. instrumental case forms formed by a suffix - ით [-it] (=with) (without postpositions); 6. Postposition and full agreement of attribute and antecedent 7. Characteristics of using inflection as a reflection of Old Georgian (გწყალობდესთ [gtskalobdet]...; გამოვჰკითხავ [gamovhkitkhav]...; ჰსვამ [hvsvam]...; ჰნიშნავს [hnishnavs]...; წარმოსთქვა [tsarmostkva]...; გასტეხე [gastekhe]...); 8. Using conjunction - ვით [vit] (=as/like) for comparison and so on. If we ask questions concerning the function of old Georgian forms in Akaki Tsereteli’s works, it becomes clear that they can be used for: 1. rhythm, emotiveness and expressiveness; 2. Preserving traditional forms, to maintain the connection between old and new Georgian. It should be mentioned that similar forms are equally reflected in Akaki’s prose and poetry which further reinforces the idea in favor of showing the connection between the old and the new and the desire to maintain this connection and always remember where we come from and who we are....This fact does not completely contradict the idea that Akaki is a representative of the generation that courageously rejected the old linguistic norms and contributed to the democratization (rapprochement process with the spoken language) of the literary language.
This paper presents and discusses the first Universal Dependencies treebank\nfor the Apurin\\~a language. The treebank contains 76 fully annotated sentences,\napplies 14 parts-of-speech, as well as seven augmented or new features - some\nof which are unique to Apurin\\~a. The construction of the treebank has also\nserved as an opportunity to develop finite-state description of the language\nand facilitate the transfer of open-source infrastructure possibilities to an\nendangered language of the Amazon. The source materials used in the initial\ntreebank represent fieldwork practices where not all tokens of all sentences\nare equally annotated. For this reason, establishing regular annotation\npractices for the entire Apurin\\~a treebank is an ongoing project.\n
PURPOSE: Medical education has been transformed during the COVID-19 pandemic, creating challenges regarding adequate training in ultrasound (US). Due to the discontinuation of traditional classroom teaching, the need to expand digital learning opportunities is undeniable. The aim of our study is to develop a tele-guided US course for undergraduate medical students and test the feasibility and efficacy of this digital US teaching method. MATERIALS AND METHODS: A tele-guided US course was established for medical students. Students underwent seven US organ modules. Each module took place in a flipped classroom concept via the Amboss platform, providing supplementary e-learning material that was optional and included information on each of the US modules. An objective structured assessment of US skills (OSAUS) was implemented as the final exam. US images of the course and exam were rated by the Brightness Mode Quality Ultrasound Imaging Examination Technique (B-QUIET). Achieved points in image rating were compared to the OSAUS exam. RESULTS: A total of 15 medical students were enrolled. Students achieved an average score of 154.5 (SD ± 11.72) out of 175 points (88.29 %) in OSAUS, which corresponded to the image rating using B-QUIET. Interrater analysis of US images showed a favorable agreement with an ICC (2.1) of 0.895 (95 % confidence interval 0.858 < ICC < 0.924). CONCLUSION: US training via teleguidance should be considered in medical education. Our pilot study demonstrates the feasibility of a concept that can be used in the future to improve US training of medical students even during a pandemic.
Sentiment Analysis (SA) aims to extract useful information from online Unstructured User-Generated Contents (UUGC) and classify them into positive and negative classes. State-of-the-art techniques for SA suffer a high dimensional feature space because of noisy and irrelevant features from the UUGC. Researchers have also proposed feature extraction and selection techniques to reduce high dimensional feature space, but they fall short in extracting and selecting the most effective sentiment features for sentiment model learning. Effective feature extraction and selection are significant for the SA because they can boost the learning algorithm’s predictive performance while reducing the high-dimensional feature space. To address these concerns, we propose an Intelligent Hybrid Feature Selection for Sentiment Analysis (IHFSSA) based on ensemble learning methods. IHFSSA first identifies sentiment features in the review text utilizing Penn Treebank part-of-speech tagset and integrated Wide Coverage Sentiment Lexicons (WCSL). The sentiment features subset is then selected employing a fast and simple rank-based ensemble of multiple filters feature selection method. The selected sentiment features are further refined by applying a wrapper-based backward feature selection method. Finally, for textual sentiment classification, the well-known classification algorithms Support Vector Machine (SVM), Naive Bayes (NB), Generalized Linear Model (GLM) are trained in the ensemble model on the refined sentiment feature set. The in-depth evaluation using heterogeneous domain benchmark datasets demonstrates that IHFSSA outperforms existing SA techniques.
Despite their continued popularity, categorical approaches to affect recognition have limitations, especially in real-life situations. Dimensional models of affect offer important advantages for the recognition of subtle expressions and more fine-grained analysis. We introduce a simple but effective facial expression analysis (FEA) system for dimensional affect, solely based on geometric features and Partial Least Squares (PLS) regression. The system jointly learns to estimate Arousal and Valence ratings from a set of facial images. The proposed approach is robust, efficient, and exhibits comparable performance to contemporary deep learning models, while requiring a fraction of the computational resources.
We propose the Recursive Non-autoregressive Graph-to-Graph Transformer architecture (RNGTr) for the iterative refinement of arbitrary graphs through the recursive application of a non-autoregressive Graph-to-Graph Transformer and apply it to syntactic dependency parsing. We demonstrate the power and effectiveness of RNGTr on several dependency corpora, using a refinement model pre-trained with BERT. We also introduce Syntactic Transformer (SynTr), a non-recursive parser similar to our refinement model. RNGTr can improve the accuracy of a variety of initial parsers on 13 languages from the Universal Dependencies Treebanks, English and Chinese Penn Treebanks, and the German CoNLL2009 corpus, even improving over the new state-of-the-art results achieved by SynTr, significantly improving the state-of-the-art for all corpora tested.
The combined use of neural scoring systems and BERT fine-tuning has led to very high results in many natural language processing (NLP) tasks. These high results raise two important questions about the contribution and the limitations of pretrained-language models: (i) what are the remaining errors in the bestperforming systems? (ii) what are the types of test examples where pretrained language models help the most? In this paper, we investigate both questions for the task of English discontinuous constituency parsing on the Penn Treebank, for which recent models obtain close to 95 F 1 score. To do so, we propose two methods for automatically analysing the errors of discontinuous parser. First, we annotate and release a test-suite focused on the syntactic phenomena responsible for discontinuities in the Penn Treebank, enabling us to obtain a per-phenomenon evaluation of a parser's output. Second, we extend the Berkeley Parser Analyser -a tool that classifies parsing errors according to predefined structural patterns -, to discontinuous trees. We apply both methods to characterize errors of a state-of-theart transition-based discontinuous parser, and to provide an overview of the contribution of BERT to this task.