Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The article examines the Universal Dependencies (UD) annotation scheme. The UD project is an international initiative to produce treebanks of the world’s languages, whereby the treebanks have been annotated in a cross-linguistically consistent manner. A central aspect of the UD annotation scheme is its analysis of function words. The scheme advocates subordinating function words to content words. This article discusses linguistic and practical motivations behind the UD decision to subordinate function words to content words. It demonstrates that UD choices in this area are not supported linguistically. At the same time, the near convertibility of the UD treebanks to a more linguistically motivated annotation format means that the UD initiative remains of great value to linguistics in general.
Abstract This chapter summarises the contributions to the volume The Normative Animal? On the Anthropological Significance of Social, Moral and Linguistic Norms. The contributions are divided into three sections in line with the tripartite division of the types of norms discussed in the volume. The key claims of the individual chapters are presented and set into relation to one another, and a number of issues raised by competition between the claims are highlighted. This prepares the ground for an assessment of the normative animal thesis in the light of the varying accounts both of specific deontic phenomena and of normativity in general. Central issues concern the concepts of social norms and conventions, the relative importance of coordination and cooperation, the nature and role of collective intentionality, the place of norms in evolutionary explanations, and the structure of normative action guidance. Decisive for the normative animal thesis are the questions as to whether moral principles and linguistic rules are correctly characterised as both real and deontic in the same senses in which these characterisations apply to social norms.
Neural parsers obtain state-of-the-art results on benchmark treebanks for constituency parsing-but to what degree do they generalize to other domains? We present three results about the generalization of neural parsers in a zero-shot setting: training on trees from one corpus and evaluating on out-of-domain corpora. First, neural and non-neural parsers generalize comparably to new domains. Second, incorporating pre-trained encoder representations into neural parsers substantially improves their performance across all domains, but does not give a larger relative improvement for out-of-domain treebanks. Finally, despite the rich input representations they learn, neural parsers still benefit from structured output prediction of output trees, yielding higher exact match accuracy and stronger generalization both to larger text spans and to out-of-domain corpora. We analyze generalization on English and Chinese corpora, and in the process obtain state-of-the-art parsing results for the Brown, Genia, and English Web treebanks.
This chapter aims to provide a large-scale collection of digital tape recordings of Cantonese speech and establishing an archive of Cantonese texts based on transcriptions of these recordings. It explains a corpus of Cantonese syllables and words together with other polysyllabic Chinese expressions. The chapter considers generation of relevant lexical information of Cantonese Chinese speech, determination of the processing and production unit of Cantonese speech, and estimation of the code-switched situation in Hong Kong. Sources of the natural Cantonese speech include dialogues of Radio call-in programs, conversations of TV programs, casual chatting among the students in canteen. The chapter also provide useful information of the pervasive code-switching situation in Hong Kong that clearly confounded the traditional language teaching methods in Hong Kong education sector.
In the present article, a novel emotional complexity marker is proposed for classification of discrete emotions induced by affective video film clips. Principal Component Analysis (PCA) is applied to full-band specific phase space trajectory matrix (PSTM) extracted from short emotional EEG segment of 6 s, then the first principal component is used to measure the level of local neuronal complexity. As well, Phase Locking Value (PLV) between right and left hemispheres is estimated for in order to observe the superiority of local neuronal complexity estimation to regional neuro-cortical connectivity measurements in clustering nine discrete emotions (fear, anger, happiness, sadness, amusement, surprise, excitement, calmness, disgust) by using Long-Short-Term-Memory Networks as deep learning applications. In tests, two groups (healthy females and males aged between 22 and 33 years old) are classified with the accuracy levels of [Formula: see text] and [Formula: see text] through the proposed emotional complexity markers and and connectivity levels in terms of PLV in amusement. The groups are found to be statistically different ( p << 0.5) in amusement with respect to both metrics, even if gender difference does not lead to different neuro-cortical functions in any of the other discrete emotional states. The high deep learning classification accuracy of [Formula: see text] is commonly obtained for discrimination of positive emotions from negative emotions through the proposed new complexity markers. Besides, considerable useful classification performance is obtained in discriminating mixed emotions from each other through full-band connectivity features. The results reveal that emotion formation is mostly influenced by individual experiences rather than gender. In detail, local neuronal complexity is mostly sensitive to the affective valance rating, while regional neuro-cortical connectivity levels are mostly sensitive to the affective arousal ratings.
BACKGROUND: The tendency to inhibit anger (anger-in) is associated with increased pain. This relationship may be explained by the negative affectivity hypothesis (anger-in increases negative affect that increases pain). Alternatively, it may be explained by the cognitive resource hypothesis (inhibiting anger limits attentional resources for pain modulation). METHODS: A well-validated picture-viewing paradigm was used in 98 healthy, pain-free individuals who were low or high on anger-in to study the effects of anger-in on emotional modulation of pain and attentional modulation of pain. Painful electrocutaneous stimulations were delivered during and in between pictures to evoke pain and the nociceptive flexion reflex (NFR; a physiological correlate of spinal nociception). Subjective and physiological measures of valence (ratings, facial/corrugator electromyogram) and arousal (ratings, skin conductance) were used to assess reactivity to pictures and emotional inhibition in the high anger-in group. RESULTS: The high anger-in group reported less unpleasantness, showed less facial displays of negative affect in response to unpleasant pictures, and reported greater arousal to the pleasant pictures. Despite this, both groups experienced similar emotional modulation of pain/NFR. By contrast, the high anger-in group did not show attentional modulation of pain. CONCLUSIONS: These findings support the cognitive resource hypothesis and suggest that overuse of emotional inhibition in high anger-in individuals could contribute to cognitive resource deficits that in turn contribute to pain risk. Moreover, anger-in likely influenced pain processing predominantly via supraspinal (e.g., cortico-cortical) mechanisms because only pain, but not NFR, was associated with anger-in.
An experiment was conducted to examine the independent and interactive influence of the audio and visual channels of information in television on viewers’ emotional experience. Audio-only, video-only, and audiovisual television content was presented as psychological stimuli, while participants completed continuous-response measures (CRMs) to index over-time changes in emotional experience of positive valence, negative valence, and arousal. Positive valence and arousal means were significantly influenced by channel over time. Participants reported the most positive emotional experience during audiovisual exposure. The channel and time interaction did not significantly affect negative valence ratings. However, positive valence, negative valence, and arousal ratings were significantly influenced by the interaction of channel and specific message content.
This paper suggests one way to enhance the ability of Chinese learners to analyze sentences. It is building a treebank and visualizing it as a syntactic tree(or parsed tree) and providing it to learners. The process of building a treebank, which is the most important key in this method, is divided into three parts and described in detail.
Animal phobias are one of the most prevalent mental disorders. We analysed how fear and disgust, two emotions involved in their onset and maintenance, are elicited by common phobic animals. In an online survey, the subjects rated 25 animal images according to elicited fear and disgust. Additionally, they completed four psychometrics, the Fear Survey Schedule II (FSS), Disgust Scale - Revised (DS-R), Snake Questionnaire (SNAQ), and Spider Questionnaire (SPQ). Based on a redundancy analysis, fear and disgust image ratings could be described by two axes, one reflecting a general negative perception of animals associated with higher FSS and DS-R scores and the second one describing a specific aversion to snakes and spiders associated with higher SNAQ and SPQ scores. The animals can be separated into five distinct clusters: (1) non-slimy invertebrates; (2) snakes; (3) mice, rats, and bats; (4) human endo- and exoparasites (intestinal helminths and louse); and (5) farm/pet animals. However, only snakes, spiders, and parasites evoke intense fear and disgust in the non-clinical population. In conclusion, rating animal images according to fear and disgust can be an alternative and reliable method to standard scales. Moreover, tendencies to overgeneralize irrational fears onto other harmless species from the same category can be used for quick animal phobia detection.
Tree-LSTMs have been used for tree-based sentiment analysis over Stanford Sentiment Treebank, which allows the sentiment signals over hierarchical phrase structures to be calculated simultaneously. However, traditional tree-LSTMs capture only the bottom-up dependencies between constituents. In this paper, we propose a tree communication model using graph convolutional neural network and graph recurrent neural network, which allows rich information exchange between phrases constituent tree. Experiments show that our model outperforms existing work on bidirectional tree-LSTMs in both accuracy and efficiency, providing more consistent predictions on phrase-level sentiments.
* Introduction This is the Myanmar ALT of the Asian Language Treebank (ALT) Corpus. Please refer to<br> http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/index.html<br> for an introduction of the ALT project. The process of building the Myanmar ALT began with sampling about 20,000 sentences from English Wikinews, and then these sentences were translated into Myanmar language.<br> <br> The English Wikinews<br> https://en.wikinews.org/wiki/Main_Page<br> is available under the terms of the Creative Commons Attribution 2.5 License.<br> https://creativecommons.org/licenses/by/2.5/ Myanmar ALT has been developed by NICT and UCSY. The license of Myanmar ALT is Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License<br> https://creativecommons.org/licenses/by-nc-sa/4.0/ <br> * Contents - data: Myanmar ALT treebank
We present our CHARLES-SAARLAND system for the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology, in task 2, Morphological Analysis and Lemmatization in Context. We leverage the multilingual BERT model and apply several fine-tuning strategies introduced by UDify demonstrating exceptional evaluation performance on morpho-syntactic tasks. Our results show that fine-tuning multilingual BERT on the concatenation of all available treebanks allows the model to learn cross-lingual information that is able to boost lemmatization and morphology tagging accuracy over fine-tuning it purely monolingually. Unlike UDify, however, we show that when paired with additional character-level and word-level LSTM layers, a second stage of fine-tuning on each treebank individually can improve evaluation even further. Out of all submissions for this shared task, our system achieves the highest average accuracy and f1 score in morphology tagging and places second in average lemmatization accuracy.
The communicative role of nonlinear vocal phenomena remains poorly understood since they are difficult to manipulate or even measure with conventional tools. In this study parametric voice synthesis was employed to add pitch jumps, subharmonics/sidebands, and chaos to synthetic human nonverbal vocalizations. In Experiment 1 (86 participants, 144 sounds), chaos was associated with lower valence, and subharmonics with higher dominance. Arousal ratings were not noticeably affected by any nonlinear effects, except for a marginal effect of subharmonics. These findings were extended in Experiment 2 (83 participants, 212 sounds) using ratings on discrete emotions. Listeners associated pitch jumps, subharmonics, and especially chaos with aversive states such as fear and pain. The effects of manipulations in both experiments were particularly strong for ambiguous vocalizations, such as moans and gasps, and could not be explained by a non-specific measure of spectral noise (harmonics-to-noise ratio) – that is, they would be missed by a conventional acoustic analysis. In conclusion, listeners interpret nonlinear vocal phenomena quite flexibly, depending on their type and the kind of vocalization in which they occur. These results showcase the utility of parametric voice synthesis and highlight the need for a more fine-grained analysis of voice quality in acoustic research.
This paper suggests one way to enhance the ability of Chinese learners to analyze sentences. It is building a treebank and visualizing it as a syntactic tree(or parsed tree) and providing it to learners. The process of building a treebank, which is the most important key in this method, is divided into three parts and described in detail.
This chapter is the first large-scale typological survey of the lexical means used in African languages to express color-related meanings. It is based on a very large sample, with data from 350 languages, most of which come from the RefLex online lexical database. It focuses on language-internal semantic sources, morphosyntactic strategies, and contact-induced terminology used for color naming. After a brief discussion of the issues raised by “basic” color terms, and “polychromatic” color terms, the chapter provides a review of the semantic sources of color terms, the origin of borrowings, colexifications and metaphorical uses of color terms, main patterns of lexicalization, and, briefly, color-related ideophones.
Callous-unemotional (CU) traits are associated with lower emotional reactivity in adolescents. However, since previous studies have focused mainly on reactivity to negative stimuli, it is unclear whether reactivity to positive stimuli is also affected. Further, few studies have addressed the link between CU traits and emotional reactivity in longitudinal community samples, which is important for determining its generalizability and developmental course. In the current study, pupil dilation and self-ratings of arousal and valence were assessed in 100 adolescents (15-17 years) from a community sample, while viewing images with negative and positive valence from the International Affective Pictures System (IAPS). Behavioral traits (CU) were assessed concurrently, as well as at ages 12-15, and 8-9 (subsample, n = 68, low levels of prosocial behavior were used as a proxy for CU traits). The results demonstrate that CU traits assessed at ages 12-15 and 8-9 predicted less pupil dilation to both positive and negative images at ages 15-17. Further, CU traits at ages 12-15 and concurrently were associated with less negative valence ratings for negative images and concurrently to less positive valence ratings for positive images. The current findings demonstrate that CU traits are related to lower emotional reactivity to both negative and positive stimuli in adolescents from a community sample.
Affective states underlie daily decision-making and pathological behaviours relevant to obsessive-compulsive disorders (OCD), mood disorders and addictions. Deep brain stimulation targeting the motor and associative-limbic subthalamic nucleus (STN) has been shown to be effective for Parkinson's disease (PD) and OCD, respectively. Cognitive and electrophysiological studies in PD showed responses of the motor STN to emotional stimuli, impairments in recognition of negative affective states and modulation of the intensity of subjective emotion. Here we studied whether the stimulation of the associative-limbic STN in OCD influences the subjective emotion to low-intensity positive and negative images and how this relates to clinical symptoms. We assessed 10 OCD patients with on and off STN DBS in a double-blind randomized manner by recording ratings of valence and arousal to low- and high-intensity positive and negative emotional images. STN stimulation increased positive ratings and decreased negative ratings to low-intensity positive and negative stimuli, respectively, relative to off stimulation. We also show that the change in severity of obsessive-compulsive symptoms pre- versus post-operatively interacts with both DBS and valence ratings. We show that stimulation of the associative-limbic STN might influence the negative cognitive bias in OCD and decreasing the negative appraisal of emotional stimuli with a possible relationship with clinical outcomes. That the effect is specific to low intensity might suggest a role of uncertainty or conflict related to competing interpretations of image intensity. These findings may have implications for the therapeutic efficacy of DBS.
The purpose of this paper is to examine the characteristics of the additions of sounds as a product of changing North Korean linguistic norms reflected in Korean and grammar textbooks. The writing of Saisiot from the Enlightenment period to the Japanese colonial period have been a topic of frequent discussion. The publication of “The Theory of Korean” in the Enlightenment period established the Hangeul writing system, so writings produced using various writing systems were found in contemporary textbooks. “A Proposal for a Unified Korean Spelling System” was later published by the Japanese Government-General of Korea, but Korean spoken in North Korea has changed since the national liberation from Japanese colonial rule. Today’s North Korean grammar textbooks differentiate between the addition of ㄴand ㄷ. These additions are known as the Saisori phenomenon. After liberation, “jeoleumbu(絶音符)” or “saipyo(’)” were written between compound words, but this is no longer the case. This difference is one example of how Korean in North and South Korea differs with regard to the Saisori phenomenon, which is known as the addition of sound, and the writing of Saisiot. This study analyzed North Korean and grammar textbooks to determine how North Korean linguistic norms have changed over time.
Abstract The goal of this study is to demonstrate how network science and graph theory tools and concepts can be effectively used for exploring and comparing semantic spaces of word embeddings and lexical databases. Specifically, we construct semantic networks based on word2vec representation of words, which is “learnt” from large text corpora (Google news, Amazon reviews), and “human built” word networks derived from the well-known lexical databases: WordNet and Moby Thesaurus. We compare “global” (e.g., degrees, distances, clustering coefficients) and “local” (e.g., most central nodes and community-type dense clusters) characteristics of considered networks. Our observations suggest that human built networks possess more intuitive global connectivity patterns, whereas local characteristics (in particular, dense clusters) of the machine built networks provide much richer information on the contextual usage and perceived meanings of words, which reveals interesting structural differences between human built and machine built semantic networks. To our knowledge, this is the first study that uses graph theory and network science in the considered context; therefore, we also provide interesting examples and discuss potential research directions that may motivate further research on the synthesis of lexicographic and machine learning based tools and lead to new insights in this area.
Global Bible Initiative (GBI) have developed Hebrew OT treebanks and Greek NT syntactic treebanks. The treebanks were first generated with a parser using computerized Hebrew and Greek grammars and then proofed verse by verse by Hebrew and Greek Scholars. All the corrections made by the scholars were kept as disambiguation data.
 The phrase structures in the trees have been used to build interlinears, concordances, and translation memories which operate not only on the word level, but on the phrase and clause levels as well. The syntactic relations (dependencies) in the trees have also been used to do smart search where we can find texts that are different in form but similar in meaning.
 Recently, we have also used the trees to improve the accuracy of automatic word alignment and explore tree-based interactive machine translation of the Bible. The auto aligner can be used to the Hebrew and Greek texts to translations in various languages. The interactive machine translation will speed up Bible translation without compromising quality by providing real time suggestions and checking.
 We have already contributed two sets of Greek trees to Creative Commons, the Nestle 1904 version and the SBLGNT version. We also have trees for NA27 and NA28, but we do not own the texts. The Hebrew OT treebank we developed was owned by the Groves Center. We are also capable of creating new treebanks with the parser, grammar, and disambiguation data we own if we are given a text that is morphologically tagged.
It is possible to include complicated structures into an individual syntactic tree, to enhance the usefulness of parsed text corpus. In this part, existing works on Thai treebank construction have been developed in order to address the lack of high-level syntactic resources. However, it has yet to be sufficient for Thai Natural Language Processing. Furthermore, Thai treebanks have either syntactic or dependency structure only. This paper presents a construction of hybrid structural Thai treebank which includes both syntactic/dependency structure, a tool for conversion between constituency and dependency parse tree, and a web-based GUI for parse tree visualization. Towards the hybrid treebank construction, hundreds of constituent tree are manually annotated with predicate header to each phrase. Once the set of annotated constituent trees are obtained, the conversion procedure will be performed by determining the annotated head and its dependents. As our experiments, features of hybrid treebank are extracted and illustrated. Finally, difficulties and issues in constructing the hybrid Thai treebank are discussed.
This paper proposes a novel Recurrent Neural Network (RNN) language model that takes advantage of character information. We focus on character n-grams based on research in the field of word embedding construction (Wieting et al. 2016). Our proposed method constructs word embeddings from character ngram embeddings and combines them with ordinary word embeddings. We demonstrate that the proposed method achieves the best perplexities on the language modeling datasets: Penn Treebank, WikiText-2, and WikiText-103. Moreover, we conduct experiments on application tasks: machine translation and headline generation. The experimental results indicate that our proposed method also positively affects these tasks
Counterconditioning (CC) is a form of retroactive interference that inhibits expression of learned behavior. But similar to extinction, CC can be a fairly weak and impermanent form of interference, and the original behavior is prone to relapse. Research on CC is limited, especially in humans, but prior studies suggest it is more effective than extinction at modifying some behaviors (e.g., preference or valence ratings) than others (e.g., physiological arousal). Here, we used a within-subjects design to compare the effects of aversive-to-appetitive CC versus standard extinction on two separate tests of long-term memory in human adults: implicit physiological arousal and explicit episodic memory. Participants underwent Pavlovian fear conditioning to two semantic categories (animals, tools) paired with an electric shock. Conditioned stimuli (i.e., category exemplars) from one category were then extinguished, while stimuli from the other category were paired with a positive outcome. Participants returned 24-h later for a test of skin conductance responses (SCR) to the conditioned exemplars, as well as a surprise recognition memory test for stimuli encoded the previous day. Results showed reduced SCRs at a test for unique stimuli from a category that had undergone CC, relative to stimuli from a category that had undergone standard extinction. Additionally, participants selectively remembered more stimuli encoded during CC than extinction. These results provide new evidence that aversive-to-appetitive CC, as compared to extinction, strengthens memory for items directly associated with a positive outcome, which may provide stronger retrieval competition against a fear memory at test to help diminish fear relapse.
xml treebank Annotated by Toon Van Hal, with student contributions by Mathieu Cuijpers; Sanderijn Gijbels; Yoran Joosten; Yordi Lenaerts; Eva Uffing; Chiara Van der Hasselt; Lisa Vanhee and Jolien Volders (KU Leuven Bachelor 3, 2018-2019). Based on a preparsed text by Alek Keersmaekers. Controlled by Toon Van Hal, Sanderijn Gijbels and Yoran Joosten.
Despite the fact that there are a number of researches working on Khmer Language in the field of Natural Language Processing along with some resources regarding words segmentation and POS Tagging, we still lack of high-level resources regarding syntax, Treebanks and grammars, for example. This paper illustrates the semi-automatic framework of constructing Khmer Treebank and the extraction of the Khmer grammar rules from a set of sentences taken from the Khmer grammar books. Initially, these sentences will be manually annotated and processed to generate a number of grammar rules with their probabilities once the Treebank is obtained. In our experiments, the annotated trees and the extracted grammar rules are analyzed in both quantitative and qualitative way. Finally, the results will be evaluated in three evaluation processes including Self-Consistency, 5-Fold Cross-Validation, Leave-One-Out Cross-Validation along with the three validation methods such as Precision, Recall, F1-Measure. According to the result of the three validations, Self-Consistency has shown the best result with more than 92%, followed by the Leave-One-Out Cross-Validation and 5-Fold Cross Validation with the average of 88% and 75% respectively. On the other hand, the crossing bracket data shows that Leave-One-Out Cross Validation holds the highest average with 96% while the other two are 85% and 89%, respectively.
Imagining fictional creatures like zombies in survival situations boosts long-term memory for words encoded in these situations more than rating words for pleasantness (zombie effect). Study 1 required word-ratings in a zombie-survival scenario; participants were told they had to protect against either possible zombie attack or contamination. The zombie-survival situations yielded identical recall levels but higher recall rates than pleasantness. Study 2 matched a zombie-survival scenario on perceived fear with scenarios involving ghosts or predators. Perceived disgust in the zombie scenario was higher than in these other survival conditions. Words were remembered better when processed in survival scenarios than when rated for pleasantness, but there was no reliable difference in recall between the scenarios. In neither study did the number of death-related words produced in a word-fragment completion task fit the mortality salience account of the zombie memory effect. Overall findings suggest that this effect relates to the fear system.
Word vectors are at the core of many natural language processing tasks. Recently, there has been interest in post-processing word vectors to enrich their semantic information. In this paper, we introduce a novel word vector post-processing technique based on matrix conceptors (Jaeger 2014), a family of regularized identity maps. More concretely, we propose to use conceptors to suppress those latent features of word vectors having high variances. The proposed method is purely unsupervised: it does not rely on any corpus or external linguistic database. We evaluate the post-processed word vectors on a battery of intrinsic lexical evaluation tasks, showing that the proposed method consistently outperforms existing state-of-the-art alternatives. We also show that post-processed word vectors can be used for the downstream natural language processing task of dialogue state tracking, yielding improved results in different dialogue domains.
Recent studies have compared tinnitus suppression, or residual inhibition, between amplitude- and frequency-modulated (AM) sounds and noises or pure tones (PT). Results are indicative, yet inconclusive, of stronger tinnitus suppression of modulated sounds especially near the tinnitus frequency. Systematic comparison of AM sounds at the tinnitus frequency has not yet been studied in depth. The current study therefore aims at further advancing this line of research by contrasting tinnitus suppression profiles of AM and PT sounds at the matched tinnitus frequency (i.e., 10 and 40 Hz AM vs. PT). Participants with chronic, tonal tinnitus (n = 29) underwent comprehensive psychometric, audiometric, tinnitus matching, and acoustic stimulation procedures. Stimuli were presented for 3 minutes in two loudness regimes (60 dB sensation level [SL], minimum masking level [MML] + 6 dB, control sound: SL -6 dB) and amplitude modulated with 0, 10, or 40 Hz. Tinnitus loudness suppression was measured after the stimulation every 30 seconds. In addition, stimuli were rated regarding their valence and arousal. Results demonstrate only trends for better tinnitus suppression for the 10 Hz modulation and presentation level of 60 dB SL compared with PT, whereas nonsignificant results are reported for 40 Hz and MML + 6 dB, respectively. Furthermore, the 10 Hz AM at 60 dB SL and the 40 Hz AM at MML + 6 dB (trend) stimuli were better tolerated as elicited by valence ratings. We conclude that 10 Hz AM sounds at the tinnitus frequency may be useful to further elucidate the phenomenon of residual inhibition.
Discourse-annotated corpora are an important resource for the community, but they are often annotated according to different frameworks. This makes joint usage of the annotations difficult, preventing researchers from searching the corpora in a unified way, or using all annotated data jointly to train computational systems. Several theoretical proposals have recently been made for mapping the relational labels of different frameworks to each other, but these proposals have so far not been validated against existing annotations. The two largest discourse relation annotated resources, the Penn Discourse Treebank and the Rhetorical Structure Theory Discourse Treebank, have however been annotated on the same texts, allowing for a direct comparison of the annotation layers. We propose a method for automatically aligning the discourse segments, and then evaluate existing mapping proposals by comparing the empirically observed against the proposed mappings. Our analysis highlights the influence of segmentation on subsequent discourse relation labelling, and shows that while agreement between frameworks is reasonable for explicit relations, agreement on implicit relations is low. We identify several sources of systematic discrepancies between the two annotation schemes and discuss consequences for future annotation and for usage of the existing resources.
In this paper, we compute the affective-aesthetic potential (AAP) of literary texts by using a simple sentiment analysis tool called SentiArt. In contrast to other established tools, SentiArt is based on publicly available vector space models (VSMs) and requires no emotional dictionary, thus making it applicable in any language for which VSMs have been made available (>150 so far) and avoiding issues of low coverage. In a first study, the AAP values of all words of a widely used lexical databank for German were computed and the VSM’s ability in representing concrete and more abstract semantic concepts was demonstrated. In a second study, SentiArt was used to predict ~2800 human word valence ratings and shown to have a high predictive accuracy (R2 > 0.5, p < 0.0001). A third study tested the validity of SentiArt in predicting emotional states over (narrative) time using human liking ratings from reading a story. Again, the predictive accuracy was highly significant: R2adj = 0.46, p < 0.0001, establishing the SentiArt tool as a promising candidate for lexical sentiment analyses at both the micro- and macrolevels, i.e., short and long literary materials. Possibilities and limitations of lexical VSM-based sentiment analyses of diverse complex literary texts are discussed in the light of these results.
Recurrent neural network language models (RNNLMs) have shown superior performance across a range of speech recognition tasks. At the heart of all RNNLMs, the activation functions play a vital role to control the information flows and tracking longer history contexts that are useful for predicting the following words. Long short-term memory (LSTM) units are well known for such ability and thus widely used in current RNNLMs. However, the deterministic parameter estimates in LSTM RNNLMs are prone to over-fitting and poor generalization when given limited training data. Furthermore, the precise forms of activations in LSTM have been largely empirically set for all cells at a global level. In order to address these issues, this paper introduces Gaussian process (GP) LSTM RNNLMs. In addition to modeling parameter uncertainty under a Bayesian framework, it also allows the optimal forms of gates being automatically learned for individual LSTM cells. Experiments were conducted on three tasks: the Penn Treebank (PTB) corpus, Switchboard conversational telephone speech (SWBD) and the AMI meeting room data. The proposed GP-LSTM RNNLMs consistently outperform the baseline LSTM RNNLMs in terms of both perplexity and word error rate.
BACKGROUND: Neighbourhood environment characteristics have been found to be associated with residents' willingness to conduct physical activity (PA). Traditional methods to assess perceived neighbourhood environment characteristics are often subjective, costly, and time-consuming, and can be applied only on a small scale. Recent developments in deep learning algorithms and the recent availability of street view images enable researchers to assess multiple aspects of neighbourhood environment perceptions more efficiently on a large scale. This study aims to examine the relationship between each of six neighbourhood environment perceptual indicators-namely, wealthy, safe, lively, depressing, boring and beautiful-and residents' time spent on PA in Guangzhou, China. METHODS: A human-machine adversarial scoring system was developed to predict perceptions of neighbourhood environments based on Tencent Street View imagery and deep learning techniques. Image segmentation was conducted using a fully convolutional neural network (FCN-8s) and annotated ADE20k data. A human-machine adversarial scoring system was constructed based on a random forest model and image ratings by 30 volunteers. Multilevel linear regressions were used to examine the association between each of the six indicators and time spent on PA among 808 residents living in 35 neighbourhoods. RESULTS: Total PA time was positively associated with the scores for "safe" [Coef. = 1.495, SE = 0.558], "lively" [1.635, 0.789] and "beautiful" [1.009, 0.404]. It was negatively associated with the scores for "depressing" [- 1.232, 0.588] and "boring" [- 1.227, 0.603]. No significant linkage was found between total PA time and the "wealthy" score. PA was further categorised into three intensity levels. More neighbourhood perceptual indicators were associated with higher intensity PA. The scores for "safe" and "depressing" were significantly related to all three intensity levels of PA. CONCLUSIONS: People living in perceived safe, lively and beautiful neighbourhoods were more likely to engage in PA, and people living in perceived boring and depressing neighbourhoods were less likely to engage in PA. Additionally, the relationship between neighbourhood perception and PA varies across different PA intensity levels. A combination of Tencent Street View imagery and deep learning techniques provides an accurate tool to automatically assess neighbourhood environment exposure for Chinese large cities.
This is the first versioned collection of.xml files containing approximately 550,000 tokens of ancient Greek prose that have been hand-analyzed into dependency syntax using Perseids/Arethusa by Prof. Vanessa Gorman of the University of Nebraska-Lincoln. CC0 1.0 license.
We present a novel semantic framework for modeling linguistic expressions of\ngeneralization---generic, habitual, and episodic statements---as combinations\nof simple, real-valued referential properties of predicates and their\narguments. We use this framework to construct a dataset covering the entirety\nof the Universal Dependencies English Web Treebank. We use this dataset to\nprobe the efficacy of type-level and token-level information---including\nhand-engineered features and static (GloVe) and contextual (ELMo) word\nembeddings---for predicting expressions of generalization. Data and code are\navailable at decomp.io.\n
In this paper, we propose a novel approach to address the newly defined needs of linguistic typology recently interested in fine-grained features underlying language diversity. In fact, we introduce a method to extract qualitative and quantitative information about a wide range of features from multilingual annotated corpora based on Natural Language Processing methods and techniques. We tested our method in a case study focusing on word order variation in two widely investigated constructions, VERB-SUBJ(ect) and NOUN-ADJ(ective), with a specific view to structural and functional factors underlying the preference for one or the other order, both intra- and cross-linguistically, and their interaction. Preliminary experiments have been carried out aimed at acquiring typological evidence from a selection of linguistically annotated treebanks for three different languages, namely Italian, Spanish and English.