Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Syntactic parsing plays a crucial role in improving the quality of natural language processing tasks. Although there have been several research projects on syntactic parsing in Vietnamese, the parsing quality has been far inferior than those reported in major languages, such as English and Chinese. In this work, we evaluated representative constituency parsing models on a Vietnamese Treebank to look for the most suitable parsing method for Vietnamese. We then combined the advantages of automatic and manual analysis to investigate errors produced by the experimented parsers and find the reasons for them. Our analysis focused on three possible sources of parsing errors, namely limited training data, part-of-speech (POS) tagging errors, and ambiguous constructions. As a result, we found that the last two sources, which frequently appear in Vietnamese text, significantly attributed to the poor performance of Vietnamese parsing.
The excellent back-of-the book index contains terms that navigate readers to find the indexed term on the important book page and related to the content of the book. The conceptual characteristic of indexed term is one of the difficulties that faced by the professional book indexers. They should ensure that the indexed term directs the reader to a book page with the relevant information so the readers can understand the meaning or information in the indexed term. This study aims to identify the relevance of the indexed term. We consider the relatedness of the text containing the indexed terms with the book domain. We calculate the semantic relatedness of the data set, using the WordNet lexical database. We have evaluated the performance of semantic-relatedness method on books from the knowledge domain of Biology, Religion and Computer Science using Kappa Statistic Values. The result implies that the relatedness to the knowledge domain of the book influences the relevance of an indexed term. The Kappa values that obtained from the results test show that this semantic relatedness approach is reliable and could contribute to identify the relevant indexed term.
is known for his bold experimentation with poetic forms and eccentric deviation from linguistic norms in his visual love poetry.The peculiar distribution of lines and unerring rhetorical skills are given special zest and prominence by virtue of his twin obsessions: poetry and painting.The current study is premised on two tenets: first, foregrounding is the dominant feature of Cummings' narrative poetry; second, the multitude of semiotic resources and their division of labor in his poetic texts, coupled with the wittily expressed attitudinal values, add new layers of discourse that are worthy of investigation.The aim of the research endeavor at hand is to unravel the poetic effects of the foregrounding devices employed in the love poem "all in green went my love riding".To this end, the meta-functions of Systemic Functional Theory and Visual Grammar are used as the basis of analysis.Various types of deviation and regular patterns, in terms of repetition and parallelism, are combined to enable smooth narrative flow while engaging the reader in the setting and the action from start to finish.Linguistic deviation effectively serves foregrounding in the poem.This, in turn, tempts readers to reach valid interpretations of the poem overall.The poem is an exquisite mixing bowl of counter-grammatical devices and regularities within the syntactic texture of each of the fourteen stanzas.The researcher argues that even with the less deviant of Cummings' poems, textual and visual resources in the semiotic ensemble of the poetic text are carriers of potential meaning serving the overall structure of the narrative genre.
Among the vast literature on discourse markers (henceforth DMs) and discourse-relational devices in general, one aspect of their behaviour has been somewhat overlooked, namely their co-occurrence. It is frequently the case that two or more DMs co-occur, as in the case of so for instance if, where DMs only co-occur or are juxtaposed, or in the case of but actually, and yet, but look, etc., where they actually combine. DM co-occurrence is a multi-faceted phenomenon, since not all cases display the same degree of integration: most authors distinguish between at least two types of co-occurrence, namely addition vs. composition, depending on a number of syntactic and functional criteria (see, e.g., Luscher 1993; Hansen 1998; Pons 2008, in press; Cuenca & Marín 2009). Discourse analysis and corpus annotation show that this phenomenon is quite pervasive: 20% of all occurrences are coded as part of a co-occurring string in Crible’s (2017) corpus study of spoken English and French. In fact, DM co-occurrence poses a challenge for corpus annotation since i) it is not always clear whether two co-occurring DMs remain independent from each other or whether they should be considered as one token, and ii) senses can be influenced by co-occurring DMs during disambiguation. This study sets out to provide clear criteria for different degrees of co-occurrence on the basis of corpus-based examples. Previous papers on the subject propose several criteria to distinguish different degrees of integration. Luscher (1993) uses syntactic and semantic scope to distinguish between “additive” and “compositional” sequences. He defines the latter as applying to two adjacent DMs which are semantically similar (e.g. French mais pourtant ‘but however’), one of them being more restricted or specific in its meaning than the other. This latter type is the focus of Fraser’s (2013) study targeting English contrastive connectives. Hansen’s (1998) distinction between summative and combinatory sequences adopts a different perspective and depends on whether the elements in the sequence retain their individual meaning (French ah bon ‘oh really’) or form a new complex one (eh bien ‘well’). She argues that most DM sequences are summative (or compositional), since it is always possible to reconstruct the meaning of each element. Similarly, Pons (2008) concludes from his analysis of the co-occurrences of the Spanish modal marker bueno with other discourse markers that discourse segmentation of oral discourse allows to differentiate two different configurations: the cases in which the two markers are simply adjacent from the cases in with they combine according to whether they apply to different or to a unique structural unit. More recently, Dostie (2013) and Crible (2015) consider other types of cues in DM use that provide evidence for stronger degrees of combination, such as phonological reduction (eh bien to eh ben), new spellings (ou sinon ‘or else’ to aussi non) and new contexts of use (initial to final position for ou sinon). Cuenca & Marín (2009) discuss and illustrate a three-fold distinction in a corpus of spoken Spanish and Catalan, namely: • juxtaposition, when the DMs do not combine syntactically nor semantically (typically two conjunctions); • addition, when the DMs combine locally but their functions remain distinct (typically conjunctions followed by parenthetical connectives that jointly connect at a local level); • composition, when the DMs function as one unit (typically two parenthetical connective units with a single global-level function). Their analysis is very fine-grained and identifies recurrent formal and functional tendencies for each of these levels. Crible (2017) attempted to apply Cuenca & Marín’s (2009) classification through systematic annotation and was confronted with problematic, borderline cases (e.g. and so or et alors ‘and then’) which raised concerns about some features, pointing especially at the fuzzy border between addition and composition. Crible also discusses the role of frequency in the definition of these levels, and suggests an additional degree to deal with cases of “reinforcement” (e.g. but in fact). Her study draws the attention to the consequences of an adequate treatment of DM co-occurrence for corpus annotation (token identification and sense disambiguation). Similarly, in the guidelines of the Penn Discourse TreeBank 2.0, Prasad et al. (2007) mention that multiple (i.e. co-occurring) connectives should ideally be annotated as such and differentiated according to the (in)dependence of their elements in order to improve predictive features and classifiers. The purpose of this study is to revisit Cuenca & Marín’s (2009) three-fold classification and refine the criteria to distinguish each degree of co-occurrence, in order to be able to apply them systematically to corpus data. To this end, we used a sample of English conversational data from the DisFrEn dataset where DMs were already identified (Crible 2017): 71 DM clusters were thus extracted, from a total of 17,479 words (about 90 minutes of recordings). We did not consider phrasal DMs (e.g. I mean) as co-occurring. We further excluded cases where two DMs belong to different units (final position of the first unit, initial position of the second one, as in I like winter actually but I prefer spring). For each cluster, we manually encoded the following features: number of elements in the cluster, syntactic category (based on Cuenca 2013: conjunction, parenthetical connective, pragmatic connective, interjection), scope (same or different), position (initial or medial). We then discussed whether the elements of the cluster expressed the same meaning (or function) or not, and what degree of co-occurrence they represented. Thanks to this qualitative analysis, we were able to distinguish between criteria (necessary conditions) and features (quantitative tendencies): we found that considerations of scope and of function are criterial in the definition of the levels, whereas prosody (i.e. contiguous pause) and syntactic categories are mere tendencies. As a result, the revised cline of co-occurrence is the following: • juxtaposition, when the DMs take scope over different units (mostly when a coordinating conjunction and a subordinating conjunction co-occur); • addition, when the DMs have the same scope but clearly distinct meanings; • composition, when the DMs have the same scope and the same overall meaning, but one of them is more specific than the other (reinforcement effect); • lexicalization: a new meaning arises from the co-occurrence which is not the sum of its parts, and removing one element changes the meaning of the cluster (often with semantic bleaching and phonological reduction). This proposal takes into account the dynamicity of language and phenomena such as layering and stratification, related to polyfunctionality and underspecification. For instance, in English the highly frequent cluster and then instantiates different configurations and degrees of integration depending on the semantic status of the temporal adverbial. The first (and most frequent) use of and then (1) is an addition of the additive conjunction and the temporal adverb. In another related use (2), the elements add to express consequence, a meaning which can be derived – but differs – from the temporal meaning of then (‘at that time’). Lastly, and then (3) can express one global function of continuity or enumeration at discourse level (i.e. not temporality between facts) with contrastive nuances, in which case the co-occurrence is somewhere in-between the space of composition and lexicalization, since the meaning of the cluster is not (strictly) the sum of its parts. (1) they buy the book say for a couple of pounds (1.420) and then return it and get half (2) I've got people coming I'll get some salmon from the stall and when you get down there you find he hasn't actually got any and then it throws you into a complete quandary (3) people do tend to describe themselves […] a lot of people describe people as jealous […] and then there are the really bland ones It can be concluded that a single co-occurrence and then can instantiate different categorical configurations and can also vary along the cline of co-occurrence, thus advocating for a flexible, context-bound approach to the issue in future annotation endeavours. These distinctions are subtle and highly context-bound, yet they can and should be systematically accounted for, especially since and then is also quite frequent in writing (cf. but then or so for instance, mentioned in the PDTB guidelines). Additional features (e.g. prosody, length and type of host unit) can be investigated to further support this flexible portrait of and then. To conclude, in line with Crible & Cuenca (2017), we suggest that DM annotation endeavours should consider including information about co-occurrence, minimally by identifying clusters, ideally by distinguishing between degrees of integration following the criteria that we have developed in this study. This is particularly crucial for sequences such as and then (and its cross-linguistic equivalents, e.g. French et puis), which do not display a unique functional profile depending on co-occurrence degree. Our criteria and analysis paves the way for fruitful comparisons across spoken and written languages.
Background: The study of the literary monuments of the New Bulgarian period is one of the challenging issues of the modern historical grammar of the Bulgarian language. It helps to trace as regular and normative the formation and consolidation of those qualitative changes of the Bulgarian linguistic system, which testify to its transition from synthetism to analyticism, as well as give it the status of Balkan and “exotic”. The growth of the scientific interest to these questions proves the emergence of numerous works devoted to a multidimensional study of the manuscripts of the New Bulgarian period, which had not been catalogued yet. Purpose: The purpose of the study is to discover, on the basis of a comparative analysis, the main linguistic features of the Bulgarian damaskins of the XVIII century on the different linguistic levels; to define a set of the most characteristic rarities (archaisms), the linguistic features inherited from the traditional literary language of the church-religious heritage of the previous era as well as “nourished” by the influence of the prestigious Church Slavonic language, which was realized due to the widespread circulation of printed liturgical Church Slavonic books; to identify a complex of progressive innovative processes and phenomena in the linguistic structure that penetrated into the written language from the folk speech; to outline the corpus of linguistic and extralinguistic factors that provoked and supported a similar symbiosis of linguistic means in the written monuments of the New Bulgarian period. Results: The ratio of the concentration of rarities and innovations in the language of the new Bulgarian damaskins was regulated by the notions of the scribes about the linguistic norm (which, on the one hand, blocked the penetration of such features of the folk speech as an article, reprise, etc., and, on the other hand, could not restrain the expansion of some other new Bulgarian innovations as the so called “da”-constructions, analytical way of expressing the relationship between words, etc.), determined by the genre, stylistic and artistic peculiarities of texts, depended on the recipient and the pragmatic purpose. Key words: archaism, innovation, damaskin, the New Bulgarian period, the traditional literary language of the Resava orthography, vernacular language.
This article describes variants of the plural genitive of masculine nouns with variant endings – ów//-y (more seldom -i). Personal nouns are syncretic with accusative forms, i.e. czarodziei // czarodziejów; kuracjuszy // kuracjuszów; przybyszy // przybyszów. 40 pairs of variants of the type krokodyli // krokodylów were analysed. They were verified in terms of linguistic norm and frequency. The first verification was based on the determinations made in The Great Dictionary of Correct Polish Usage (GDCPU) published by PWN (in Polish: Wielki słownik poprawnej polszczyzny PWN), while the second one was based on The National Corpus of the Polish Language available in an electronic form.Out of 40 pairs of nouns, only 9 pairs have the ending -ów, which accounts for 22.5%. These are personal nouns (cywilów, czarodziejów, pedofilów, przybyszów, uczniów) and inanimate impersonal nouns (napojów, rodzajów, słojów, zwojów). Some of them are more frequent, as they are: 1. the only correct forms (pedofilów, zwojów); 2. non-colloquial, meaning they are recommended to be used in a carefully spoken and written Polish language (rodzajów, uczniów; colloquially rodzai, uczni); 3. non-jargonistic (cywilów; jargonistic – cywili); 4. regarded as seldom (przybyszy). Out of 8 pairs of forms treated by GDCPU as totally variant, only 3 pairs have the ending -ów (czarodziejów, napojów, słojów). The frequency of two first pairs is more than 90%, while the frequency of the form słoi is 77%. The share of napoi is only 3.6%, although it is also a form of a verb napoić (i.e. On napoi konie). It may seem that a scope of variants of the analysed forms is large. However, a considerable part of doublets are forms that are considered seldom, archaic, colloquial, jargonistic or inconsistent with the language norm. Thus, the thesis of a substantial variant occurrence of such types of forms should be verified.
Demand-withdraw is an ineffective communication pattern frequently experienced by distressed couples. Therapists often attempt to address this pattern by helping partners understand and regulate the emotions that underlie these behaviors. To date, there is a lack of research focusing on the emotional experiences underlying the demand-withdraw pattern of interaction in couples. Related lines of research focus on emotional arousal and the expression of hard and soft emotions, but this research does not specifically investigate demand-withdraw interactions. The purpose of this study is to identify what emotions underlie demanding behavior in both men and women during marital demand-withdraw conflict interactions. Six couples were chosen from a five-year longitudinal randomized clinical trial that compared Integrative Behavioral Couple Therapy (IBCT) and Traditional Behavioral Couple Therapy (TBCT). Researchers viewed 10-minute pre-treatment problem-solving interactions to observe the demand-withdraw pattern in vivo among couples seeking therapy. The Behavioral Affective Rating Scale (BARS) was used to code the emotions observed during the interactions. The results indicated that the types of emotions varied not only depending on who initiated the problem-solving interaction (e.g., wife topic-husband topic) but also between the different couples, and when comparing gender. Anxiety (#2) and aggression (#4) were in the top four most commonly observed emotions for husbands, while they were two of the least observed emotions for wives. Moreover, frustration and hurt were the two most observed emotions for wives, while they were the least observed emotions for husbands.
Fibromyalgia syndrome (FMS), a common chronic pain condition, is often incompletely treated by conventional medical therapies. It can cause disability, psychological distress, work-related absenteeism, increased use of healthcare resources, and result in the inability to carry out the tasks of daily living. The purpose of this quantitative, correlational study was to investigate the potential influence of laughter on affect and pain in individuals with FMS. Laughter produces beneficial effects on acute pain and on chronic pain in general and has been found to improve temporary affective states, but there have been no studies testing the effects of laughter on the pain and affect of fibromyalgia patients. Informing this study were the gate control and neuromatrix theories of pain, as well as the dynamic model of affect theory. The research questions addressed whether laughter frequency is associated with affect and or with perceived chronic pain levels in these individuals. Forty-one adult fibromyalgia patients documented all laughter episodes daily and assessed their pain and affective states 3 times per day for 14 days. Hierarchical regressions revealed that increased overall laughter frequency was significantly associated with decreases in overall pain and increases in overall positive affect but was not associated with measures of negative affect. Also, morning laughter frequency was predictive of increased afternoon and evening positive affect ratings, as well as with decreased afternoon pain ratings, but was not significantly associated with evening pain ratings. The knowledge gained from these results may have positive social change implications at the individual level, within those individuals' larger social networks, and within the research and medical communities.
The Brulex directory includes the different versions fo the BRULEX database. BRULEX was developed about ten years ago, for the purpose of facilitating selection and control of experimental materials in psycholinguistic experiments on lexical processing. It contains a large number of informations that were typed manually by benevolent collaborators. Potential users should be aware that the current version, which hasn't been changed since '89, contains a number of transcription errors and inaccuracies."<br>reference_to_cite: "Content, A., Mousty, P., & Radeau, M. (1990). Brulex. Une base de données lexicales informatisée pour le français écrit et parlé [Brulex, A lexical database for written and spoken French]. <i>L'Année Psychologique, 90</i>, 551-566."<br>url_information: "ULB server ( ftp://ftp.ulb.ac.be/pub/packages/psyling/ )"
Recent neuroimaging studies have revealed the involvement of the prefrontal cortex in the process of making an aesthetic judgment. One study reported that transcranial direct current stimulation (tDCS) of the left dorsolateral prefrontal cortex (lDLPFC) led to the significant enhancement of subjects' aesthetic judgment of visual images. We tested whether subjects' aesthetic judgments of visual and sensuous stimuli can be modulated by applying tDCS on the lDLPFC with the same parameters (2 mA, 20 min) used in a previous study. Sixteen subjects underwent the stimulation and sham conditions. We also measured subjects' feelings of pleasure, which has been reported to correlate with the feeling of beauty. We used six images from the International Affective Picture System (IAPS) with high valence ratings, and six IAPS images with middle valence ratings for visual stimuli; candies with six different kinds of flavors for gustatory stimuli, and stuffed animals with different kinds of textures for tactile stimuli (Brielmann & Pelli, 2017). Subjects rated their feelings of pleasure (1–10) during 30 s of stimulus exposure and in the following 60 s. At the end, they were asked to rate the beauty of the stimuli (0–4). The results showed that only the beauty of visual stimuli tended to receive lower ratings after stimulation. Bayes factor, an indicator of which among two competing models is a better data predictor, supported the decrease in beauty ratings through tDCS with high-valence and middle-valence IAPS images 1.4 and 1.9 times, respectively, more than the model that predicted that tDCS had no effect on rating. The ineffectiveness of tDCS on pleasure was 2.6 times higher than the model that tDCS had an effect on pleasure. Overall, our study suggests that effect of tDCS may be limited to certain modalities and that judgments of feelings of beauty and pleasure might involve independent processes. Meeting abstract presented at VSS 2018
Studies of human social perception become more persuasive when the behavior of raters can be separated from the variability of the stimuli they are rating. We prototype such a rigorous analysis for a set of five social ratings of faces varying by body fat percentage (BFP). 274 raters of both sexes in three age groups (adolescent, young adult, senior) rated five morphs of the same averaged facial image warped to the positions of 72 landmarks and semilandmarks predicted by linear regression on BFP at five different levels (the average, ±2 SD, ±5 SD). Each subject rated all five morphs for maturity, dominance, masculinity, attractiveness, and health. The patterns of dependence of ratings on the BFP calibration differ for the different ratings, but not substantially across the six groups of raters. This has implications for theories of social perception, specifically, the relevance of individual rater scale anchoring. The method is also highly relevant for other studies on how biological facial variation affects ratings.
Sentiment analysis determines the polarities and strength of the sentiment‐bearing expressions, and it has been an important and attractive research area. In the past decade, resources and tools have been developed for sentiment analysis in order to provide subsequent vital applications, such as product reviews, reputation management, call center robots, automatic public survey, etc. However, most of these resources are for the English language. Being the key to the understanding of business and government issues, sentiment analysis resources and tools are required for other major languages, e.g., Chinese.To overcome this obstacle, we introduce CSentiPackage, where resources for retrieving sentiment from texts in the Chinese language, are provided. The related sentiment analysis technologies and datasets are described to give the readers the opportunities to use resources and tools to process Chinese sentiment texts from the very basic to the advanced, i.e., applying sentiment dictionaries, obtaining sentiment scores, and analyzing stance of social media posts using the deep learning model. The introduced resources and tools in this paper include NTUSD, ANTUSD, the Chinese Morphological Dataset, the Chinese Opinion Treebank, CopeOpi, and UTCNN. These resources are all available at http://academiasinicanlplab.github.io/ and they are free for the research purpose.
The way information spreads through society has changed significantly over the past decade with the advent of online social networking. \nTwitter, one of the most widely used social networking websites, is known as the real-time, public microblogging network where news \nbreaks first. Most users love it for its iconic 140-character limitation and unfiltered feed that show them news and opinions in the \nform of tweets. Tweets are usually multilingual in nature and of varying quality. However, machine translation (MT) of twitter data \nis a challenging task especially due to the following two reasons: (i) tweets are informal in nature (i.e., violates linguistic norms), and \n(ii) parallel resource for twitter data is scarcely available on the Internet. In this paper, we develop FooTweets, a first parallel corpus of \ntweets for English–German language pair. We extract 4, 000 English tweets from the FIFA 2014 world cup and manually translate them \ninto German with a special focus on the informal nature of the tweets. In addition to this, we also annotate sentiment scores between 0 \nand 1 to all the tweets depending upon the degree of sentiment associated with them. This data has recently been used to build sentiment \ntranslation engines and an extensive evaluation revealed that such a resource is very useful in machine translation of user generated \ncontent.
Semantic classification and annotation of satellite images are of great importance and require knowledge resources. The complexity of satellite scenes makes its classification and annotation hard tasks and we are still far from totally resolving the semantic gap problem. There are several knowledge resources such as semantic networks, taxonomies and ontologies. In this paper, we propose to enrich the SatelliteScene-Ontology using real hyperspectral scenes, the USGS spectral library and the WordNet lexical database. The resulting ontology would be published online for further exploitation by researchers.
Research in the field of text analysis will always be related to words, either the selection of words to be used or the position of the words in a sentence. Furthermore, a hypothesis that each language difference can cause different meanings, makes some researchers interested in doing research classifying words based on emotion or affective words. Research focuses on affective states as a continuous numerical value to the dimensions of valence and arousal. Sentiment analysis that is usually done with positive and negative category approaches, nowadays, the dimensional approach can provide more analysis of grained sentiments. On the other hand, the affective words dataset with valence and arousal rating are still very rare, especially for the Indonesian language. Therefore, this research does an affective lexicon dataset called Indonesian Valence and Arousal Words (IVAW) containing 1024 words by Self-Assessment Manikin (SAM) surveys. Furthermore, for the next study, we will also crawls status in twitter based on selected words from IVAW to get Indonesian Valence and Arousal Text (IVAT). To predict VA rating for obtaining the advance of annotation quality, experiment will be compared by brain signal using EEG tool.
Down-regulation of negative emotions has been shown to reliably inhibit the emotion-modulated startle reflex, but it remains unclear whether the timing of the startle probe influences the quantification of emotion regulation with this measure. Moreover, it is not known whether the degree of startle inhibition corresponds to the subjective attenuation of negative emotions. Therefore, the two main goals of the study were, first, to systematically analyze the effect of probe time on startle inhibition and, second, to explore the association between subjectively perceived down-regulation of arousal and valence and the degree of startle inhibition. We presented negative and neutral pictures to N = 47 participants. Pictures were paired with the instruction to reappraise or to maintain the emotions elicited by these pictures. Probes were delivered at three different times during a 12.5-s regulation phase, and the startle response was measured with electromyography. Valence and arousal ratings were assessed after each trial. Results revealed no significant impact of probe time on startle inhibition during reappraisal. Startle inhibition and perceived down-regulation of arousal were significantly and positively correlated, whereas perceived down-regulation of valence was not. The results provide important implications for future studies in terms of startle probe timing and shed light onto the interpretation of startle inhibition as an indicator of subjective attenuation of negative emotions. (PsycINFO Database Record (c) 2018 APA, all rights reserved).
We introduce TED-Multilingual Discourse Bank, a corpus of TED talks transcripts in 6 languages (English, German, Polish, EuropeanPortuguese, Russian and Turkish), where the ultimate aim is to provide a clearly described level of discourse structure and semanticsin multiple languages. The corpus is manually annotated following the goals and principles of PDTB, involving explicit and implicitdiscourse connectives, entity relations, alternative lexicalizations and no relations. In the corpus, we also aim to capture the character-istics of spoken language that exist in the transcripts and adapt the PDTB scheme according to our aims; for example, we introducehypophora. We spot other aspects of spoken discourse such as the discourse marker use of connectives to keep them distinct from theirdiscourse connective use. TED-MDB is, to the best of our knowledge, one of the few multilingual discourse treebanks and is hoped tobe a source of parallel data for contrastive linguistic analysis as well as language technology applications. We describe the corpus, theannotation procedure and provide preliminary corpus statistics.
We encounter metaphors every day, but only a few jump out on us and make us stumble. However, little effort has been devoted to investigating more novel metaphors in comparison to general metaphor detection efforts. We attribute this gap primarily to the lack of larger datasets that distinguish between conventionalized, i.e., very common, and novel metaphors. The goal of this paper is to alleviate this situation by introducing a crowdsourced novel metaphor annotation layer for an existing metaphor corpus. Further, we analyze our corpus and investigate correlations between novelty and features that are typically used in metaphor detection, such as concreteness ratings and more semantic features like the Potential for Metaphoricity. Finally, we present a baseline approach to assess novelty in metaphors based on our annotations.
International Journal of Exercise Science 11(5): 609-624, 2018. An aversion to the sensations of physical exertion can deter engagement in physical activity. This is due in part to an associative focus in which individuals are attending to uncomfortable interoceptive cues. The purpose of this study was to test the effect of mindfulness on affective valence, ratings of perceived exertion (RPE), and enjoyment during treadmill walking. Participants (N=23; Mage=19.26, SD = 1.14) were only included in the study if they engaged in no more than moderate levels of physical activity and reported low levels of intrinsic motivation. They completed three testing sessions including a habituation session to determine the grade needed to achieve 65% of heart rate reserve (HRR); a control condition in which they walked at 65% of HRR for 10 minutes and an experimental condition during which they listened to a mindfulness track that directed them to attend to the physical sensations of their body in a nonjudgmental manner during the 10-minute walk. ANOVA results showed that in the mindfulness condition, affective valence was significantly more positive (p =.02, np2 =.22), enjoyment and mindfulness of the body were higher (p <.001, np2 =.36 and.40, respectively), attentional focus was more associative (p <.001, np2 =.67) and RPE was minimally lower (p =.06, np2 =.15). Higher mindfulness of the body was moderately associated with higher enjoyment (p <.05, r =.44) in the mindfulness but not the control condition. Results suggest that mindfulness during exercise is associated with more positive affective responses.
In dimensional affect recognition, the machine learning methods, which are used to model and predict affect, are mostly classification and regression. However, the annotation in the dimensional affect space usually takes the form of a continuous real value which has an ordinal property. The aforementioned methods do not focus on taking advantage of this important information. Therefore, we propose an affective rating ranking framework for affect recognition based on face images in the valence and arousal dimensional space. Our approach can appropriately use the ordinal information among affective ratings which are generated by discretizing continuous annotations. Specifically, we first train a series of basic cost-sensitive binary classifiers, each of which uses all samples relabeled according to the comparison results between corresponding ratings and a given rank of a binary classifier. We obtain the final affective ratings by aggregating the outputs of binary classifiers. By comparing the experimental results with the baseline and deep learning based classification and regression methods on the benchmarking database of the AVEC 2015 Challenge and the selected subset of SEMAINE database, we find that our ordinal ranking method is effective in both arousal and valence dimensions.
We report on a pilot study involving emotion elicitation in virtual reality (VR) and assessment of emotional responses with a consumer-grade EEG device. The stimulation used HTC Vive VR system showing pictures from NAPS database within a specifically designed virtual environment. The stimulation consisted of two distinct sequences with 10 pictures of happiness and 10 pictures of fear. Each picture was contained in a separate virtual room that the participants traveled through along a preset path. The estimation employed EMOTIV EPOC+ 14-channel EEG headset and a custom-developed application. The software wirelessly received EEG signals from alpha, beta low, beta high, gamma and theta bands, time-stamped them and dynamically stored in a relational database for subsequent analysis. Our preliminary results show that statistically significant correlations between valence and arousal ratings of pictures and EEG bands are present but highly personalized. Simultaneous correct placement of VR and EEG headsets is demanding and precise localization of electrodes is difficult. In fact, if emotion estimation is not strictly necessary we recommend using devices with fewer electrodes. Nevertheless, we found the EEG to be effective. By acknowledging its limitations, and using the headset in the correct context, experiments involving emotions may be significantly amended.
The recognition of emotional facial expressions is a central aspect for an effective interpersonal communication. This study aims to investigate whether changes occur in emotion recognition ability and in the affective reactions (self-assessed by participants through valence and arousal ratings) associated with the viewing of basic facial expressions during preadolescence (n = 396, 206 girls, aged 11–14 years, Mage = 12.73, DS =.91). Our results confirmed that happiness is the best recognised emotion during preadolescence. However, a significant decrease in recognition accuracy across age emerged for fear expressions. Moreover, participants’ affective reactions elicited by the vision of happy facial expressions resulted to be the most pleasant and arousing compared to the other emotional expressions. On the contrary, the viewing of sadness was associated with the most negative affective reactions. Our results also revealed a developmental change in participants’ affective reactions to the stimuli. Implications are discussed by taking into account the role of emotion recognition as one of the main factors involved in emotional development.
The emotional valence of target information has been a centerpiece of recent false memory research, but in most experiments, it has been confounded with emotional arousal. We sought to clarify the results of such research by identifying a shared mathematical relation between valence and arousal ratings in commonly administered normed materials. That relation was then used to (a) decide whether arousal as well as valence influences false memory when they are confounded and to (b) determine whether semantic properties that are known to affect false memory covary with valence and arousal ratings. In Study 1, we identified a quadratic relation between valence and arousal ratings of words and pictures that has 2 key properties: Arousal increases more rapidly as function of negative valence than positive valence, and hence, a given level of negative valence is more arousing than the same level of positive valence. This quadratic function predicts that if arousal as well as valence affects false memory when they are confounded, false memory data must have certain fine-grained properties. In Study 2, those properties were absent from norming data for the Cornell-Cortland Emotional Word Lists, indicating that valence but not arousal affects false memory in those norms. In Study 3, we tested fuzzy-trace theory's explanation of that pattern: that valence ratings are positively related to semantic properties that are known to increase false memory, but arousal ratings are not. (PsycINFO Database Record (c) 2019 APA, all rights reserved).
Research on food experience is typically challenged by the way questions are worded. We therefore developed the EmojiGrid: a graphical (language-independent) intuitive self-report tool to measure food-related valence and arousal. In a first experiment participants rated the valence and the arousing quality of 60 food images, using either the EmojiGrid or two independent visual analog scales (VAS). The valence ratings obtained with both tools strongly agree. However, the arousal ratings only agree for pleasant food items, but not for unpleasant ones. Furthermore, the results obtained with the EmojiGrid show the typical universal U-shaped relation between the mean valence and arousal that is commonly observed for a wide range of (visual, auditory, tactile, olfactory) affective stimuli, while the VAS tool yields a positive linear association between valence and arousal. We hypothesized that this disagreement reflects a lack of proper understanding of the arousal concept in the VAS condition. In a second experiment we attempted to clarify the arousal concept by asking participants to rate the valence and intensity of the taste associated with the perceived food items. After this adjustment the VAS and EmojiGrid yielded similar valence and arousal ratings (both showing the universal U-shaped relation between the valence and arousal). A comparison with the results from the first experiment showed that VAS arousal ratings strongly depended on the actual wording used, while EmojiGrid ratings were not affected by the framing of the associated question. This suggests that the EmojiGrid is largely self-explaining and intuitive. To test this hypothesis, we performed a third experiment in which participants rated food images using the EmojiGrid without an associated question, and we compared the results to those of the first two experiments. The EmojiGrid ratings obtained in all three experiments closely agree. We conclude that the EmojiGrid appears to be a valid and intuitive affective self-report tool that does not rely on written instructions and that can efficiently be used to measure food-related emotions.
This paper (1) presents the first partially manually verified treebank for Dutch CHILDES corpora, the AnnCor CHILDES Treebank; (2) argues explicitly that it is useful to assign adult grammar syntactic structures to utterances of children who are still in the process of acquiring the language; (3) argues that human annotation and automatic checks on this annotation must go hand in hand; (4) argues that explicit annotation guidelines and conventions must be developed and adhered to and emphasises consistency of the annotations as an important desirable property for annotations. It also describes the tools used for annotation and automated checks on edited syntactic structures, as well as extensions to an existing treebank query application (GrETEL) and the multiple formats in which the resources will be made available
espanolEste articulo se centra en la lexicografia del ingles antiguo y el analisis de corpus. El objetivo es definir un procedimiento de lematizacion para un tipo de corpus del ingles antiguo anotado y parseado conocido como treebank. Este estudio se centra en dos cuestiones, concretamente en indicar donde se encuentran los datos con los que se puede lematizar el treebank del ingles antiguo; y que procedimiento debe adoptarse para enlazar la lematizacion disponible en las fuentes con el treebank. A partir de las bases de conocimiento del Proyecto Nerthus, se disena, pone en practica y evalua un procedimiento semiautomatico para dotar The York-Toronto-Helsinki Parsed Corpus of Old English Prose de etiquetas de lemas. EnglishThis article deals with Old English lexicography and corpus analysis. It aims at devising a lemmatisation procedure for a type of annotated and parsed corpus of Old English known as treebank. This study addresses two questions, namely where to find the data with which an Old English treebank can be lemmatised; and what procedure should be adopted to link the lemmatisation available from the sources to the treebank. On the grounds of the set of knowledge bases compiled by the Nerthus Project, a semi-automatic procedure for annotating The York-Toronto-Helsinki Parsed Corpus of Old English Prose with lemma tags is devised, illustrated and assessed.
This article presents the LIA treebank of transcribed spoken Norwegian dialects. It consists of dialect recordings made in the period between 1950--1990, which have been digitised, transcribed, and subsequently annotated with morphological and dependency-style syntactic analysis as part of the LIA (Language Infrastructure made Accessible) project at the University of Oslo. In this article, we describe the LIA material of dialect recordings and its transcription, transliteration and further morphosyntactic annotation. We focus in particular on the extension of the native NDT annotation scheme to spoken language phenomena, such as pauses and various types of disfluencies, and present the subsequent conversion of the treebank to the Universal Dependencies scheme. The treebank currently consists of 13,608 tokens, distributed over 1396 segments taken from three different dialects of spoken Norwegian. The LIA treebank annotation is an on-going effort and future releases will extend on the current data set.
This paper presents the Coptic Universal Dependency Treebank, the first dependency treebank within the Egyptian subfamily of the Afro-Asiatic languages. We discuss the composition of the corpus, challenges in adapting the UD annotation scheme to existing conventions for annotating Coptic, and evaluate inter-annotator agreement on UD annotation for the language. Some specific constructions are taken as a starting point for discussing several more general UD annotation guidelines, in particular for appositions, ambiguous passivization, incorporation and object-doubling.
In view of the differences between the annotations of micro and macro discourse rela-tionships, this paper describes the relevant experiments on the construction of the Macro Chinese Discourse Treebank (MCDTB), a higher-level Chinese discourse corpus. Fol-lowing RST (Rhetorical Structure Theory), we annotate the macro discourse information, including discourse structure, nuclearity and relationship, and the additional discourse information, including topic sentences, lead and abstract, to make the macro discourse annotation more objective and accurate. Finally, we annotated 720 articles with a Kappa value greater than 0.6. Preliminary experiments on this corpus verify the computability of MCDTB.
In geographic information science, semantic relatedness is important for Geographic Information Retrieval (GIR), Linked Geospatial Data, geoparsing, and geo-semantics. But computing the semantic similarity/relatedness of geographic terminology is still an urgent issue to tackle. The thesaurus is a ubiquitous and sophisticated knowledge representation tool existing in various domains. In this article, we combined the generic lexical database (WordNet or HowNet) with the Thesaurus for Geographic Science and proposed a thesaurus–lexical relatedness measure (TLRM) to compute the semantic relatedness of geographic terminology. This measure quantified the relationship between terminologies, interlinked the discrete term trees by using the generic lexical database, and realized the semantic relatedness computation of any two terminologies in the thesaurus. The TLRM was evaluated on a new relatedness baseline, namely, the Geo-Terminology Relatedness Dataset (GTRD) which was built by us, and the TLRM obtained a relatively high cognitive plausibility. Finally, we applied the TLRM on a geospatial data sharing portal to support data retrieval. The application results of the 30 most frequently used queries of the portal demonstrated that using TLRM could improve the recall of geospatial data retrieval in most situations and rank the retrieval results by the matching scores between the query of users and the geospatial dataset.
The arbitrary relation between sound and meaning is a fundamental assumption of modern linguistic theory. However, psycholinguistic literature also reports evidence for iconicity of phonological symbols. Here, we focus on phonological iconicity or sound-meaning mappings with regard to affective word content. Analyses of affective ratings for a large-scale database of German words suggest potential sublexical affective values for certain graphemes, as they occur more frequently in words of certain affective meaning. Using a letter-search task, we investigate how these systematic mappings between phonology and affective word content influence online language processing. Responses were generally shorter for high-arousal target graphemes-involving crucial interactions with affective word content. Iconic form-meaning mappings regarding affective content seem to influence both the organization of the vocabulary and the processing of language using phonological units of high perceptual salience to iconically encode threat or alert. (PsycINFO Database Record (c) 2018 APA, all rights reserved).
German is a language with complex morphological processes.Its long and often ambiguous word forms present a bottleneck problem in natural language processing.As a step towards morphological analyses of high quality, this paper introduces a morphological treebank for German.It is derived from the linguistic database CELEX which is a standard resource for German morphology.We build on its refurbished, modernized and partially revised version.The derivation of the morphological trees is not trivial, especially for such cases of conversions which are morpho-semantically opaque and merely of diachronic interest.We develop solutions and present exemplary analyses.The resulting database comprises about 40,000 morphological trees of a German base vocabulary whose format and grade of detail can be chosen according to the requirements of the applications.The Perl scripts for the generation of the treebank are publicly available on github.In our discussion, we show some future directions for morphological treebanks.In particular, we aim at the combination with other reliable lexical resources such as GermaNet.
Imageability and subjective frequency of the 500 rated nouns in the Croatian Lexical DatabaseProperties such as word class, length, phonological and morphological complexity, concreteness, frequency, age of acquisition, and imageability have to be controlled in research and clinical practice, since they strongly affect the speed and accuracy of language processing by monolinguals and bilinguals as well as by speakers with language disorders.The purpose of this paper is to present the online Croatian Lexical Database (Cro.Hrvatska leksi~ka baza [HLB], http://polin-hlb.erf.hr/) that contains different (psycho)linguistic word properties, and to use the HLB to provide the first analyses about (1) the relationship between frequency and imageability for the rated 500 nouns, and (2) the influence of raters' age, gender and education on their judgement.The results indicate a significant positive correlation between noun frequency and imageability, but no significant influence of the three non-linguistic rater factors on judgements about (psycho)linguistic property.
Dependency treebank is an important resource in any language. In this paper, we present our work on building BKTreebank, a dependency treebank for Vietnamese. Important points on designing POS tagset, dependency relations, and annotation guidelines are discussed. We describe experiments on POS tagging and dependency parsing on the treebank. Experimental results show that the treebank is a useful resource for Vietnamese language processing.
The purpose of this study is to determine how hiring managers perceive applicants who have U.S. military job experience (i.e., military veterans) and how these perceptions affect ratings of perceived job fit. The study was conducted using an experimental design in which participants reviewed a job d
One of the goals of the Russian language course in the primary school is the formation of the communicative literacy. The content of the course should be aimed at understanding the wealth of linguistic means by primary school children; the formation of the ability to detect a violation of linguistic norms and the inadequacy of the linguistic means used in the speech situation; the accumulation of the experience in choosing of linguistic means in accordance with the peculiarities of the speech situation; the creation of oral and written texts that meet the criteria of content, connectivity, compliance with the norms of the Russian literary language. The article considers the classification of exercises that contribute to the formation of communicative literacy. The author gives the examples of exercises where the student acts in different roles: the student is an observer of the speech situation and analyzes the adequacy of the choice of linguistic means; the student is a direct participant in the given speech situation and makes a choice of language facilities; the student is offered to create the speech situation himself, to independently construct an oral and written text.
Part-of-speech (POS) tagging is the foundation of many natural language processing applications. Rule-based POS tagging is a wellknown solution, which assigns tags to the words using a set of predefined rules. Many researchers favor statistical-based approaches over rule-based methods for better empirical accuracy. However, until now, the computational cost of rule-based POS tagging has made it difficult to study whether more complex rules or larger rulesets could lead to accuracy competitive with statistical approaches. In this paper, we leverage two hardware accelerators, the Automata Processor (AP) and Field Programmable Gate Arrays (FPGA), to accelerate rule-based POS tagging by converting rules to regular expressions and exploiting the highly-parallel regular-expressionmatching ability of these accelerators. We study the relationship between rule set size and accuracy, and observe that adding more rules only poses minimal overhead on the AP and FPGA. This allows a substantial increase in the number and complexity of rules, leading to accuracy improvement. Our experiments on Treebank and Brown corpora achieve up to 2,600X and 1,914X speedups on the AP and on the FPGA respectively over rule-based methods on the CPU in the rule-matching stage, up to 58× speedup over the Perceptron POS tagger on the CPU in total testing time, and up to 253× speedup over the LSTM tagger on the GPU in total testing time, while showing a competitive accuracy compared to neural-network and statistical solutions.
We explore whether it is possible to build lighter parsers, that are statistically equivalent to their corresponding standard version, for a wide set of languages showing different structures and morphologies. As testbed, we use the Universal Dependencies and transitionbased dependency parsers trained on feedforward networks. For these, most existing research assumes de facto standard embedded features and relies on pre-computation tricks to obtain speed-ups. We explore how these features and their size can be reduced and whether this translates into speed-ups with a negligible impact on accuracy. The experiments show that grand-daughter features can be removed for the majority of treebanks without a significant (negative or positive) LAS difference. They also show how the size of the embeddings can be notably reduced.
The quantity of music content is rapidly increasing and automated affective tagging of music video clips can enable the development of intelligent retrieval, music recommendation, automatic playlist generators, and music browsing interfaces tuned to the users' current desires, preferences, or affective states. To achieve this goal, the field of affective computing has emerged, in particular the development of so-called affective brain-computer interfaces, which measure the user's affective state directly from measured brain waves using non-invasive tools, such as electroencephalography (EEG). Typically, conventional features extracted from the EEG signal have been used, such as frequency subband powers and/or inter-hemispheric power asymmetry indices. More recently, the coupling between EEG and peripheral physiological signals, such as the galvanic skin response (GSR), have also been proposed. Here, we show the importance of EEG amplitude modulations and propose several new features that measure the amplitude-amplitude cross-frequency coupling per EEG electrode, as well as linear and non-linear connections between multiple electrode pairs. When tested on a publicly available dataset of music video clips tagged with subjective affective ratings, support vector classifiers trained on the proposed features were shown to outperform those trained on conventional benchmark EEG features by as much as 6, 20, 8, and 7% for arousal, valence, dominance and liking, respectively. Moreover, fusion of the proposed features with EEG-GSR coupling features showed to be particularly useful for arousal (feature-level fusion) and liking (decision-level fusion) prediction. Together, these findings show the importance of the proposed features to characterize human affective states during music clip watching.
The current study examined the effect of hearing loss and hearing aids in older adults on the experience of music emotion. Fifty-four participants aged 60 and above were recruited, including 18 with normal hearing (NH), 18 with hearing impairment (HI), and 18 with hearing impairment who use hearing aids (HA). Arousal ratings and skin conductance responses were obtained from participants across 24 extracts of film music that were previously validated as conveying one of four emotions: happy, sad, fearful, and tender. These four emotions represent a crossing of arousal (high and low) and valence (positive and negative). For each group, an “arousal range” was calculated as the average arousal ratings of happy and fearful excerpts, minus the average arousal ratings of tender and sad excerpts. A similar scheme was used for each group to determine a “skin conductance range”. For negatively-valenced music, groups did not differ with regard to arousal of skin conductance. For positively valenced music, NH and HA yielded a larger arousal range than did HI. A similar pattern emerged from the analysis of skin conductance range. Results will be discussed with regard to the acoustic cues to emotion in music and signal processing methods in hearing aids.
Musicians’ body movements are considered to be a source of communication and to contribute to the perceived musical impression of a performance. The aim of this study was to analyse ancillary body movements of clarinettists in order to identify characteristic motion types used in performance and to investigate their influence on the perception of the musical performance. The body movements of 22 clarinettists were recorded using 3D motion capture. A cluster analysis was performed on the variances of the angles in the shoulders, arms, knees, the back and the instrument hold, in order to examine commonalities in movement patterns. Four different motion types were classified: predominant knee motion (PkneeM), predominant arm motion (ParmM), no specific prominent motion pattern (NoSMP) and overall low motion performance (LowMP). In a perceptual experiment, in which all audio-visual recordings of performances were presented with the same audio track, 154 participants were asked to rate 8 performances (2 in each motion type) on 5 evaluation scales. The results showed that different motion types had a significant effect on perceptual ratings, with the highest affective ratings given to those who performed with predominant motion patterns in either the arms (ParmM) or in the knees (PkneeM), and the lowest ratings were given for performances with overall low motion (LowMP). In summary, the categorization of four motion types supports a systematic analysis of commonalities in motions made during clarinet performances across players. In addition, these different motion types have been shown to effectively influence perception of musical performance.