Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper presents the first release of LICO, a Lexicon for Italian COnnectives. LICO includes about 170 discourse connectives used in Italian, together with their orthographical variants, part of speech(es), semantic relation(s) (according to the Penn Discourse Treebank relation catalogue), and a number of usage examples.
Quantitative analysis of organismal form is an important component for almost every branch of biology. Although generally considered an easily-measurable structure, the quantification of gastropod shell form is still a challenge because many shells lack homologous structures and have a spiral form that is difficult to capture with linear measurements. In view of this, we adopt the idea of theoretical modelling of shell form, in which the shell form is the product of aperture ontogeny profiles in terms of aperture growth trajectory that is quantified as curvature and torsion, and of aperture form that is represented by size and shape. We develop a workflow for the analysis of shell forms based on the aperture ontogeny profile, starting from the procedure of data preparation (retopologising the shell model), via data acquisition (calculation of aperture growth trajectory, aperture form and ontogeny axis), and data presentation (qualitative comparison between shell forms) and ending with)
Стаття присвячена дослідженню англо-американських запозичень-термінів сучасної німецької мови, а також дослідженню особливостей їх функціонування в мові та проблем, які вони створюють при перекладі. Звертається увага на сучасні підходи до дослідження прагматичного потенціалу як самого політичного тексту так і використання запозичень англо-американського походження. (This article deals with the Anglo-american loan words, the terms in the modern german language and scientific researches of their functioning and the problems by their translate. Terminus is the lexical unit, it plays special functions. For analysis of termini are used semiotical/ terminological methods. All components of structure must be studied. The study of terms, the formation of which is attributed as extralinguistics factors and structural-linguistic norms assumes the duties of the structural-semantic analysis of these unit. To research the content structure of the term important all the elements of this scheme. You should start with a consideration of the meaning of the term, that is, the value of the lexical units serving in the term, if it has such a function. It can be argued that in this case the lexical unit has the nomìnativne value, it directly calls a special concept, which corresponds to the term. Complex terms form the main arsenal of the nominative means terminology elektrovimìrûval′noï technique. The model of complex terms shall be constructive function plays the ratio between turns.)
Since its emergence as an academic discipline in the early 1970s, feminist commentary and scholarship has prosecuted a critique of androcentric or sexist (gender exclusive) language, which has to some extent been successful. The struggle by women to occupy a positive linguistic space is continually being challenged by the endemic nature of masculine bias, which is realized through “indirect” or “subtle” sexism in the community. Seemingly innocuous words, like guy/guys, are frequently used to represent both men and women, reminiscent of the previous use of man/men as gender-inclusive common nouns. This raises the question of how to account for the persistence of such language use in spite of the fact that attention is regularly drawn to its problematic character. In this paper we approach the matter in a novel way, by appealing to work in the field of cognitive semantics, in particular the conceptual theory of metonymy. We propose that the relationship between the concepts of masculine and feminine as these are typically structured through language is indicative of a metonymy THE MASCULINE FOR THE FEMININE, in which the masculine “stands for” the feminine and in which lexical items are given as inclusive yet in effect refer to one (normative) gender. A corollary is that the feminine is subsumed (really or virtually) by the presence of the masculine and is made to disappear, and only reappears when she needs to be specified within the contextual frame.
Objective: This study examines reading aloud in patients with amyotrophic lateral sclerosis (ALS) and those with frontotemporal dementia (FTD) in order to determine whether differences in patterns of speaking and pausing exist between patients with primary motor vs. primary cognitive-linguistic deficits, and in contrast to healthy controls. Design: 136 participants were included in the study: 33 controls, 85 patients with ALS, and 18 patients with either the behavioural variant of FTD (FTD-BV) or progressive nonfluent aphasia (FTD-PNFA). Participants with ALS were further divided into 4 non-overlapping subgroups—mild, respiratory, bulbar (with oral-motor deficit) and bulbar-respiratory—based on the presence and severity of motor bulbar or respiratory signs. All participants read a passage aloud. Custom-made software was used to perform speech and pause analyses, and this provided measures of speaking and articulatory rates, duration of speech, and number and duration of pauses. Thes)
The article deals with two most important aspects of linguistic ecology – interlingual and translingual. The first aspect is connected with culture of speech, stylistics and rhetoric and includes analysis of violations of norms in speech – stylistic, lexical, grammatical – and their possible correction. The second aspect is studied in connection with the problems of adequacy of translation of fiction as a unity of “ecosystems” in contact of languages and cultures. Liguoecological approach allowed to study the role of language as an instrument of supporting community, functioning in certain situations of communication which are presented by pupils’ and students’ speech. Rhetorical, stylistic and aspects of culture of speech in the sphere of linguoecology have been examined from the point of view of the norm of any speech activity.
Life satisfaction refers to a somewhat stable cognitive assessment of one’s own life. Life satisfaction is an important component of subjective well being, the scientific term for happiness. The other component is affect: the balance between the presence of positive and negative emotions in daily life. While affect has been studied using social media datasets (particularly from Twitter), life satisfaction has received little to no attention. Here, we examine trends in posts about life satisfaction from a two-year sample of Twitter data. We apply a surveillance methodology to extract expressions of both satisfaction and dissatisfaction with life. A noteworthy result is that consistent with their definitions trends in life satisfaction posts are immune to external events (political, seasonal etc.) unlike affect trends reported by previous researchers. Comparing users we find differences between satisfied and dissatisfied users in several linguistic, psychosocial and other features. For )
This release adds the remaining parts of Sphrantzes' <em>Chronicles</em> along with a few annotation corrections to other texts.
The dictionary resources are very important for Natural Language Processing (NLP). Generating high quality dictionary resources is a crucial step for the success and effectiveness of NLP application. Linguistic information about lexical database is complex, large size and various (ie, phonological, morphological, syntactic, semantic and pragmatic). Among such lexical database entries, we find conjugated verbs. To this end, we present in this paper the open source mobile application of our conjugator that we developed in Java platform under the Android. This Conjugator allowed us to generate a lexicon of more than 18667 conjugated verbs. This lexicon will be used to generate textual words. The resultant lexicon can be used in various applications such as morphological analysis (lexical approach), text indexing, etc
In der vorliegenden Arbeit wurde die Wirksamkeit einer Expositionstherapie in virtueller Realitat bei Zahnbehandlungsphobikern untersucht. Uber eine Vorher- und Nachher-Analyse sollte herausgefunden werden, inwieweit die Angst vor phobischen Stimuli reduziert werden kann. Die Untersuchungen dieser Studie stutzten sich auf zwei empirische EEG-Studien von Kenntner-Mabiala & Pauli (2005, 2008), die evaluierten, dass Emotionen, die Schmerzwahrnehmung und die Toleranz der Schmerzschwelle modulieren konnen. Zudem konnte in einer EEG-Studie von Leutgeb et al. (2011) gezeigt werden, dass Zahnbehandlungsphobiker eine Erhohung der EKPs auf phobisches Stimulusmaterial aufwiesen. Die Frage nach dem Einfluss von emotionalen und phobischen Bildern auf die neuronale Verarbeitung sollte hier untersucht werden. Auserdem sollte herausgefunden werden welche Auswirkung emotionale und phobische Gerausche auf die Schmerzverarbeitung vor und nach der Therapie haben. Die Probanden wurden an drei aufeinanderfolgenden Terminen untersucht. Der erste Termin beinhaltete die Diagnostik zur Zahnbehandlungsphobie und den experimentellen Teil, der sich in drei Teile pro Termin gliederte. Der erste Teil enthielt die Aufzeichnung des EEG unter Schmerzreizapplikation im Kontext emotionaler Gerausche (neutral, negativ, positiv & zahn) und das Bewerten dieser Schmerzreize bezuglich der Intensitat und der Unangenehmheit des Schmerzes. Der zweite Teil enthielt Ratings zu Valenz und Arousal bezuglich dieser emotionalen Gerauschkategorien. Der dritte Teil enthielt die Aufzeichnung des EEG und das Rating zu Valenz und Arousal bezuglich emotionaler Bildkategorien (neutral, negativ, zahn). Am zweiten Termin folgte die Expositionstherapie unter psychologischer Betreuung. Der dritte Termin diente zur Erfolgsmessung und verlief wie Termin eins. Als Erfolgsmase der Therapie dienten Selbstbeurteilungsfragebogen, Valenz- und Arousal-Ratings des Stimulusmaterials, Schmerzratings und die durch das EEG aufgezeichneten visuell Ereigniskorrelierten- und Somatosensorisch-Evozierten-Potentialen. Die Ergebnisse zeigten, dass Gerausche mit unterschiedlichen emotionalen Kategorien zu eindeutig unterschiedlichen Valenz- und Arousalempfindungen bei Zahnbehandlungsphobikern fuhren. Die Studie konnte bestatigen, dass phobische Gerauschstimuli einen Einfluss auf die erhohte Erregung bei Zahnbehandlungsphobikern haben, die nach der Intervention als weniger furchterregend empfunden werden. Zudem konnte erwiesen werden, dass Personen mit Zahnbehandlungsphobie durch das Horen phobischer Zahnbehandlungsgerausche eine starkere Schmerzempfindung aufwiesen als durch positive, neutrale und negative Gerausche. Die Ergebnisse der Somatosensorisch-Evozierten-Potenziale (N150, P260) im Vergleich der Vorher und Nachher-Analyse zeigten tendenzielle Modulationen, die jedoch nicht signifikant waren. Im Vergleich zur Pra-Messung nahm die N150 Amplitude in der Post-Messung fur die schmerzhaften Stimuli wahrend der phobischen und negativen Gerausche ab. Auserdem wurden in dieser Studie parallel zum Gerauschparadigma weitere Sinnesmodalitaten mit phobie-relevanten Reizen anhand von Bildern getestet. Parallel zu den Ergebnissen der Studie von Leutgeb et al. (2011) fanden wir eine verstarkte elektrokortikale Verarbeitung im Late-Positive-Potential (LPP) auf phobische Bilder bei Zahnbehandlungsphobikern. Die Erwartung, dass die verstarkte elektrokortikale Verarbeitung des LPPs auf phobische Bilder bei Zahnbehandlungsphobikern durch Intervention reduziert werden kann, konnte nicht belegt werden. Rein deskriptiv gehen die Ergebnisse aber in diese Richtung. Auch das Verhalten anderte sich durch die Teilnahme an der Studie. Die Probanden gaben an, dass sich ihre Zahnbehandlungsangst nach der Expositionstherapie signifikant verringert hat. Das telefonische Follow-Up 6 Monate nach der Post-Messung zeigte, dass sich einige Probanden nach mehreren Jahren wieder in zahnarztliche Behandlung begeben haben. Insgesamt kann diese Studie zeigen, dass Zahnbehandlungsphobie durch psychologische Intervention reduziert werden kann und auch die Angst vor phobischem Stimulusmaterial durch eine wiederholte Reizkonfrontation abnimmt. Jedoch konnte auf elektrokortikaler Ebene keine Modulation der Schmerzempfindung uber emotionale Gerausche festgestellt werden.
The project aims to provide a semi-supervised approach to identify Multiword Expressions in a multilingual context consisting of English and most of the major Indian languages. Multiword expressions are a group of words which refers to some conventional or regional way of saying things. If they are literally translated from one language to another the expression will lose its inherent meaning. To automatically extract multiword expressions from a corpus, an extraction pipeline have been constructed which consist of a combination of rule based and statistical approaches. There are several types of multiword expressions which differ from each other widely by construction. We employ different methods to detect different types of multiword expressions. Given a POS tagged corpus in English or any Indian language the system initially applies some regular expression filters to narrow down the search space to certain patterns (like, reduplication, partial reduplication, compound nouns, compound verbs, conjunct verbs etc.). The word sequences matching the required pattern are subjected to a series of linguistic tests which include verb filtering, named entity filtering and hyphenation filtering test to exclude false positives. The candidates are then checked for semantic relationships among themselves (using Wordnet). In order to detect partial reduplication we make use of Wordnet as a lexical database as well as a tool for lemmatising. We detect complex predicates by investigating the features of the constituent words. Statistical methods are applied to detect collocations. Finally, lexicographers examine the list of automatically extracted candidates to validate whether they are true multiword expressions or not and add them to the multiword dictionary accordingly.
Kamusi has been developing a system to analyze texts on the source side and present users with sense-specified dictionary options. Similarly to spellcheck, the user selects the intended meaning. We then use a multilingual lexical database to bridge to matching vocabulary in other languages. When paired with Freeling, additional pre-processing is possible for several languages. Integration with MT via Moses and Apertium is planned, but not yet undertaken. MWEs treatment is important. An MWE is lexicalized in the Kamusi database and marked for separability, with a definition and translation equivalents (one or more words) in other languages. When the initial term of an MWE appears in the source text, Pre:D queries the database and scans the sentence for all MWEs that could follow. The user can select the relevant MWE rather than the component words. A user can submit a missing sense or MWE for inclusion in the lexicon. Named entities can also be identified from data sources or by users and rendered appropriately across languages. When users agree, we will also use sense-tagged sentences for machine learning. A prototype of the core system is already functional.
Parsing Arabic language is a difficult task given the specificities of the language and given the scarcity of linguistic resources. Linguistic resources such as grammars are very important to any natural language processing application. Unfortunately, the manual construction of these resources is laborious and time-consuming. The use of annotated corpora as a knowledge database might be a solution to a fast construction of a grammar for a given language. In this paper, we began by presenting an overview of our method to automatically induce a probabilistic context free grammar from an Arabic annotated corpus (The Penn Arabic TreeBank). Then we tested the obtained grammar in the parsing task and we expose the evaluation results. Finally we present our vision of a hybrid method for parsing Modern Standard Arabic (MSA) that we believe that it could enhance obtained results.
This article deals with the problem of development of the dialectological base of the Machine fund of the Bashkir language. Along with the written monuments, folklore material, dialect is one of the most important sources for the study of the historical development and formation of the literary language. The main task of dialectologists is not only the collection but also the storage of dialect materials of the Bashkir language. Taking into account the importance of translation of dialect materials in electronic format and opening of free access to a wide audience to the information provided, the staff of the Laboratory of Linguistics and Information Technologies of the Institute of History, Language and Literature of the Ufa Scientific Center established a dialectological base as a part of the Machine fund of the Bashkir language consisting of three separate databases — lexical database, the database of dialectological atlas and textological database. The lexical database includes information about the dialectal lexicon of the Bashkir language. The database was developed on the basis of dialectal dictionaries compiled and published by the staff of the Institute of History, Language and Literature. The volume of the database is more than 52 000 dialectal units. The database of dialectological atlas is developed on the basis of materials of Dialectological atlas of the Bashkir language. The database allows to select the types of linguistic phenomena (phonetic, morphological, syntactic, lexical), for each type specific isoglosses were identified. Isoglosses allocated 250 strong points of the republic and neighboring regions. The textological database represents illustrative materials collected during numerous expeditions by the staff of the Institute of History, Language and Literature. To date, the database contains more than 500 texts in all dialects of the Bashkir language. Introduced in dialectological base material forms the basis of synchronic and diachronic study of the language features of dialects and sub-dialects of the Bashkir language.
This paper describes our construction of named-entity recognition (NER) systems in two Western Iranian languages, Sorani Kurdish and Tajik, as a part of a pilot study of “Linguistic Rapid Response” to potential emergency humanitarian relief situations. In the absence of large annotated corpora, parallel corpora, treebanks, bilingual lexica, etc., we found the following to be effective: exploiting distributional regularities in monolingual data, projecting information across closely related languages, and utilizing human linguist judgments. We show promising results on both a four-month exercise in Sorani and a two-day exercise in Tajik, achieved with minimal annotation costs.
We present a dependency to constituent tree conversion technique that aims to improve constituent parsing accuracies by leveraging dependency treebanks available in a wide variety in many languages. The technique works in two steps. First, a partial constituent tree is derived from a dependency tree with a very simple deterministic algorithm that is both language and dependency type independent. Second, a complete high accuracy constituent tree is derived with a constraint-based parser, which uses the partial constituent tree as external constraints. Evaluated on Section 22 of the WSJ Treebank, the technique achieves the state-of-the-art conversion F-score 95.6. When applied to English Universal Dependency treebank and German CoNLL2006 treebank, the converted treebanks added to the human-annotated constituent parser training corpus improve parsing F-scores significantly for both languages.
In this paper a methodology for disambiguating the word senses of polysemous words using Lexical Categories present in WordNet is presented. WordNet is a commonly used English lexical database. The algorithm is applied to the data scraped from Wikipedia articles. The representative context used in the algorithm is extracted from the Wikipedia pages of the words belonging to the category. The lexical category of the given word is determined using the words in its neighbouring context. After finding the lexical category of the word, the correct sense is found using a modified version of Lesks Algorithm. The output word sense correspond to those available in WordNet.
Treebank is one of important resources in the natural language processing. Compared with the rich and mature Chinese corpus, Vietnamese Syntactic Analysis is much more difficult. This paper presents a new approach which uses Chinese-Vietnamese bilingual word alignment corpus to build Vietnamese Dependency Treebank. Firstly, the aligned word processing was made by Chinese-Vietnamese sentence alignment; Secondly, the dependency parsing was done with Chinese sentences. Finally, Vietnamese Dependency Parsing Treebank was generated by Chinese-Vietnamese Languages align relationship and Chinese Dependency Tree, At the same time, The Vietnamese phrase tree converted into dependency Treebank can significantly improve the accuracy of dependency analysis. Experimental results show that this approach can simplify the process of manual collection and annotation of Vietnamese Treebank, and it can save manpower and time to build the Vietnamese Treebank. Experimental results show that the accuracy of this method compared to machine learning methods has improved significantly.
Recent HCI research has looked at conveying emotions through non-visual modalities, such as vibrotactile and thermal feedback. However, emotion is primarily conveyed through visual signals, and so this research aims to support the design of emotional visual feedback. We adapt and extend the design of the "pulsing amoeba" [29], and measure the emotion conveyed through the abstract visual designs. It is a first step towards more holistic multimodal affective feedback combining visual, auditory and tactile stimuli. An online survey garnered valence and arousal ratings of 32 stimuli that varied in colour, contour, pulse size and pulse speed. The results support previous research but also provide new findings and highlight the effects of each individual visual parameter on perceived emotion. We present a mapping of all stimulus combinations onto the common two-dimensional valence-arousal model of emotion.
Emotional expressions are an essential element of human interactions. Recent work has increasingly recognized that emotional vocalizations can color and shape interactions between individuals. Here we present data on the psychometric properties of a recently developed database of authentic nonlinguistic emotional vocalizations from human adults and infants (the Oxford Vocal 'OxVoc' Sounds Database; Parsons, Young, Craske, Stein, & Kringelbach, 2014). In a large sample (n = 562), we demonstrate that adults can reliably categorize these sounds (as 'positive,' 'negative,' or 'sounds with no emotion'), and rate valence in these sounds consistently over time. In an extended sample (n = 945, including the initial n = 562), we also investigated a number of individual difference factors in relation to valence ratings of these vocalizations. Results demonstrated small but significant effects of (a) symptoms of depression and anxiety with more negative ratings of adult neutral vocalizations (R2 =.011 and R2 =.008, respectively) and (b) gender differences in perceived valence such that female listeners rated adult neutral vocalizations more positively and infant cry vocalizations more negatively than male listeners (R2 =.021, R2 =.010, respectively). Of note, we did not find evidence of negativity bias among other affective vocalizations or gender differences in perceived valence of adult laughter, adult cries, infant laughter, or infant neutral vocalizations. Together, these findings largely converge with factors previously shown to impact processing of emotional facial expressions, suggesting a modality-independent impact of depression, anxiety, and listener gender, particularly among vocalizations with more ambiguous valence. (PsycINFO Database Record
English grammars indicate a variety of relations holding between conjoined VPs. VPs conjoined by and evince such senses as Result, Temporal Sequence and Concession. Although all these senses are ones associated with discourse relations, conjoined VPs have not been fully included in discourse annotation. Because of the value of discourse-annotated corpora for developing approaches to automated sense recognition, we have added their annotation to the Penn Discourse TreeBank. This paper describes how tokens were identified; how the process of span and sense annotation was modified and extended in order to keep the annotation of intra-sentential multi-clausal structures consistent with the rest of the corpus; and what the resulting corpus looks like, in terms of token frequency and common sense patterns.
This study investigated whether age and/or differences in hearing sensitivity influence the perception of the emotion dimensions arousal (calm vs. aroused) and valence (positive vs. negative attitude) in conversational speech. To that end, this study specifically focused on the relationship between participants' ratings of short affective utterances and the utterances' acoustic parameters (pitch, intensity, and articulation rate) known to be associated with the emotion dimensions arousal and valence. Stimuli consisted of short utterances taken from a corpus of conversational speech. In two rating tasks, younger and older adults either rated arousal or valence using a 5-point scale. Mean intensity was found to be the main cue participants used in the arousal task (i.e., higher mean intensity cueing higher levels of arousal) while mean F 0 was the main cue in the valence task (i.e., higher mean F 0 being interpreted as more negative). Even though there were no overall age group differences in arousal or valence ratings, compared to younger adults, older adults responded less strongly to mean intensity differences cueing arousal and responded more strongly to differences in mean F 0 cueing valence. Individual hearing sensitivity among the older adults did not modify the use of mean intensity as an arousal cue. However, individual hearing sensitivity generally affected valence ratings and modified the use of mean F 0. We conclude that age differences in the interpretation of mean F 0 as a cue for valence are likely due to age-related hearing loss, whereas age differences in rating arousal do not seem to be driven by hearing sensitivity differences between age groups (as measured by pure-tone audiometry).
Animated characters are expected to fulfill a variety of social roles across different domains. To be successful and effective, these characters must display a wide range of personalities. Designers and animators create characters with appropriate personalities by using their intuition and artistic expertise. Our goal is to provide evidence-based principles for creating social characters. In this article, we describe the results of two experiments that show how exaggerated and damped facial motion magnitude influence impressions of cartoon and more realistic animated characters. In our first experiment, participants watched animated characters that varied in rendering style and facial motion magnitude. The participants then rated the different animated characters on extroversion, warmth, and competence, which are social traits that are relevant for characters used in entertainment, therapy, and education. We found that facial motion magnitude affected these social traits in cartoon and realistic characters differently. Facial motion magnitude affected ratings of cartoon characters’ extroversion and competence more than their warmth. In contrast, facial motion magnitude affected ratings of realistic characters’ extroversion but not their competence nor warmth. We ran a second experiment to extend the results of the first. In the second experiment, we added emotional valence as a variable. We also asked participants to rate the characters on more specific aspects of warmth, such as respectfulness, calmness, and attentiveness. Although the characters’ emotional valence did not affect ratings, we found that facial motion magnitude influenced ratings of the characters’ respectfulness and calmness but not attentiveness. These findings provide a basis for how animators can fine-tune facial motion to control perceptions of animated characters’ personalities.
Recently, several sets of standardized food pictures have been created, supplying both food images and their subjective evaluations. However, to date only the OLAF (Open Library of Affective Foods), a set of food images and ratings we developed in adolescents, has the specific purpose of studying emotions toward food. Moreover, some researchers have argued that food evaluations are not valid across individuals and groups, unless feelings toward food cues are compared with feelings toward intense experiences unrelated to food, that serve as benchmarks. Therefore the OLAF presented here, comprising a set of original food images and a group of standardized highly emotional pictures, is intended to provide valid between-group judgments in adults. Emotional images (erotica, mutilations, and neutrals from the International Affective Picture System/IAPS) additionally ensure that the affective ratings are consistent with emotion research. The OLAF depicts high-calorie sweet and savory foods and low-calorie fruits and vegetables, portraying foods within natural scenes matching the IAPS features. An adult sample evaluated both food and affective pictures in terms of pleasure, arousal, dominance, and food craving, following standardized affective rating procedures. The affective ratings for the emotional pictures corroborated previous findings, thus confirming the reliability of evaluations for the food images. Among the OLAF images, high-calorie sweet and savory foods elicited the greatest pleasure, although they elicited, as expected, less arousal than erotica. The observed patterns were consistent with research on emotions and confirmed the reliability of OLAF evaluations. The OLAF and affective pictures constitute a sound methodology to investigate emotions toward food within a wider motivational framework. The OLAF is freely accessible at digibug.ugr.es.
Self-assessment methods are broadly employed in emotion research for the collection of subjective affective ratings. The Self-Assessment Manikin (SAM), a pictorial scale developed in the eighties for the measurement of pleasure, arousal, and dominance, is still among the most popular self-reporting tools, despite having been conceived upon design principles which are today obsolete. By leveraging on state-of-the-art user interfaces and metacommunicative pictorial representations, we developed the Affective Slider (AS), a digital self-reporting tool composed of two slider controls for the quick assessment of pleasure and arousal. To empirically validate the AS, we conducted a systematic comparison between AS and SAM in a task involving the emotional assessment of a series of images taken from the International Affective Picture System (IAPS), a database composed of pictures representing a wide range of semantic categories often used as a benchmark in psychological studies. Our results show that the AS is equivalent to SAM in the self-assessment of pleasure and arousal, with two added advantages: the AS does not require written instructions and it can be easily reproduced in latest-generation digital devices, including smartphones and tablets. Moreover, we compared new and normative IAPS ratings and found a general drop in reported arousal of pictorial stimuli. Not only do our results demonstrate that legacy scales for the self-report of affect can be replaced with new measurement tools developed in accordance to modern design principles, but also that standardized sets of stimuli which are widely adopted in research on human emotion are not as effective as they were in the past due to a general desensitization towards highly arousing content.
We developed a new lexical database named as `PolyWordNet'. The PolyWordNet organizes multiple senses of a polysemy word in such a way that each sense of the polysemy word is linked with its related words by dividing these related words into verbs, nouns, adverbs and adjectives. Each related word in the PolyWordNet is linked only with a single sense of a polysemy word except for the case of bridging related word. This is because such related word will lead to the multiple senses of a polysemy word during the sense disambiguation process if the related word is liked to more than one sense of the same polysemy word introducing the ambiguity in ambiguity as in the case of contextual overlap count WSD approaches that use the Princeton WordNet for sense disambiguation. The PolyWordNet resolves this problem which is produced due to the common information collected from Princeton WordNet. The results obtained from the experiments show exceptionally high accuracy (96.11%) of our Word Sense Disambiguation algorithm that uses our lexical database PolyWordNet. This accuracy is significantly higher than that of the accuracy (58.33%) of the other contextual overlap count Word Sense Disambiguation method that used the Princeton WordNet for sense disambiguation.
In the Big Data era, the visualization of large data sets is becoming an increasingly relevant task due to the great impact that data have from a human perspective. Since the visualization is the closer phase to the users within the data life cycles phases, there is no doubt that an effective, efficient and impressive representation of the analyzed data may result as important as the analytic process itself. Starting from previous experiences in importing, querying and visualizing WordNet database within Neo4J and Cytoscape, this work aims at improving the WordNet Graph visualization by exploiting the features and concepts behind tag clouds. The objective of this study is twofold: firstly, we argue that the proposed visualization strategy is able to put order in the messy and dense structure of nodes and edges of large knowledge bases as WordNet, showing as much as possible information from this knowledge source and in a clearer way; secondly, we think that the tag cloud approach applied to the synonyms rings reinforces the human cognition in recognizing the different usages of words in natural languages like English. In this regard, we also propose a formal strategy in order to evaluate the information perception in the use of our methodology by means of a questionnaire asked to a group of users. Finally, we compare these results with those resulting from the adoption of well known representations of WordNet within existing GUIs.
How do we parse the languages for which no treebanks are available? This contribution addresses the cross-lingual viewpoint on statistical dependency parsing, in which we attempt to make use of resource-rich source language treebanks to build and adapt models for the under-resourced target languages. We outline the benefits, and indicate the drawbacks of the current major approaches. We emphasize synthetic treebanking: the automatic creation of target language treebanks by means of annotation projection and machine translation. We present competitive results in cross-lingual dependency parsing using a combination of various techniques that contribute to the overall success of the method. We further include a detailed discussion about the impact of part-of-speech label accuracy on parsing results that provide guidance in practical applications of cross-lingual methods for truly under-resourced languages.
The mere exposure effect refers to the phenomenon where previous exposures to stimuli increase subsequent affective preference for those stimuli. It has been indicated that with specific stimulus-category( i.e., paintings, matrices, and photographs of scene), repeated exposure has little or opposite effect on affective ratings. In this study, two experiments were conducted in order to explore the effect of stimulus-category on the mere exposure effects. Photographs of young woman’s(Experiment 1)and photographs of scene(Experiment 2)were used as stimuli. The experimental methods and affective rating score for the stimuli as novel stimuli were almost equated between two experiments. The results showed that the repeated exposure increased affective ratings not only for the face stimuli but also for the scene stimuli. More over, this could be seen across the type of scene(natural or artificial).
The aim of this article is to discuss the advances already made as well as the issues that have arisen in the process of lemmatization of Old English weak verbs on a lexical database. A list of lemmas of the second class weak verbs of Old English is compiled by using the latest version of the lexical database Nerthus, which incorporates the texts of the Dictionary of Old English Corpus. A number of issues are discussed, mainly related to queries and spelling. The conclusion insists of the ways in which the queries as defined so far should be refined.
ABSTRACT This research encompasses the construction of a multilingual lexical database for cross-lingual information retrieval in the Indonesian legal domain. Multilingual lexical database featuring lexically and legally grounded conceptual representation can fit the cross-lingual information retrieval. Lexical database use Ontology Web Language (OWL) representation language. This representation is useful to provide application developers a high-quality resource and to promote interoperability.
This paper presents the IULA Spanish LSP Treebank, an open-source treebank of over 40,000 sentences, developed in the framework of the European project METANET4U. The IULA Spanish LSP Treebank is the first technical corpus of Spanish annotated at surface syntactic level, following the dependency grammar theory. We present the method we used to create the resource and the linguistic annotations that the treebank provides, using examples and comparing with similar resources. We also provide the statistics of the treebank and the evaluation results.
HamleDT (HArmonized Multi-LanguagE Dependency Treebank) is a compilation of existing dependency treebanks (or dependency conversions of other treebanks), transformed so that they all conform to the same annotation style. This version uses Universal Dependencies as the common annotation style. Update (November 1017): for a current collection of harmonized dependency treebanks, we recommend using the Universal Dependencies (UD). All of the corpora that are distributed in HamleDT in full are also part of the UD project; only some corpora from the Patch group (where HamleDT provides only the harmonizing scripts but not the full corpus data) are available in HamleDT but not in UD.
Recent years have seen a serious increase in the number and the quality of available online linguistic databases. Yet the use of databases for research purposes is still in its infancy. I intend to show how, by using different databases, one can test scientific hypotheses and how different databases can be used for different purposes. The examples will be drawn from phonology, the domain where the most comprehensive datasets are to be found. Specifically, I will examine in some details the distribution of labial-velars in Africa, since this feature has been considered typical of the 'Macro-Sudan Belt' (Clements & Rialland 2008, Guldemann 2008), hence showing an areal distribution pattern instead of a genealogical one. The following online databases have been explored: WALS (World Atlas of Language Structures), PHOIBLE, LAPSyD (Lyon-Albuquerque Phonological Systems Databases, Version 1.0.). All of them are freely available and provide maps for selected features. First I will show how these databases differ from each other: scope and quantity of data, their quality (i.e. reliability), presence vs absence of explicit curation, query interface, etc. Then the specific case of labial-velars will be explored through these databases. The geographical distribution of labial-velars is known to be restricted to an area that has been labelled ‘Macro-Sudan Belt’ (Guldemann 2008). Languages that have one or more labial-velar consonant(s) in their inventory are all situated within this area, and no language outside the area seem to exhibit any labial-velar consonant (though there probably are a handful of exceptions in the Pacific region). The use of various online databases to assess this claim yields no big surprise, if one considers the general pattern only. In the details, though, lie a few interesting things, not every one of which are captured by online databases. The most important is the local prevalence of labial-velars, i.e. their status in each language. If the actual distribution of labial-velars is the result of contact and diffusion, one would expect that their status be more marginal at the edges of the domain. The goals of this talk will therefore be: i) to present arguments to verify (or not) the above prediction; to show that even when the use of databases is not in itself enough to answer some scientific questions, they can be most helpful in suggesting directions that could have been very difficult to take without these tools.
We propose a linguistically driven approach to represent discourse relations in Chinese text as sequences. We observe that certain surface characteristics of Chinese texts, such as the order of clauses, are overt markers of discourse structures, yet existing annotation proposals adapted from formalism constructed for English do not fully incorporate these characteristics. We present an annotated resource consisting of 325 articles in the Chinese Treebank. In addition, using this annotation, we introduce a discourse chunker based on a cascade of classifiers and report 70% top-level discourse sense accuracy.
The present study investigates how sequential coherence in sentence pairs (events in sequence vs. unrelated events) affects the perceived ability to form a mental image of the sentences for both auditory and visual presentations. In addition, we investigated how the ease of event imagery affected online comprehension (word reading times) in the case of sequentially coherent and incoherent sentence pairs. Two groups of comprehenders were identified based on their self-reported ability to form vivid mental images of described events. Imageability ratings were higher and faster for pairs of sentences that described events in coherent sequences rather than non-sequential events, especially for high imagers. Furthermore, reading times on individual words suggested different comprehension patterns with respect to sequence coherence for the two groups of imagers, with high imagers activating richer mental images earlier than low imagers. The present results offer a novel link between research on imagery and discourse coherence, with specific contributions to our understanding of comprehension patterns for high and low imagers.
BACKGROUND: As lip augmentation becomes more popular, validated measures of lip fullness for quantification of outcomes are needed. OBJECTIVE: Develop a scale for rating lip fullness and establish its reliability and sensitivity for assessing clinically meaningful differences. METHODS: The initial Allergan Lip Fullness Scale (iLFS; a four-point photographic scale with verbal descriptions) was validated by eight physicians rating 55 live subjects during two rounds, conducted on one day. In addition, subjects performed self-evaluations. The revised Allergan Lip Fullness Scale (LFS), a five-point scale with a broader range of lip presentations, was validated by 21 clinicians in two online image rating sessions, ≥14 days apart, in which they used the LFS to rate overall, upper, and lower lip fullness of 144 3-dimensional (3D) images. Physician inter- and intra-rater agreement, subject intra-rater agreement (iLFS), and subject-physician agreement (iLFS) were evaluated. Additionally, during online rating session 1, raters ranked 38 pairs of 3D images, taken before and after lip augmentation, as "clinically different" or "not clinically different." The median LFS score difference for clinically different pairs was calculated to determine the clinically meaningful difference. RESULTS: Clinician inter- and intra-rater agreement for the iLFS and LFS was substantial to almost perfect. Subject self-assessments (iLFS) had substantial intra-rater reliability and a high level of agreement with physician assessments. Median LFS score differences for overall, upper, and lower lip fullness were 1 (mean: 0.63-0.69) for "clinically different" and 0 (mean: 0.28-0.36) for "not clinically different" image pairs; thus, clinical significance of a 1-point difference in LFS score was established. CONCLUSIONS: The LFS is a reliable instrument for physician classification of lip fullness. A 1-point score difference can detect clinically meaningful differences in lip fullness.
This thesis presents open source resources in the form of annotated corpora and modules for automatic morphosyntactic processing and analysis of Persian texts. More specifically, the resources consist of an improved part-of-speech tagged corpus and a dependency treebank, as well as tools for text normalization, sentence segmentation, tokenization, part-of-speech tagging, and dependency parsing for Persian. In developing these resources and tools, two key requirements are observed: compatibility and reuse. The compatibility requirement encompasses two parts. First, the tools in the pipeline should be compatible with each other in such a way that the output of one tool is compatible with the input requirements of the next. Second, the tools should be compatible with the annotated corpora and deliver the same analysis that is found in these. The reuse requirement means that all the components in the pipeline are developed by reusing resources, standard methods, and open source state-of-the-art tools. This is necessary to make the project feasible. Given these requirements, the thesis investigates two main research questions. The first is how can we develop morphologically and syntactically annotated corpora and tools while satisfying the requirements of compatibility and reuse? The approach taken is to accept the tokenization variations in the corpora to achieve robustness. The tokenization variations in Persian texts are related to the orthographic variations of writing fixed expressions, as well as various types of affixes and clitics. Since these variations are inherent properties of Persian texts, it is important that the tools in the pipeline can handle them. Therefore, they should not be trained on idealized data. The second question concerns how accurately we can perform morphological and syntactic analysis for Persian by adapting and applying existing tools to the annotated corpora. The experimental evaluation of the tools shows that the sentence segmenter and tokenizer achieve an F-score close to 100%, the tagger has an accuracy of nearly 97.5%, and the parser achieves a best labeled accuracy of over 82% (with unlabeled accuracy close to 87%).
Compound words are cross-linguistic morphological phenomena that occur in all languages. Compound words are widely accepted to be stored in the lexicon but their constituents need to be accessed during both language learning and production processes. In this study, the use of corpora was investigated for how to differentiate single-stem words from single-word compounds and then how to segment compound words when no phonological information is available. Stems and morphs discovered in manual segmentations of the METU-Sabanci Turkish Treebank and the CHILDES were employed in the compound word recognition task and the results were compared. The METU Turkish Corpus (with about 2 million words) and a webcorpus (with about 490 million of Turkish words) were utilized in the segmentation task. The results emphasize that the lexicon can be morpheme-based; and lexical frequencies are effective heuristics in compound word recognition and segmentation.
In this paper, we introduce an ongoing project for the development of a parallel treebank for Italian, English and French. The treebank is annotated in a dependency format, namely the one designed in the Turin University Treebank (TUT), hence the choice to call such new resource Par(allel)TUT. The project aims at creating a resource which can be useful in particular for translation research. Therefore, beyond constantly enriching the treebank with new and heterogeneous data, so as to build a dynamic and balanced multilingual treebank, the current stage of the project is devoted to the design of a tool for the alignment of data, which takes into account syntactic knowledge as annotated in this kind of resource. The paper focuses in particular on the study of translational divergences and their implications for the development of the alignment tool. The paper provides an overview of the treebank, with its current content and the peculiarities of the annotation format, the description of the classes of translational divergences which could be encountered in the treebank, together with a proposal for their alignment.
Despite extensive research on the neural basis of empathic responses for pain and disgust, there is limited data about the brain regions that underpin affective response to other people's emotional facial expressions. Here, we addressed this question using event-related functional magnetic resonance imaging to assess neural responses to emotional faces, combined with online ratings of subjective state. When instructed to rate their own affective response to others' faces, participants recruited anterior insula, dorsal anterior cingulate, inferior frontal gyrus, and amygdala, regions consistently implicated in studies investigating empathy for disgust and pain, as well as emotional saliency. Importantly, responses in anterior insula and amygdala were modulated by trial-by-trial variations in subjective affective responses to the emotional facial stimuli. Furthermore, overall task-elicited activations in these regions were negatively associated with psychopathic personality traits, which are characterized by low affective empathy. Our findings suggest that anterior insula and amygdala play important roles in the generation of affective internal states in response to others' emotional cues and that attenuated function in these regions may underlie reduced empathy in individuals with high levels of psychopathic traits.