Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Bilingual Base Noun Phrase (BaseNP) extraction is one of the key tasks of Natural Language Processing (NLP). This task is more challenging for the pair of English-Vietnamese due to the lack of available Vietnamese language resources such as treebanks, part-of-speech taggers, and parsers. In this paper, we propose a combination model that uses language characteristics based on statistics and the projection method to extract BaseNP correspondences from a bilingual corpus. The language characteristics used in this model include the word segmentation, word order and word classification [1]. Our model overcomes not only the lack of resources of Vietnamese, but also improves the performance of miss-alignment, null-alignment, overlap and conflict projection of the existing methods. The proposed model can be easily applied to other language pairs. Experiment on 66,646 pairs of sentences in the English-Vietnamese bilingual corpus shows that our proposed model is very satisfactory.
The corpus for training a parser consists of sentences of heterogeneous grammar usages. Previous parser domain adaptation work has concentrated on adaptation to the shifts in vocabulary rather than grammar usage. In this paper, we focus on exploiting the diversity of training date separately and then accumulates their advantages. We propose an approach that grammar is biased toward relevant syntactic style, and the complementary grammar usage are combined for inference. Multiple grammars with partly complementary points of strength are induced individually. They capture complementary data representation, and we accumulates their advantages in a joint model to assemble the complementary depicting powers. Despite its compatibility with many other methods, out product model achieves 85.20% F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> score on Penn Chinese Treebank, higher than previous systems.
Talking about emotion and sharing emotional experiences is a key component of human interaction. Specifically, individuals often consider the reactions of other people when evaluating the meaning and impact of an emotional stimulus. It has not yet been investigated, however, how emotional arousal ratings and physiological responses elicited by affective stimuli are influenced by the rating of an interaction partner. In the present study, pairs of participants were asked to rate and communicate the degree of their emotional arousal while viewing affective pictures. Strikingly, participants adjusted their arousal ratings to match up with their interaction partner. In anticipation of the affective picture, the interaction partner's arousal ratings correlated positively with activity in anterior insula and prefrontal cortex. During picture presentation, social influence was reflected in the ventral striatum, that is, activity in the ventral striatum correlated negatively with the interaction partner's ratings. Results of the study show that emotional alignment through the influence of another person's communicated experience has to be considered as a complex phenomenon integrating different components including emotion anticipation and conformity.
The Digital Revolution has significantly impacted lexicography on several levels. First, freedom from the traditional paper format has removed constraints on size and format, paving the way for the construction of ever larger lexical databases with multi-faceted, flexible and rich representations of word meaning and use that have been unfeasible for print dictionaries. Second, access to electronic text corpora provide a solid empirical base and allow the lexicographer to craft entries that reflect actual speaker usage, variations across genres and the dynamics of continuous language change. Corpora have moreover opened the possibility to explore and statistically measure on the distributional properties of word and encode their syntagmatic properties. Third, different resources (lexicons, annotated corpora, language-independent, formal ontologies, syntactic and frame-based resources, Wikipedia) can be interlinked and harmonized. Fourth, electronic dictionaries can be continuously updated by a large community of both experts and volunteers, independently of official releases of new editions. All these make it possible to test, on a large-scale and crosslingually, the viability of different theories of lexical meaning and the structure of the lexicon. We discuss these aspects of digital lexicography with a particular emphasis on WordNet.
We present a novel toolkit that implements the long short-term memory (LSTM) neural network concept for language modeling. The main goal is to provide a software which is easy to use, and which allows fast training of standard recurrent and LSTM neural network language models. The toolkit obtains state-of-the-art performance on the standard Treebank corpus. To reduce the training time, BLAS and related libraries are supported, and it is possible to evaluate multiple word sequences in parallel. In addition, arbitrary word classes can be used to speed up the computation in case of large vocabulary sizes. Finally, the software allows easy integration with SRILM, and it supports direct decoding and rescoring of HTK lattices. The toolkit is available for download under an open source license.
This article is concerned with the data structures, properties of query languages, and visualization facilities required for the generic representation of richly annotated, heterogeneous linguistic corpora. We propose that above and beyond a general graph-based data model, which is becoming increasingly popular in many complex annotation formats, a well-defined concept of multiple, potentially conflicting segmentation layers must be introduced to deal with different sources and applications of corpus data flexibly. We also propose a generic solution for specialized corpus visualizations in a Web interface using annotation-triggered style sheets, which leverage the power of modern browsers and CSS for multiple and highly customizable views of primary data. We offer an implementation and evaluation of our architecture in ANNIS3, an open-source browser-based architecture for corpus search and visualization. We present three case studies to test the coverage of the system, encompassing core linguistic and digital humanities use-cases including richly annotated newspaper treebanks, multilingual diplomatic and normalized manuscript materials edited in TEI, and analysis of multimodal recordings of spoken language.
In the last decades, food pictures have been repeatedly employed to investigate the emotional impact of food on healthy participants as well as individuals who suffer from eating disorders and obesity. However, despite their widespread use, food pictures are typically selected according to each researcher's personal criteria, which make it difficult to reliably select food images and to compare results across different studies and laboratories. Therefore, to study affective reactions to food, it becomes pivotal to identify the emotional impact of specific food images based on wider samples of individuals. In the present paper we introduce the Open Library of Affective Foods (OLAF), which is a set of original food pictures created to reliably select food pictures based on the emotions they prompt, as indicated by affective ratings of valence, arousal, and dominance and by an additional food craving scale. OLAF images were designed to allow simultaneous use with affective images from the International Affective Picture System (IAPS), which is a well-known instrument to investigate emotional reactions in the laboratory. The ultimate goal of the OLAF is to contribute to understanding how food is emotionally processed in healthy individuals and in patients who suffer from eating and weight-related disorders. The present normative data, which was based on a large sample of an adolescent population, indicate that when viewing affective non-food IAPS images, valence, arousal, and dominance ratings were in line with expected patterns based on previous emotion research. Moreover, when viewing food pictures, affective and food craving ratings were consistent with research on food cue processing. As a whole, the data supported the methodological and theoretical reliability of the OLAF ratings, therefore providing researchers with a standardized tool to reliably investigate the emotional and motivational significance of food. The OLAF database is publicly available at zenodo.org.
We present a gold standard annotation of syntactic dependencies in the English Web Treebank corpus using the Stanford Dependencies formalism. This resource addresses the lack of a gold standard dependency treebank for English, as well as the limited availability of gold standard syntactic annotations for English informal text genres. We also present experiments on the use of this resource, both for training dependency parsers and for evaluating the quality of different versions of the Stanford Parser, which includes a converter tool to produce dependency annotation from constituency trees. We show that training a dependency parser on a mix of newswire and web data leads to better performance on that type of data without hurting performance on newswire text, and therefore gold standard annotations for non-canonical text can be a valuable resource for parsing. Furthermore, the systematic annotation effort has informed both the SD formalism and its implementation in the Stanford Parser’s dependency converter. In response to the challenges encountered by annotators in the EWT corpus, the formalism has been revised and extended, and the converter has been improved.
Kashmiri is a resource poor language with very less computational and language resources available for its text processing. As the main contribution of this paper, we present an initial version of the Kashmiri Dependency Treebank. The treebank consists of 1,000 sentences (17,462 tokens), annotated with part-of-speech (POS), chunk and dependency information. The treebank has been manually annotated using the Pān. inian Computational Grammar (PCG) formalism (Begum et al., 2008; Bharati et al., 2009). This version of Kashmiri treebank is an extension of its earlier version of 500 sentences (Bhat, 2012), a pilot experiment aimed at defining the annotation guidelines on a small subset of Kashmiri corpora. In this paper, we have refined the guidelines with some significant changes and have carried out inter-annotator agreement studies to ascertain its quality. We also present a dependency parsing pipeline, consisting of a tokenizer, a stemmer, a POS tagger, a chunker and an inter-chunk dependency parser. It, therefore, con-stitutes the first freely available, open source dependency parser of Kashmiri, setting the initial baseline for Kashmiri dependency parsing. Keywords:Treebanks, Computational Resources, Dependency Parsing 1.
In this paper, we report our preliminary efforts in building an English-Turkish parallel treebank corpus for statistical machine translation. In the corpus, we manually generated parallel trees for about 5,000 sentences from Penn Treebank. English sentences in our set have a maximum of 15 tokens, including punctuation. We constrained the translated trees to the reordering of the children and the replacement of the leaf nodes with appropriate glosses. We also report the tools that we built and used in our tree translation task.
Automatic prediction of emotions requires reliably annotated data which can be achieved using scoring or pairwise ranking. But can we predict an emotional score using a ranking-based annotation approach? In this paper, we propose to answer this question by describing a regression analysis to map crowdsourced rankings into affective scores in the induced valence-arousal emotional space. This process takes advantages of the Gaussian Processes for regression that can take into account the variance of the ratings and thus the subjectivity of emotions. Regression models successfully learn to fit input data and provide valid predictions. Two distinct experiments were realized using a small subset of the publicly available LIRIS-ACCEDE affective video database for which crowdsourced ranks, as well as affective ratings, are available for arousal and valence. It allows to enrich LIRIS-ACCEDE by providing absolute video ratings for the whole database in addition to video rankings that are already available.
The Norwegian Dependency Treebank is a new syntactic treebank for Norwegian Bokmål and Nynorsk with manual syntactic and morphological annotation, developed at the National Library of Norway in collaboration with the University of Oslo. It is the first publically available treebank for Norwegian. This paper presents the core principles behind the syntactic annotation and how these principles were employed in certain specific cases. We then present the selection of texts and distribution between genres, as well as the annotation process and an evaluation of the inter-annotator agreement. Finally, we present the first results of data-driven dependency parsing of Norwegian, contrasting four state-of-the-art dependency parsers trained on the treebank. The consistency and the parsability of this treebank is shown to be comparable to other large treebank initiatives.\n\nProceedings of the LREC 2014, Ninth International Conference on Language Resources and Evaluation, Reykjavik, Iceland. http://www.lrec-conf.org/proceedings/lrec2014/index.html.
Proof theory began in the 1920’s as a part of Hilbert’s program. That program aimed to secure the foundations of mathematics by modeling infinitary mathematics with formal axiomatic systems, and proving those systems consistent using restricted, “finitary” means. The program thus viewed mathematics as a system of reasoning with precise linguistic norms, governed by rules that can be described and studied in concrete terms. Such a viewpoint, today, has applications in mathematics, computer science, and the philosophy of mathematics.
On the basis of studying the lexicographic sources a definition of cynicism the term “linguocinism” is given in the article. The meaning of this term is specified by its correlation with similar to its meaning, but not synonymous to the terms “vulgarisms” and “obscenism”. It became clear that vulgarisms and obscenisms being violation of ethic-linguistic norm are not always linguocinisms. Linguistic means of expressing linguocinism and the role of context in their realization are revealed.
The main objective of the Rhapsodie project (ANR Rhapsodie 07 Corp-030-01) was to define rich, explicit, and reproducible schemes for the annotation of prosody and syntax in different genres (± spontaneous, ± planned, face-to-face interviews vs. broadcast, etc.), in order to study the prosody/syntax/discourse interface in spoken French, and their roles in the segmentation of speech into discourse units (Lacheret, Kahane, & Pietrandrea forthcoming). We here describe the deliverable, a syntactic and prosodic treebank of spoken French, composed of 57 short samples of spoken French (5 minutes long on average, amounting to 3 hours of speech and 33000 words), orthographically and phonetically transcribed. The transcriptions and the annotations are all aligned on the speech signal: phonemes, syllables, words, speakers, overlaps. This resource is freely available at www.projet-rhapsodie.fr. The sound samples (wav/mp3), the acoustic analysis (original F0 curve manually corrected and automatic stylized F0, pitch format), the orthographic transcriptions (txt), the microsyntactic annotations (tabular format), the macrosyntactic annotations (txt, tabular format), the prosodic annotations (xml, textgrid, tabular format), and the metadata (xml and html) can be freely downloaded under the terms of the Creative Commons licence Attribution - Noncommercial - Share Alike 3.0 France. The metadata are encoded in the IMDI-CMFI format and can be parsed on line.
The Penn Discourse Treebank (PDTB) was released to the public in 2008. It remains the largest manually annotated corpus of discourse relations to date. Its focus on discourse relations that are either lexically-grounded in explicit discourse connectives or associated with sentential adjacency has not only facilitated its use in language technology and psycholinguistics but also has spawned the annotation of comparable corpora in other languages and genres. Given this situation, this paper has four aims: (1) to provide a comprehensive introduction to the PDTB for those who are unfamiliar with it; (2) to correct some wrong (or perhaps inadvertent) assumptions about the PDTB and its annotation that may have weakened previous results or the performance of decision procedures induced from the data; (3) to explain variations seen in the annotation of comparable resources in other languages and genres, which should allow developers of future comparable resources to recognize whether the variations are relevant to them; and (4) to enumerate and explain relationships between PDTB annotation and complementary annotation of other linguistic phenomena. The paper draws on work done by ourselves and others since the corpus was released.
Although exposure therapy is an effective treatment for anxiety disorders, fear sometimes returns following successful therapy. The Rescorla–Wagner model predicts that presenting two fear-provoking stimuli simultaneously (compound extinction) will maximize learning during exposure and reduce the likelihood of relapse. Participants were presented with either single extinction trials only or single extinction trials followed by compound extinction trials. In addition, participants within each extinction group were randomized to caffeine or placebo ingestion prior to extinction to investigate the mechanism by which compound extinction may maximize learning (enhanced associative change or enhanced responding). Participants presented with compound trials demonstrated significantly less fear responding at spontaneous recovery compared with participants who received single extinction trials only. Ingestion of caffeine also provided some protection from spontaneous recovery (as measured by valence ratings). At the reinstatement test, only compound extinction trials predicted less fear responding; caffeine ingestion prior to extinction did not attenuate reinstatement effects.
This web service performs dependency parsing in Spanish using a Malt Parser instance.It parses plain texts introduced by the user and generates linguistically annotated Treebank instances based on a data-driven parsing model.The parsing model is induced from de dependency-annotated IULA Treebank (Marimon et al, 2012) using the language-independent MaltParser 2 system as a dependency model trainer (Nivre et al, 2007).This Treebank contains 589,542 tokens in 42,099.In order to achieve optimal performance, the training corpus was previously analyzed with MaltOptimizer 3 (Ballesteros and Nivre, 2012), a tool developed to set the best parameters for MaltParser. Inputs, outputs and formats InputsThe input to be parsed is a plain text encoded in UTF-8.It can be introduced directly as a text instance in the dialogue box, as a text file or as a URL.The input language available at this moment in the web service is Spanish (es).
To solve the problem of lower precision caused by traditional query expansion technology, a new query expansion technique based on semantic context was proposed. The semantic context is constructed by WordNet knowledge base and related feedback documents. Firstly, the query words senses are confirmed by disambiguation with WordNet lexical database. Secondly, the initial expansion words are obtained according to the WordNet semantic hierarchy structure. Finally, the weight of the expansion terms is determined according to the overall correlation of candidate expansion terms and all the query words. These words whose weight is higher than weight threshold will be chosen as the final query expansion word. The experimental results show that the proposed method obviously improves the retrieval precision while preserving higher recall.
Prior research suggests that repeatedly approaching or avoiding a certain stimulus changes the liking of this stimulus. We investigated whether these effects of approach and avoidance training occur also when participants do not perform these actions but are merely instructed about the stimulus-action contingencies. Stimulus evaluations were registered using both implicit (Implicit Association Test and evaluative priming) and explicit measures (valence ratings). Instruction-based approach-avoidance effects were observed for relatively neutral fictitious social groups (i.e., Niffites and Luupites), but not for clearly valenced well-known social groups (i.e., Blacks and Whites). We conclude that instructions to approach or avoid stimuli can provide sufficient bases for establishing both implicit and explicit evaluations of novel stimuli and discuss several possible reasons for why similar instruction-based approach-avoidance effects were not found for valenced well-known stimuli.
In this paper, we present our work of humor recognition on Twitter, which will facilitate affect and sentimental analysis in the social network. The central question of what makes a tweet (Twitter post) humorous drives us to design humor-related features, which are derived from influential humor theories, linguistic norms, and affective dimensions. Using machine learning techniques, we are able to recognize humorous tweets with high accuracy and F-measure. More importantly, we single out features that contribute to distinguishing non-humorous tweets from humorous tweets, and humorous tweets from other short humorous texts (non-tweets). This proves that humorous tweets possess discernible characteristics that are neither found in plain tweets nor in humorous non-tweets. We believe our novel findings will inform and inspire the burgeoning field of computational humor research in the social media.
We present a novel approach for induc-ing unsupervised dependency parsers for languages that have no labeled training data, but have translated text in a resource-rich language. We train probabilistic pars-ing models for resource-poor languages by transferring cross-lingual knowledge from resource-rich language with entropy reg-ularization. Our method can be used as a purely monolingual dependency parser, requiring no human translations for the test data, thus making it applicable to a wide range of resource-poor languages. We perform experiments on three Data sets — Version 1.0 and version 2.0 of Google Universal Dependency Treebanks and Treebanks from CoNLL shared-tasks, across ten languages. We obtain state-of-the art performance of all the three data sets when compared with previously studied unsupervised and projected pars-ing systems. 1
Dans le but de présenté notre projet de licence nous avons réalisé un moteur de recherche sémantique afin de trouver les documents pertinent pour l’utilisateur pour cella on a utilisé la base de donnée lexical WordNet Comme principe afin d’apporter les sacs de mots. L’implémentation et la conception sont faites à l’aide du langage java en utilisant l’IDE NetBeans6.8. Abstract In order to present our project of licence we conducted a semantic search engine to find the relevant documents for the user that’s why we used the lexical database WordNet as a principle to provide bags of words. The implementation and design are made with the java language using the IDE NetBeans6.8.
In Dutch V-final clauses the verbs tend to form a cluster which cannot be split up by nonverbal \nmaterial. However, Haeseryn et al. (1997) as well as other studies on the phenomenon list several \ncases in which the verb cluster may be interrupted by \ncluster creepers. \nThe most common examples are constructions with separable verb particles, but examples with nouns, adjectives, and adverbs are attested as well. \nSince the majority of the data in previous studies is collected by introspection and elicitation, \nit is interesting to compare those findings to corpus data. The corpus analysis is based on data \nfrom two Dutch treebanks (CGN and LASSY), which allow to take into account regional and/or \nstylistic variation. This is an important aspect for the analysis, since cluster creeping is reported \nto be a typical property of spoken and regional variants of Dutch. \nThe goal of this corpus-based investigation is on the one hand to provide insight in the frequency \nof the phenomenon, and on the other hand to classify the types of cluster creepers. Besides the \nlinguistic analysis, methodological issues regarding the extraction of the relevant data from the \ntreebanks will be addressed as well.
We develop an instance (token) based extension of the state of the art word (type) based part-ofspeech induction system introduced in (Yatbaz et al., 2012). Each word instance is represented by a feature vector that combines information from the target word and probable substitutes sampled from an n-gram model representing its context. Modeling ambiguity using an instance based model does not lead to significant gains in overall accuracy in part-of-speech tagging because most words in running text are used in their most frequent class (e.g. 93.69% in the Penn Treebank). However it is important to model ambiguity because most frequent words are ambiguous and not modeling them correctly may negatively affect upstream tasks. Our main contribution is to show that an instance based model can achieve significantly higher accuracy on ambiguous words at the cost of a slight degradation on unambiguous ones, maintaining a comparable overall accuracy. On the Penn Treebank, the overall many-to-one accuracy of the system is within 1% of the state-of-the-art (80%), while on highly ambiguous words it is up to 70% better. On multilingual experiments our results are significantly better than or comparable to the best published word or instance based systems on 15 out of 19 corpora in 15 languages. The vector representations for words used in our system are available for download for further experiments.
My study assessed the relationship between the colour of depicted clothing and ratings of emotional intensity in drawings of emotional scenarios. Participants (N = 42) viewed a set of drawings in one of seven colours, labeled the Actor and Cause, and rated the intensity of the emotions depicted on 11 emotional scales. Participants also listed colours they associated with specific emotions. Colour of clothing in drawings was not found to affect ratings of emotional intensity, but in drawings of sadness, embarrassment, and empathy the Actor obtained higher ratings than did the Cause. Colour-emotion associations obtained separately were, however, generally consistent with previous findings. COLOUR AND EMOTION 3 Colour and Emotional Intensity Emotion recognition is an evolutionary advantage that can inform behavioural decisions (Elfenbein, Foo, White, Tan & Aik, 2007) and help individuals to effectively navigate social situations (Van Kleef, 2010). It is a process that is largely automatic and unconscious, thus requiring limited cognitive resources (Tracy & Robins, 2008). However, researchers in the areas of social, cognitive and perceptual psychology are still attempting to determine the factors contributing to the automaticity of emotion perception.
Facial expressions are one of the most important ways of non-verbal communication for humans. To date, most research in this field has focused solely on emotional aspects, largely neglecting the communicative and conversational aspects of expressions. Furthermore, although there is evidence for some degree of cross-cultural universality among emotional expressions, much less is known about how facial expressions in general are perceived across cultures. Here, we investigate the structure of the complex space of both emotional and conversational expressions in a cross-cultural context. The two experiments reported here used matching video sequences of 27 expressions from both the KU (Korean) facial expression database and the MPI (German) facial expression database (each expression was shown by 6 actors, totaling 162 videos from each database). In the first experiment, four groups (each n=20) of native German and Korean participants were asked to group the sequences of the German or Korean databases into clusters based on similarity. This grouping data yielded four different confusion matrices. In the second experiment, another four groups of participants (each n=20) from both cultures were asked to rate each video according to 13 emotional and conversational attributes. This rating data yielded an averaged 13-dimensional vector for each sequence. For each of the four grouping/rating data-pairs, we then used kernel canonical correlation analysis (KCCA) to determine a two-dimensional embedding of expressions that best explained both grouping and rating data. Although other attributes contributed as well, the two dimensions recovered by KCCA showed maximal correlation with valence and arousal ratings – this was true regardless of participants' cultural backgrounds or of the database that was used. Our results show that evaluative dimensions for both German and Korean cultural contexts are highly similar, confirming that cultural universals exist even in this complex space of emotional and conversational facial expressions. Meeting abstract presented at VSS 2014
Research in Sentiment Analysis has shown rapid progress since late 90s. It is an important research area as analyzing user’s feedback is useful for business analysis, product comparison, counter intelligence, and poll prediction. Despite the rapid surge of Sentiment Analysis research, many unresolved research questions remain. One of the biggest concerns is the Semantic Gap, which involves translating machine understandable form to human understandable form. Though research has been carried out for machine to understand human language, it is still not capable to address the problem mentioned as human languages are diverse and complex. WordNet, for example, attempt to address this issue by incorporating large lexical database for English, with various functionalities to manipulate this database. Recently, WordNet provides multilingual support, which is very helpful to address the diverse human languages. In this paper, we propose a novel multilingual common ontology tool to analyze user’s feedback and opinion. Unlike other existing state of the art tools, our tool is capable of handling multi languages regardless of the webpage layout. Experimental results show that our tool is highly efficient in analyzing opinion from social networking sites.
Helping behavior as a prosocial action emerges early in childhood and is of interest for psychologists in a broad range of sub disciplines as well as for society. One necessary precondition for active helping is the ability to recognize that somebody needs help. The NeoHelp stimulus set used in this study was developed to enable the assessment and quantification of need-of-help recognition abilities. Previous research with the NeoHelp stimuli has shown that children of different ages are able to recognize their content. Specific effects of age and gender on need-of-help recognition have also been observed. How children subjectively experience these stimuli and thus how they rate depictions of need-of-help and no-need-of-help situations emotionally has not been assessed before. Here we report analyses of valence and arousal ratings for the complete NeoHelp stimulus set obtained from a diverse sample of 46 children. We employed the SAM-scales because they are an established rating instrument validated for diverse populations of adults. However, their use with children still needs further investigation. Thus, there were two main goals of the presented study: 1) Validating that the SAM arousal and valence scales may be used with young children below school age, and 2) investigating children's subjective emotional experience of need-of-help depictions. Our study demonstrates that the SAM scales, if properly explained, may be used reliably with children at and above five years of age. Ratings of younger and older children covered the whole range of the 5-point scales used. There was a linear relationship between arousal and valence ratings across all pictures: the higher the arousal ratings, the lower the valence ratings. Pictures showing a child in need-of-help were rated as lower in valence and higher in arousal than the corresponding no-need-of-help-stimuli regardless of children’s age or gender. With increasing age, arousal ratings for no-need-of-help depictions decreased, but arousal ratings for need-of-help depictions remained on the same higher level across ages. We thus provide first evidence that need-of-help depictions elicit differential subjective emotional responses in children on both, valence and arousal dimensions. This emotional component of need-of-help recognition has to be considered when assessing children's need-of-help recognition abilities.
Lexical Knowledge base such as WordNet has been used as a valuable tool for measuring semantic similarity in various Information Retrieval (IR) applications. It is a domain independent lexical database. Since, the quality of semantic relationship in WordNet has not upgraded appropriately for the current usage in the modern IR. Building the WordNet from scratch is not an easy task for keeping updated with current terminology and concepts. Therefore, this paper undergoes a different perspective that automatically updates an existing lexical ontology uses knowledge resources such as the Wikipedia and the Web search engine. This methodology has established the recently evolving relations and also aligns the existing relations between concepts based on its usage over time. It consists of three main phases such as candidate article generation, lexical relationship extraction and generalization and WordNet alignment. In candidate article generation, disambiguation mapping disambiguates ambiguous links between WordNet concepts and Wikipedia articles and returns a set of word-article pairings. Lexical relationship extraction phase includes two algorithms, Lexical Relationship Retrieval (LRR) algorithm discovers the set of lexical patterns exists between concepts and sequential pattern grouping algorithm generalizes lexical patterns and computes corresponding weights based on its frequencies. Furthermore, Sequential Minimal Optimization (SMO) selects the suitable good pattern using the optimal combination of weight of lexical patterns and page count based concurrence measures. WordNet alignment phase establishes a new relationship that is not available in WordNet and also aligns the existing patterns based on computed weight. Experimental results illustrate that the proposed approach better than existing mechanisms on benchmark datasets and achieves a correlation value of 0.87. Moreover, the extended WordNet returns high accuracy results in query expansion.
The present paper explored the relationship between emotional facial response and electromyographic modulation in children when they observe facial expression of emotions. Facial responsiveness (evaluated by arousal and valence ratings) and psychophysiological correlates (facial electromyography, EMG) were analyzed when children looked at six facial expressions of emotions (happiness, anger, fear, sadness, surprise and disgust). About EMG measure, corrugator and zygomatic muscle activity was monitored in response to different emotional types. ANOVAs showed differences for both EMG and facial response across the subjects, as a function of different emotions. Specifically, some emotions were well expressed by all the subjects (such as happiness, anger and fear) in terms of high arousal, whereas some others were less level arousal (such as sadness). Zygomatic activity was increased mainly for happiness, from one hand, corrugator activity was increased mainly for anger, fear and surprise, from the other hand. More generally, EMG and facial behavior were highly correlated each other, showing a "mirror" effect with respect of the observed faces.
WordNet is an electronic lexical database available on-line as a powerful resource to the researchers in the area of computational linguistics, text processing and other related areas. WordNet for Hindi language has already been developed by IIT, Bombay. The Indian languages WordNets are being created using expansion approach from Hindi WordNet under IndoWordNet project. In expansion approach, semantic relations are borrowed from the reference language, while the lexical relations need to be created for each language, as these relations are language dependent. This paper describes the process of creation of lexical relations like antonym, compounding, conjunction and gradation for IndoWordNet. A lexical creation tool has been presented in this paper with provision to create lexical relations in target language on the basis of relations created in Hindi WordNet and with another provision to create lexical relations in target language without referring to Hindi WordNet. It has been observed that lexical relations for target language can be created easily on the basis of relations created in Hindi WordNet for Hindi in-family languages, while for the languages that do not fall in the same family provision of creation of lexical relation without referring to Hindi WordNet can be used.
Introduction: The last decade has witnessed a growing interest in human social cognition. Social neuroscience has differentiated at least two functions and the respective neural networks that support successful interaction: a network that allows for affect sharing (empathizing) and underlies the understanding of others’ affective states and a network that enables the inference of thoughts, beliefs, and goals and underlies understanding others’ mental states (Theory of Mind (ToM)). Empathy paradigms have mainly focused on witnessing pain or suffering and have revealed increased activation in the anterior insula (AI) and the anterior midcingulate cortex (aMCC) [1]. ToM tasks typically involve inferring others’ thoughts or intentions and induce increased activation in the temporoparietal junction (TPJ), the temporal poles and the dorsomedial prefrontal cortex (dmPFC) [2]. Although these different routes of social cognition show differential time courses during development [3] and are differentially impaired in psychopathology [4], these routes usually interact in healthy brains. So far, fMRI paradigms are lacking, which can reliably dissociate mentalizing and empathizing routes within a single person. The aim of the present study was to develop and validate an fMRI task in which demands on the affective and the cognitive social cognition route are manipulated independently and that therefore allows the investigation of their interaction as well as differential impairment. Methods: We developed the EmpaToM, a 30 minute paradigm presenting participants with naturalistic video stimuli (~15 seconds) in which people recount auto-biographical episodes that are either emotionally negative (e.g. loss of a loved one) or neutral (control condition; e.g., commuting to work). Each video is followed by ratings of affect and of empathic concern. Subsequently, specific questions about the content of the video are probing either ToM (questions about the mental states of people in the video) or no ToM (control condition; factual reasoning). The task, hence, follows a 2 (negative emotional load versus no emotional load) x 2 (ToM requirements versus factual reasoning) factorial design. In a first validation study, 19 participants were tested in the EmpaToM and classical tasks measuring empathy (the Social Video Task (SoVT) [5]) and ToM [2] (3T Siemens Verio Scanner; TR = 2000; TE =; 37 slices (2 mm)). Results: Viewing negative emotional compared to neutral videos increased activation in a distinct network including the bilateral AI and aMCC. This empathy network overlapped with activity elicited by the SoVT [5] and coordinates derived from a meta-analysis on empathy for pain paradigms [1]. Affect ratings and emotional concern ratings after the emotionally negative videos were correlated with the IRI subscale ‘empathic concern’ and with SoVT concern ratings. Furthermore, activity in the aMCC and right AI was parametrically modulated by the subjective rating of negative emotion experienced after each video. Contrasting ToM- and factual reasoning during the questions activated the well-described network including bilateral TPJ, temporal poles and mPFC. Again, these clusters largely overlapped with a classical false-belief ToM-task [2]. A subsequent study with 200 participants replicated these patterns. Conclusions: The present results confirm that the EmpaTom is able to reliably identify and separate the networks underlying our abilities to empathize (affective route) and to mentalize (cognitive route) within individuals and one task. This makes the EmpaToM a suitable task for the efficient differentiation of empathy and ToM related networks and for investigating their interaction. The EmpaToM could be of tremendous use for both the comprehensive characterization of different types of deficits in psychopathology as well as for the measurement of differential training effects and interventions both on the level of behavior and brain functioning. References [1] Lamm, C., Decety, J., & Singer, T. (2011). ‘Meta-analytic evidence for common and distinct neural networks associated with directly experienced pain and empathy for pain’, NeuroImage, vol. 54, pp. 2492-2502. [2] Dodell-Feder, D., Koster-Hale, J., Bedny, M., Saxe, R. (2011). ‘fMRI item analysis in a theory of mind task’, Neuroimage, vol. 55, pp. 705-12. [3] Singer, T. (2006). ‘The neuronal basis and ontogeny of empathy and mind reading: Review of literature and implications for future research’, Neuroscience & Biobehavioral Reviews, vol. 30, pp. 855-863. [4] Bird, G., Silani, G., Brindley, R., White, S., Frith, U., & Singer, T. (2010). ‘Empathic brain responses in insula are modulated by levels of alexithymia but not autism’, vol., 133, pp. 1515-1525. [5] Klimecki, O.M., Leiberg, S., Lamm, C., & Singer, T. 2013). ‚Functional Neural Plasticity and Associated Changes in Positive Affect After Compassion Training’. Cerebral Cortex, vol. 23, pp. 1552-1561.
Signal detection in clinical trials relies on ratings reliability. We conducted a reliability analysis of site-independent rater scores derived from audio-digital recordings of site-based rater interviews of the structured Brief Psychiatric Rating Scale (BPRS) in a schizophrenia study. "Dual" ratings assessments were conducted as part of a quality assurance program in a 12-week, double-blind, parallel-group study of PF-02545920 compared to placebo in patients with sub-optimally controlled symptoms of schizophrenia (ClinicalTrials.gov identifier NCT01939548). Blinded, site-independent raters scored the recorded site-based BPRS interviews that were administered in relatively stable patients during two visits prior to the randomization visit. We analyzed the impact of BPRS interview length on "dual" scoring variance and discordance between trained and certified site-based raters and the paired scores of the independent raters. Mean total BPRS scores for 392 interviews conducted at the screen and stabilization visits were 50.4±7.2 (SD) for site-based raters and 49.2±7.2 for site-independent raters (t=2.34; p=0.025). "Dual" rated total BPRS scores were highly correlated (r=0.812). Mean BPRS interview length was 21:05±7:47min ranging from 7 to 59min. 89 interviews (23%) were conducted in less than 15min. These shorter interviews had significantly greater "dual" scoring variability (p=0.0016) and absolute discordance (p=0.0037) between site-based and site-independent raters than longer interviews. In-study ratings reliability cannot be guaranteed by pre-study rater certification. Our findings reveal marked variability of BPRS interview length and that shorter interviews are often incomplete yielding greater "dual" scoring discordance that may affect ratings precision.
ILSP Dependency Parser is a tool trained on the Greek Dependency Treebank, a resource which comprises data annotated at several linguistic levels. Training data at the level of syntax consisted of ~70 KWords annotated using a dependency-based syntactic scheme that includes 25 main relations.
We investigate the feasibility of aligning Chinese and English parse trees by examining cases of incompatibility between Chinese-English parallel parse trees. This work is done in the context of an annotation project where we construct a parallel treebank by doing word and phrase alignments simultaneously. We discuss the most common incompatibility patterns identified within VPs and NPs and show that most cases of incompatibility are caused by divergent syntactic annotation standards rather than inherent cross-linguistic differences in language itself. This suggests that in principle it is feasible to align the parallel parse trees with some modification of existing syntactic annotation guidelines. We believe this has implications for the use of parallel parse trees as an important resource for Machine Translation models.
Temperature and chemesthesis interact, but this interaction has not been fully examined for most irritants. The current experiments focus on oral pungency from carbonation. Previous work showed that cooling carbon dioxide (CO2) solutions to below tongue temperature enhanced rated bite. However, to the best of our knowledge, the effects of warming to above tongue temperature have not been examined. In Experiment 1, subjects sampled CO2 solutions at 4 nominal concentrations (0.0, 2.0, 2.8, and 4.0 v/v) × 5 temperatures (18.3, 24.5, 29.9, 34.5, and 39.6 (o)C). Subjects dipped their tongue tips into samples and rated bite. As in previous work, subjects rated cool solutions (25.0 (o)C and lower) as more intense. Warming solutions above tongue temperature (39.6 (o)C) did not affect ratings. Experiment 2 examined warmer temperatures (18.3, 33.9, 39.0, 44.9, and 48.2 ºC). Bite was enhanced only at 48.2 ºC, and a follow-up experiment suggested that enhancement was probably due to confusion between carbonation bite and mild heat pain. Experiment 3 examined the effect of menthol cooling by pretreating the tongue with menthol. Unlike physical cooling, menthol cooling had little or no effect on rated bite. The results are discussed in the context of candidate transduction mechanisms for carbonation sensation.
poster
Introduction Previous research has suggested that visual images are more easily generated, more vivid, and more memorable than other sensory modalities. This research examined whether or not imagery is experienced in similar ways by people with and without sight. Specifically, the imageability of visual, auditory, and tactile cue words was compared. The degree to which images were multimodal or unimodal was also examined. Methods Twelve participants who were totally blind from early infancy and 12 sighted participants generated images in response to 53 sensory and nonsensory words, rating imageability and the sensory modality, and describing images. From these 53 items, 4 subgroups of words that stimulated images that were predominantly visual, tactile, auditory, and low-imagery were created. Results T-tests comparing imageability ratings from blind and sighted participants found no differences for auditory and tactile words (both p >.1). Nevertheless, although participants without sight found auditory and tactile images equally imageable, sighted participants found images in response to tactile cue words harder to generate than visual cue words (mean difference: −0.51, p =.025). Participants with sight were also more likely to develop multisensory images than were participants without sight (both U ≥ 15.0, N 1 = 12, N 2 = 12, p ≤.008). Discussion For both the blind and sighted groups, auditory and tactile images were rich and varied, and similar language was used. Sighted participants were more likely to generate multimodal images, and this was particularly the case for tactile words. Nevertheless, cue words that resulted in multisensory images were not necessarily rated as more imageable. The discussion considers whether or not multimodal imagery represents a method of compensating for impoverished unimodal imagery. Implications for practitioners Imagery is important not only as a mnemonic in memory rehabilitation, but also in everyday uses for modes such as autobiographical memory. This research emphasizes the importance of not only auditory and tactile sensory imagery, but also spatial imagery for people without sight.
Similarity between words is becoming a generic problem for many applications of computational linguistics, and computing word similarities is determined by word representations. Inspired by the analogies between words and lymphocytes, a lymphocyte-style word representation is proposed. The word representation is built on the basis of dependency syntax of sentences and represent word context as head properties and dependent properties of the word. Lymphocyte-style word representations are evaluated by computing the similarities between words, and experiments are conducted on the Penn Chinese Treebank 5.1. Experimental results indicate that the proposed word representations are effective.
The Open Library of Affective Foods (OLAF) is a set of original food pictures created to reliably select food pictures based on the emotions they prompt, as indicated by affective ratings of valence, arousal, and dominance and by an additional food craving scale. OLAF images were designed to allow simultaneous use with affective images from the International Affective Picture System (IAPS), which is a well-known instrument to investigate emotional reactions in the laboratory. The ultimate goal of the OLAF is to contribute to understanding how food is emotionally processed in healthy individuals and in patients who suffer from eating and weight-related disorders. The present normative data, which was based on a large sample of an adolescent population, indicate that when viewing affective non-food IAPS images, valence, arousal, and dominance ratings were in line with expected patterns based on previous emotion research. Moreover, when viewing food pictures, affective and food craving ratings were consistent with research on food cue processing. As a whole, the data supported the methodological and theoretical reliability of the OLAF ratings, therefore providing researchers with a standardized tool to reliably investigate the emotional and motivational significance of food.
We present HamleDT - a HArmonized Multi-LanguagE Dependency Treebank. HamleDT is a compilation of existing dependency treebanks (or depen- dency conversions of other treebanks), transformed so that they all conform to the same annotation style. In the present article, we provide a thorough investigation and discussion of a number of phenomena that are comparable across languages, though their annotation in treebanks often differs. We claim that transformation procedures can be designed to automatically identify most such phenomena and convert them to a unified annotation style. This unification is beneficial both to comparative corpus linguistics and to machine learning of syntactic parsing.
The authors use a self-built Chinese Discourse Treebank(80% relations are implicit) to recognize implicit relations. In this corpus, discourse relations are divided into three layers, the first layer has four types: causality, coordination, transition and explanation. Based on this corpus, maximum entropy classifier is employed to identify four types relations with context, lexical and dependency parse features. Experimental results show that total accuracy is 62.15% and the identification effect of coordination is the best, F1 reaches 75.26%.
The article is devoted to the problem of translation from French into Polish from the perspective of the object oriented approach proposed by Wiesław Banyś. The author takes into consideration some of the problems, both of theoretical and practical nature, which appear while working on the formation of contrastive lexical database for automatic translation of texts. While analyzing specific chosen examples which may cause different kinds of problems in the description, the author offers their interpretations in the target language in accordance with the adopted approach.
En este artículo se describe el proceso que se siguió para crear el componente anotado con información lingüística (treebank) del Corpus de Mensajes Presidenciales Costarricenses (CODIMEP-CR), en el marco del proyecto No. 745-B1-244 Interfaz para el procesamiento de corpus lingüísticos digitales-IPROCOLDI. Ambos corpus se albergan en la interfaz IPROCOLDI (http://163.178.116.145/iprocoldi/).
Many languages, including Modern Stan-dard Arabic (MSA), insert resumptive pro-nouns in relative clauses, whereas many others, such as English, do not, using empty categories instead. This discrep-ancy is a source of difficulty when trans-lating between these languages because there are words in one language that cor-respond to empty categories in the other, and these words must either be inserted or deleted—depending on translation di-rection. In this paper, we first examine challenges presented by resumptive pro-nouns in MSA-English translations and re-view resumptive pronoun translations gen-erated by a popular online MSA-English MT engine. We then present what is, to the best of our knowledge, the first system for automatic identification of resumptive pronouns. The system achieves 91.9 F1 and 77.8 F1 on Arabic Treebank data when using gold standard parses and automatic parses, respectively. 1
Prior research suggested the possibility of establishing systematic linkages between some intrinsic features of a presupposition and textual and pragmatic functions that it can carry out with greater probability. This study aims, firstly, to provide an organic view of semantics of presupposition triggers, thanks to a lexical database comprising 19,500 entries. Secondly, the database was used to investigate a corpus of chat conversations including about 200,000 tokens. The results show that triggers occur mainly as non-informative, maintaining an information already known by all participants of the communication; but, depending on their different features, some of them are systematically associated to a function of anaphora and textual cohesion; others to strengthen social conventions and stereotypes. The informative function, although in a minority proportion, is quantitatively significant only in correspondence to a single class of presupposition triggers.
The French Lexical Network (fr-LN) is a global model of the French lexicon presently under construction. The fr-LN accounts for lexical knowledge as a lexical network structured by paradigmatic and syntagmatic relations holding between lexical units. This paper describes how morphological knowledge is presently being introduced into the fr-LN through the implementation and lexicographic exploitation of a dynamic morphological model. Section 1 presents theoretical and practical justifications for the approach which we believe allows for a cognitively sound description of morphological data within semantically-oriented lexical databases. Section 2 gives an overview of the structure of the dynamic morphological model, which is constructed through two complementary processes: a Morphological Process--section 3--and a Lexicographic Process--section 4.