Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
BACKGROUND: Cystic fibrosis (CF) is a progressive, life-shortening disease currently treated symptomatically. Among the consequences of daily treatments and long-term progression of the disease is significant discomfort and pain. This study measured pain systematically in adolescents with CF using a pain diary, and evaluated its associations with treatment adherence, psychological symptoms, and health-related quality of life (HRQOL). METHODS: This study was part of a larger, multi-center study. Ninety-five adolescents with a mean age of 15.8 were enrolled. A total of 413 online pain diaries were completed for 6 days following a routine clinic visit. Diaries measured pain intensity, location, duration, affective ratings, and coping responses. Adherence to pulmonary medications was measured using prescription refill data; other measures included ratings of depression, anxiety and HRQOL. RESULTS: Pain was reported by 74.5% of participants, generally in the mild range, averaging 2.1 on a 10-point scale. Daily pain ratings were highly variable both within and between participants. Pain was significantly associated with worse adherence, more psychological distress, and worse HRQOL. CONCLUSIONS: Results indicated that pain is a common problem for adolescents with CF and negatively affects their disease management, psychological symptoms and health outcomes. Routine assessment of pain and systematic studies of interventions to treat pain are recommended. Pediatr Pulmonol. 2015; 50:244-251. © 2014 Wiley Periodicals, Inc.
Objectives: This study seeks to provide normative data on the production frequency and semantic typicality and familiarity of noun exemplars by semantic category for elderly individuals. Methods: A total of 198 individuals participated in the study. In Experiment 1, participants were categorized into two groups: professionals involved in dementia care and primary care-givers of patients with dementia. They ranked 17 semantic categories based on the priority of the stimuli required for naming treatment. Forty-five normal elderly individuals (NEIs) participated in Experiment 2, in which they were administered a generative naming task. Seventy-eight NEIs participated in Experiment 3, where they rated the semantic typicality and familiarity of each noun exemplar obtained from Experiment 2. Results: Experiment 1 revealed that 'food (dish)' was ranked as a top priority followed by 'clothes', 'fruits', and 'body parts'. In Experiment 2, 'body parts' generated the highest mean of production frequency among 17 semantic categories. Results from Experiment 3 suggested that the semantic typicality was positively and strongly correlated with the familiarity ratings. Conclusion: These results are clinically important and meaningful given that they provide preliminary guidelines to select stimuli for the naming treatment of individuals with neurogenic language disorders.
AIMS: To explore the enhancing effect of alcohol consumption on attractiveness ratings, in that few studies on the Beer Goggles effect control the stimuli attractiveness level and researchers have seldom considered extending the effect to stimuli other than faces. METHODS: Male and female participants (n = 103) were randomly assigned to alcohol consumption or placebo groups. Both groups were asked to assess the attractiveness of two types of pictures (faces and landscapes) with three levels of attractiveness for each stimulus category (high, moderate and low). RESULTS: We found significant interactions between beverage type and attractiveness level. Attractiveness ratings for moderate- and low-attractiveness faces were significantly higher in the alcohol compared with placebo condition, while there was no significant difference for high-attractiveness stimuli between these two conditions. As for landscapes, only low-attractiveness stimuli were rated significantly higher in the alcohol condition. CONCLUSION: Whether or not alcohol consumption leads to an increase in attractiveness ratings depends on the initial attractiveness of the stimulus materials. Alcohol consumption tends to affect ratings for stimuli with relatively low attractiveness. Furthermore, this effect is not limited to faces; it extends to other types of stimuli like landscapes.
Key to fast adaptation of language technologies for any language hinges on the availability of fundamental tools and resources such as monolingual/parallel corpora, annotated corpora, part-of-speech (POS) taggers, parsers and so on. The languages which lack those fundamental resources are often referred as under-resourced\nlanguages.\n\nIn this thesis, we address the problem of cross-lingual dependency parsing of under-resourced languages. We apply three methodologies to induce dependency structures: (i) projecting dependencies from a resource-rich language to under-resourced languages via parallel corpus word alignment links (ii) parsing under-\nresourced languages using parsers whose models are trained on treebanks of other\nlanguages, and do not look at actual word forms, but only on POS categories. Here\nwe address the problem of incompatibilities in annotation styles between source side parsers and target side evaluation treebanks by harmonizing annotations to a common standard; and finally (iii) we add a new under-resourced scenario in which we use machine translated parallel corpora instead of human translated corpora for\nprojecting dependencies to under-resourced languages.\n\nWe apply the aforementioned methodologies to five Indian languages (ILs): Hindi, Urdu, Telugu, Bengali and Tamil (in the order of high to low availability of treebank data). To make the evaluation possible for Tamil, we develop a depen\ndency treebank resource for Tamil from scratch and we use the created data in\nevaluation and as a source in parsing other ILs. Finally, we list out strategies that\ncan be used to obtain dependency structures for target languages under different\nresource-poor scenarios.
We report of the procedures of developing a large representative corpus of 50,000 sentences taken from clinical notes. Previous reports of annotated corpus of clinical notes have been small and they do not represent the whole domain of clinical notes. The sentences included in this corpus have been selected from a very large raw corpus of ten thousand documents. These ten thousand documents are sampled from an internal repository of more than 700,000 documents taken from multiple health care providers. Each of the documents is de-identified to remove any PHI data. Using the Penn Treebank tagging guidelines with a bit of modifications, we annotate this corpus manually with an average inter-annotator agreement of more than 98%. The goal is to create a parts of speech annotated corpus in the clinical domain that is comparable to the Penn Treebank and also represents the totality of the contemporary text as used in the clinical domain. We also report the output of the TnT tagger trained on the initial 21,000 annotated sentences reaching a preliminary accuracy of above 96%.]
The authors use a self-built Chinese Discourse Treebank(80% relations are implicit) to recognize implicit relations. In this corpus, discourse relations are divided into three layers, the first layer has four types: causality, coordination, transition and explanation. Based on this corpus, maximum entropy classifier is employed to identify four types relations with context, lexical and dependency parse features. Experimental results show that total accuracy is 62.15% and the identification effect of coordination is the best, F1 reaches 75.26%.
The Open Library of Affective Foods (OLAF) is a set of original food pictures created to reliably select food pictures based on the emotions they prompt, as indicated by affective ratings of valence, arousal, and dominance and by an additional food craving scale. OLAF images were designed to allow simultaneous use with affective images from the International Affective Picture System (IAPS), which is a well-known instrument to investigate emotional reactions in the laboratory. The ultimate goal of the OLAF is to contribute to understanding how food is emotionally processed in healthy individuals and in patients who suffer from eating and weight-related disorders. The present normative data, which was based on a large sample of an adolescent population, indicate that when viewing affective non-food IAPS images, valence, arousal, and dominance ratings were in line with expected patterns based on previous emotion research. Moreover, when viewing food pictures, affective and food craving ratings were consistent with research on food cue processing. As a whole, the data supported the methodological and theoretical reliability of the OLAF ratings, therefore providing researchers with a standardized tool to reliably investigate the emotional and motivational significance of food.
Similarity between words is becoming a generic problem for many applications of computational linguistics, and computing word similarities is determined by word representations. Inspired by the analogies between words and lymphocytes, a lymphocyte-style word representation is proposed. The word representation is built on the basis of dependency syntax of sentences and represent word context as head properties and dependent properties of the word. Lymphocyte-style word representations are evaluated by computing the similarities between words, and experiments are conducted on the Penn Chinese Treebank 5.1. Experimental results indicate that the proposed word representations are effective.
Temperature and chemesthesis interact, but this interaction has not been fully examined for most irritants. The current experiments focus on oral pungency from carbonation. Previous work showed that cooling carbon dioxide (CO2) solutions to below tongue temperature enhanced rated bite. However, to the best of our knowledge, the effects of warming to above tongue temperature have not been examined. In Experiment 1, subjects sampled CO2 solutions at 4 nominal concentrations (0.0, 2.0, 2.8, and 4.0 v/v) × 5 temperatures (18.3, 24.5, 29.9, 34.5, and 39.6 (o)C). Subjects dipped their tongue tips into samples and rated bite. As in previous work, subjects rated cool solutions (25.0 (o)C and lower) as more intense. Warming solutions above tongue temperature (39.6 (o)C) did not affect ratings. Experiment 2 examined warmer temperatures (18.3, 33.9, 39.0, 44.9, and 48.2 ºC). Bite was enhanced only at 48.2 ºC, and a follow-up experiment suggested that enhancement was probably due to confusion between carbonation bite and mild heat pain. Experiment 3 examined the effect of menthol cooling by pretreating the tongue with menthol. Unlike physical cooling, menthol cooling had little or no effect on rated bite. The results are discussed in the context of candidate transduction mechanisms for carbonation sensation.
We investigate the feasibility of aligning Chinese and English parse trees by examining cases of incompatibility between Chinese-English parallel parse trees. This work is done in the context of an annotation project where we construct a parallel treebank by doing word and phrase alignments simultaneously. We discuss the most common incompatibility patterns identified within VPs and NPs and show that most cases of incompatibility are caused by divergent syntactic annotation standards rather than inherent cross-linguistic differences in language itself. This suggests that in principle it is feasible to align the parallel parse trees with some modification of existing syntactic annotation guidelines. We believe this has implications for the use of parallel parse trees as an important resource for Machine Translation models.
Signal detection in clinical trials relies on ratings reliability. We conducted a reliability analysis of site-independent rater scores derived from audio-digital recordings of site-based rater interviews of the structured Brief Psychiatric Rating Scale (BPRS) in a schizophrenia study. "Dual" ratings assessments were conducted as part of a quality assurance program in a 12-week, double-blind, parallel-group study of PF-02545920 compared to placebo in patients with sub-optimally controlled symptoms of schizophrenia (ClinicalTrials.gov identifier NCT01939548). Blinded, site-independent raters scored the recorded site-based BPRS interviews that were administered in relatively stable patients during two visits prior to the randomization visit. We analyzed the impact of BPRS interview length on "dual" scoring variance and discordance between trained and certified site-based raters and the paired scores of the independent raters. Mean total BPRS scores for 392 interviews conducted at the screen and stabilization visits were 50.4±7.2 (SD) for site-based raters and 49.2±7.2 for site-independent raters (t=2.34; p=0.025). "Dual" rated total BPRS scores were highly correlated (r=0.812). Mean BPRS interview length was 21:05±7:47min ranging from 7 to 59min. 89 interviews (23%) were conducted in less than 15min. These shorter interviews had significantly greater "dual" scoring variability (p=0.0016) and absolute discordance (p=0.0037) between site-based and site-independent raters than longer interviews. In-study ratings reliability cannot be guaranteed by pre-study rater certification. Our findings reveal marked variability of BPRS interview length and that shorter interviews are often incomplete yielding greater "dual" scoring discordance that may affect ratings precision.
WordNet is an electronic lexical database available on-line as a powerful resource to the researchers in the area of computational linguistics, text processing and other related areas. WordNet for Hindi language has already been developed by IIT, Bombay. The Indian languages WordNets are being created using expansion approach from Hindi WordNet under IndoWordNet project. In expansion approach, semantic relations are borrowed from the reference language, while the lexical relations need to be created for each language, as these relations are language dependent. This paper describes the process of creation of lexical relations like antonym, compounding, conjunction and gradation for IndoWordNet. A lexical creation tool has been presented in this paper with provision to create lexical relations in target language on the basis of relations created in Hindi WordNet and with another provision to create lexical relations in target language without referring to Hindi WordNet. It has been observed that lexical relations for target language can be created easily on the basis of relations created in Hindi WordNet for Hindi in-family languages, while for the languages that do not fall in the same family provision of creation of lexical relation without referring to Hindi WordNet can be used.
The present paper explored the relationship between emotional facial response and electromyographic modulation in children when they observe facial expression of emotions. Facial responsiveness (evaluated by arousal and valence ratings) and psychophysiological correlates (facial electromyography, EMG) were analyzed when children looked at six facial expressions of emotions (happiness, anger, fear, sadness, surprise and disgust). About EMG measure, corrugator and zygomatic muscle activity was monitored in response to different emotional types. ANOVAs showed differences for both EMG and facial response across the subjects, as a function of different emotions. Specifically, some emotions were well expressed by all the subjects (such as happiness, anger and fear) in terms of high arousal, whereas some others were less level arousal (such as sadness). Zygomatic activity was increased mainly for happiness, from one hand, corrugator activity was increased mainly for anger, fear and surprise, from the other hand. More generally, EMG and facial behavior were highly correlated each other, showing a "mirror" effect with respect of the observed faces.
Lexical Knowledge base such as WordNet has been used as a valuable tool for measuring semantic similarity in various Information Retrieval (IR) applications. It is a domain independent lexical database. Since, the quality of semantic relationship in WordNet has not upgraded appropriately for the current usage in the modern IR. Building the WordNet from scratch is not an easy task for keeping updated with current terminology and concepts. Therefore, this paper undergoes a different perspective that automatically updates an existing lexical ontology uses knowledge resources such as the Wikipedia and the Web search engine. This methodology has established the recently evolving relations and also aligns the existing relations between concepts based on its usage over time. It consists of three main phases such as candidate article generation, lexical relationship extraction and generalization and WordNet alignment. In candidate article generation, disambiguation mapping disambiguates ambiguous links between WordNet concepts and Wikipedia articles and returns a set of word-article pairings. Lexical relationship extraction phase includes two algorithms, Lexical Relationship Retrieval (LRR) algorithm discovers the set of lexical patterns exists between concepts and sequential pattern grouping algorithm generalizes lexical patterns and computes corresponding weights based on its frequencies. Furthermore, Sequential Minimal Optimization (SMO) selects the suitable good pattern using the optimal combination of weight of lexical patterns and page count based concurrence measures. WordNet alignment phase establishes a new relationship that is not available in WordNet and also aligns the existing patterns based on computed weight. Experimental results illustrate that the proposed approach better than existing mechanisms on benchmark datasets and achieves a correlation value of 0.87. Moreover, the extended WordNet returns high accuracy results in query expansion.
Helping behavior as a prosocial action emerges early in childhood and is of interest for psychologists in a broad range of sub disciplines as well as for society. One necessary precondition for active helping is the ability to recognize that somebody needs help. The NeoHelp stimulus set used in this study was developed to enable the assessment and quantification of need-of-help recognition abilities. Previous research with the NeoHelp stimuli has shown that children of different ages are able to recognize their content. Specific effects of age and gender on need-of-help recognition have also been observed. How children subjectively experience these stimuli and thus how they rate depictions of need-of-help and no-need-of-help situations emotionally has not been assessed before. Here we report analyses of valence and arousal ratings for the complete NeoHelp stimulus set obtained from a diverse sample of 46 children. We employed the SAM-scales because they are an established rating instrument validated for diverse populations of adults. However, their use with children still needs further investigation. Thus, there were two main goals of the presented study: 1) Validating that the SAM arousal and valence scales may be used with young children below school age, and 2) investigating children's subjective emotional experience of need-of-help depictions. Our study demonstrates that the SAM scales, if properly explained, may be used reliably with children at and above five years of age. Ratings of younger and older children covered the whole range of the 5-point scales used. There was a linear relationship between arousal and valence ratings across all pictures: the higher the arousal ratings, the lower the valence ratings. Pictures showing a child in need-of-help were rated as lower in valence and higher in arousal than the corresponding no-need-of-help-stimuli regardless of children’s age or gender. With increasing age, arousal ratings for no-need-of-help depictions decreased, but arousal ratings for need-of-help depictions remained on the same higher level across ages. We thus provide first evidence that need-of-help depictions elicit differential subjective emotional responses in children on both, valence and arousal dimensions. This emotional component of need-of-help recognition has to be considered when assessing children's need-of-help recognition abilities.
Research in Sentiment Analysis has shown rapid progress since late 90s. It is an important research area as analyzing user’s feedback is useful for business analysis, product comparison, counter intelligence, and poll prediction. Despite the rapid surge of Sentiment Analysis research, many unresolved research questions remain. One of the biggest concerns is the Semantic Gap, which involves translating machine understandable form to human understandable form. Though research has been carried out for machine to understand human language, it is still not capable to address the problem mentioned as human languages are diverse and complex. WordNet, for example, attempt to address this issue by incorporating large lexical database for English, with various functionalities to manipulate this database. Recently, WordNet provides multilingual support, which is very helpful to address the diverse human languages. In this paper, we propose a novel multilingual common ontology tool to analyze user’s feedback and opinion. Unlike other existing state of the art tools, our tool is capable of handling multi languages regardless of the webpage layout. Experimental results show that our tool is highly efficient in analyzing opinion from social networking sites.
Facial expressions are one of the most important ways of non-verbal communication for humans. To date, most research in this field has focused solely on emotional aspects, largely neglecting the communicative and conversational aspects of expressions. Furthermore, although there is evidence for some degree of cross-cultural universality among emotional expressions, much less is known about how facial expressions in general are perceived across cultures. Here, we investigate the structure of the complex space of both emotional and conversational expressions in a cross-cultural context. The two experiments reported here used matching video sequences of 27 expressions from both the KU (Korean) facial expression database and the MPI (German) facial expression database (each expression was shown by 6 actors, totaling 162 videos from each database). In the first experiment, four groups (each n=20) of native German and Korean participants were asked to group the sequences of the German or Korean databases into clusters based on similarity. This grouping data yielded four different confusion matrices. In the second experiment, another four groups of participants (each n=20) from both cultures were asked to rate each video according to 13 emotional and conversational attributes. This rating data yielded an averaged 13-dimensional vector for each sequence. For each of the four grouping/rating data-pairs, we then used kernel canonical correlation analysis (KCCA) to determine a two-dimensional embedding of expressions that best explained both grouping and rating data. Although other attributes contributed as well, the two dimensions recovered by KCCA showed maximal correlation with valence and arousal ratings – this was true regardless of participants' cultural backgrounds or of the database that was used. Our results show that evaluative dimensions for both German and Korean cultural contexts are highly similar, confirming that cultural universals exist even in this complex space of emotional and conversational facial expressions. Meeting abstract presented at VSS 2014
My study assessed the relationship between the colour of depicted clothing and ratings of emotional intensity in drawings of emotional scenarios. Participants (N = 42) viewed a set of drawings in one of seven colours, labeled the Actor and Cause, and rated the intensity of the emotions depicted on 11 emotional scales. Participants also listed colours they associated with specific emotions. Colour of clothing in drawings was not found to affect ratings of emotional intensity, but in drawings of sadness, embarrassment, and empathy the Actor obtained higher ratings than did the Cause. Colour-emotion associations obtained separately were, however, generally consistent with previous findings. COLOUR AND EMOTION 3 Colour and Emotional Intensity Emotion recognition is an evolutionary advantage that can inform behavioural decisions (Elfenbein, Foo, White, Tan & Aik, 2007) and help individuals to effectively navigate social situations (Van Kleef, 2010). It is a process that is largely automatic and unconscious, thus requiring limited cognitive resources (Tracy & Robins, 2008). However, researchers in the areas of social, cognitive and perceptual psychology are still attempting to determine the factors contributing to the automaticity of emotion perception.
Abstract: Similarity is criteria of measuring nearness or proximity between two concepts. Several algorithmic approaches for computing similarity have been proposed. Among the existing Similarity measure, majority of them utilize WordNet as an underlying ontology for calculating semantic similarity. WordNet is a lexical database for English Language which was created and maintained by Congnitive Science Laboratory at Princeton University under the supervision of Professor George A. Miller. It is organized as a network which consists of concepts or terms called Synsets (list of synonyms terms) and the relationship between them. There are different type of relationship exists in WordNet such as is-a, part-of, synonym and antonym. It has thdatabases, one for noun, one for verb and one for adverb and adjective. This project work proposes a metric for semantic relatedness calculation between pair of concepts which uses Tversky’s feature based approach which takes into account the common and distinct feature of the two terms or concepts. If commonality is more as compared to differences the similarity between concepts is high otherwise similarity is low. Tversky’s theory is quantified by information content of two concepts and the Information content of most specific common ancestor of two concepts. As we move down in the WordNet hierarchy, more specific and more Informative concept are there, where as when we move up in the hierarchy more Generalized and less Informative concepts are there. So depth of a concept in the WordNet hierarchy is a critical factor in similarity calculation. We take into consideration the depth of the specific concept in the WordNet hierarchy which is the deciding factor for determining the relevance of distinct feature specific to a concept in similarity calculation. Introduction of depth reduces the impact of the less relevant dissimilarity indulge in similarity calculation thereby increase precision. We carried out our experiment of 28
Abstract Discourse parsing has become an inevitable task to process information in the natural language processing arena. Parsing complex discourse structures beyond the sentence level is a significant challenge. This article proposes a discourse parser that constructs rhetorical structure (RS) trees to identify such complex discourse structures. Unlike previous parsers that construct RS trees using lexical features, syntactic features and cue phrases, the proposed discourse parser constructs RS trees using high‐level semantic features inherited from the Universal Networking Language (UNL). The UNL also adds a language‐independent quality to the parser, because the UNL represents texts in a language‐independent manner. The parser uses a naive Bayes probabilistic classifier to label discourse relations. It has been tested using 500 Tamil‐language documents and the Rhetorical Structure Theory Discourse Treebank, which comprises 21 English‐language documents. The performance of the naive Bayes classifier has been compared with that of the support vector machine (SVM) classifier, which has been used in the earlier approaches to build a discourse parser. It is seen that the naive Bayes probabilistic classifier is better suited for discourse relation labeling when compared with the SVM classifier, in terms of training time, testing time, and accuracy.
In recent years, the emergence of English as an International Language (EIL) has paved the way for its global speakers to use it as a means of interacting globally, and representing themselves and their cultures internationally. Although English is globally considered as an international language and as a tool to be used in cross-cultural communication with people having various first languages from different parts of the world, native-speakers’ norms and cultures still dominate the language materials that are developed to be globally used. In fact, English language coursebooks insists on bombarding the ELT world with culturally-loaded native-speaker themes, such as actors in Hollywood (Coskun, 2009). Prodromou (1988) similarly underlines the issue that the majority of English language coursebooks are published by major Anglo-American publishers in Inner Circle countries and these coursebooks include cultural situations that most students will never come across, such as ‘finding a flat in London’ (p. 80). Considering the importance given to the growing role of EIL, the issue of linguistic norms and cultural content in language learning materials has remained one of the unresolved problems in the process of materials development. A group of scholars argues in favor of localizing the materials by using the learners’ experiences and making English language coursebooks culturally responsive to their needs. The opponents solely favor the integration of the linguistic and cultural norms of the native speakers of English in language learning materials. As far as EIL is concerned, there are several aspects that need to be taken into close account when language teaching materials are being prepared to be globally used. In a nutshell, in EIL era, while preparing English language coursebooks, rather than just integrating English of Specific Cultures, the linguistic and cultural norms of the native speakers of English, as the sole reference in the contents of the English language coursebook, at least a due attention should be paid to English for Specific Cultures, the linguistifc and cultural norms of non-native speakers of English. This study recommends a group of essential features for the future English language coursebooks in EIL era.
Dependency parsers, which are widely used in natural language processing tasks, employ a representation of syntax in which the structure of sentences is expressed in the form of directed links (dependencies) between their words. In this article, we introduce a new approach to transition‐based dependency parsing in which the parsing algorithm does not directly construct dependencies, but rather undirected links, which are then assigned a direction in a postprocessing step. We show that this alleviates error propagation, because undirected parsers do not need to observe the single‐head constraint, resulting in better accuracy. Undirected parsers can be obtained by transforming existing directed transition‐based parsers as long as they satisfy certain conditions. We apply this approach to obtain undirected variants of three different parsers (the Planar, 2‐Planar, and Covington algorithms) and perform experiments on several data sets from the CoNLL‐X shared tasks and on the Wall Street Journal portion of the Penn Treebank, showing that our approach is successful in reducing error propagation and produces improvements in parsing accuracy in most of the cases and achieving results competitive with state‐of‐the‐art transition‐based parsers.
Ratings of previously ignored visual stimuli reveal affective devaluation of such items when compared to ratings of novel items or the targets of attention. Growing evidence suggests this effect may reflect negative affective associations elicited by attentional inhibition of visual distractors. Here we investigate whether such 'inhibitory devaluation' is limited to situations involving visual-spatial selection of environmental stimuli (i.e., external attention) or extends to the selection of competing visual representations held solely in memory (i.e., internal attention). A two-item target-localization task in Experiment 1 utilized a delayed target-category cue ('circles' or 'squares') to ensure attentional selection occurred from the contents of working memory. An n-back task in Experiment 2 was used to examine the affective consequences of rejecting continually-updated visual representations when items held in memory did not match the corresponding visual display. And a Think/No-think paradigm employed in Experiment 3 was designed to explore the affective consequences of actively suppressing longer-term visual object memories. Across this relatively-wide range of memory-based selection tasks, the ignored/rejected/suppressed visual patterns consistently received more negative affective ratings than target items. Our results are consistent with prior suggestions that similar mechanisms are involved in the attentional selection of environmental stimuli and the selection of internally-maintained information that occurs even in the absence of external sensory stimulation. The similarity in these mechanisms appears to extend not only to processes of attentional selection, per se, but also to their affective consequences. Meeting abstract presented at VSS 2014
Automatically acquiring semantic verb classes from corpora is a challenging task, especially with no existing treebank. Building a high-performing parser for a language is still crucially depends on the existence of large, in-domain texts as training data. While previous work has focused primarily on major languages, how to extend these results to other languages is the way to avoid working start from scratch. In general, a large monolingual corpus in a resource-rich source language labeled with lexico-syntactic information, and a very limited bilingual corpus are available. This paper addresses the problem of verb classification automatically in Tibetan using bilingual lexicon and translation information.
Opinion mining is becoming of high importance with the availability of opinionated data on the Internet and the different applications it can be used for. Intensive efforts have been made to develop opinion mining systems, and in particular for the English language. However, models for opinion mining in Arabic remain challenging due to the complexity and rich morphology of the language. Previous approaches can be categorized into supervised approaches that use linguistic features to train machine learning classifiers, and unsupervised approaches that make use of sentiment lexicons. Different features have been exploited such as surface-based, syntactic, morphological, and semantic features. However, the semantic extraction remains shallow. In this paper, we propose to go deeper into the semantics of the text when considered for opinion mining. We propose a model that is inspired by the cognitive process that humans follow to infer sentiment, where humans rely on a database of preconceived notions developed throughout their life experiences. A key aspect for the proposed approach is to develop a semantic representation of the notions. This model consists of a combination of a set of textual representations for the notion (Ti), and a corresponding sentiment indicator (Si). Thus <Ti, Si> denotes the representation of a notion. However, notions can be constructed at different levels of text granularity ranging from ideas covered by words to ideas covered in full documents. The range also includes clauses, phrases, sentences, and paragraphs. To demonstrate the use of this new semantic model of preconceived notions, we develop the full representation of one-word notions by including the following set of syntactic features for Ti: word surfaces, stems, and lemmas represented by binary presence and TFIDF. We also include morphological features such as part of speech tags, aspect, person, gender, mood, and number. As for the notion sentiment indicator Si, we create a new set of features that indicate the words' sentiment scores based on an internally-developed Arabic sentiment lexicon called ArSenL, and using a third-party lexicon called Sifaat. The aforementioned features are extracted at the word-level, and are considered as raw features. We also investigate the use of additional "engineered" features that reflect the aggregated semantics of a sentence. Such features are derived from word-level information, and include count of subjective words, average of sentiment scores per sentence. Experiments are conducted on a benchmark dataset collected from the Penn Arabic TreeBank (PATB) already annotated with sentiment labels. Results reveal that raw word-level features do not achieve satisfactory performance in sentiment classification. Feature reduction was also explored to evaluate the relative importance of the raw features, where the results showed low correlations between individual raw features and sentiment labels. On the other hand, the inclusion of engineered features had a significant impact on classification accuracy. The outcome of these experiments is a comprehensive set of features that reflect the one-word notion or idea representation in a human mind. The results from one-word also show promises towards higher level context with multi-word notions.
Search engines have become the main way for people to get expected information, most of them are based on keyword search. However, keyword search is based on computing the similarity of letters of the keywords, instead of semantic meaning, therefore the searching results often include irrelevant information to user intention. This paper aims to find a way on improving keyword search efficiency. Using Wikipedia, which is the largest online encyclopedia, this paper explores the relations of terms through computing the semantic relatedness between words, and presents an algorithm called WLA in the light of link structure and text message in Wikipedia. What is more, we design a terms query platform through which users will be able to get all the meanings about the concepts. By making a comparison with lexical database WordNet, it has demonstrated the feasibility on our methods.
The well-established memory bias for arousing-negative stimuli seems to be enhanced in high trait-anxious persons and persons suffering from anxiety disorders. We monitored the emergence and development of such a bias during and after learning, in high and low trait anxious participants. A word-learning paradigm was applied, consisting of spoken pseudowords paired either with arousing-negative or neutral pictures. Learning performance during training evidenced a short-lived advantage for arousing-negative associated words, which was not present at the end of training. Cued recall and valence ratings revealed a memory bias for pseudowords that had been paired with arousing-negative pictures, immediately after learning and two weeks later. This held even for items that were not explicitly remembered. High anxious individuals evidenced a stronger memory bias in the cued-recall test, and their ratings were also more negative overall compared to low anxious persons. Both effects were evident, even when explicit recall was controlled for. Regarding the memory bias in anxiety prone persons, explicit memory seems to play a more crucial role than implicit memory. The study stresses the need for several time points of bias measurement during the course of learning and retrieval, as well as the employment of different measures for learning success.
The intent of this study is to determine what sorts of images are considered more interesting by which demographic groups. Specifically, we attempt to identify images whose interestingness ratings are influenced by the demographic attribute of the viewer’s gender. To that end, we use the data from an experiment where 18 participants (9 women and 9 men) rated several hundred images based on “visual interest” or preferences in viewing images. The images were selected to represent the consumer “photo-space” - typical categories of subject matter found in consumer photo collections. They were annotated using perceptual and semantic descriptors. In analyzing the image interestingness ratings, we apply a multivariate procedure known as forced classification, a feature of dual scaling, a discrete analogue of principal components analysis (similar to correspondence analysis). This particular analysis of ratings (i.e., ordered-choice or Likert) data enables the investigator to emphasize the effect of a specific item or collection of items. We focus on the influence of the demographic item of gender on the analysis, so that the solutions are essentially confined to subspaces spanned by the emphasized item. Using this technique, we can know definitively which images’ ratings have been influenced by the demographic item of choice. Subsequently, images can be evaluated and linked, on one hand, to their perceptual and semantic descriptors, and, on the other hand, to the preferences associated with viewers’ demographic attributes.
We propose a novel approach for learning image representation based on qualitative assessments of visual aesthetics. It relies on a multi-node multi-state model that represents image attributes and their relations. The model is learnt from pair wise image preferences provided by annotators. To demonstrate the effectiveness we apply our approach to fashion image rating, i.e., comparative assessment of aesthetic qualities. Bag-of-features object recognition is used for the classification of visual attributes such as clothing and body shape in an image. The attributes and their relations are then assigned learnt potentials which are used to rate the images. Evaluation of the representation model has demonstrated a high performance rate in ranking fashion images.
Resumen Este artículo analiza las actitudes lingüísticas de hablantes nativos de español de la Ciudad Autónoma de Buenos Aires, hacia al español de la Argentina y el español de los otros países hispanohablantes. El artículo es parte de los resultados del Proyecto LIAS (Linguistic Identity and Attitudes in Spanish-speaking Latin America), financiado por El Consejo Noruego de Investigación (RCN). La recolección de los datos se realizó en la capital del país, entrevistando a una muestra de 400 informantes previamente estratificada con las variables de edad, sexo y nivel socioeconómico. El procesamiento estadístico de los datos de campo recolectados arrojó resultados de interés en torno a la mayoría de los tópicos analizados y especialmente en lo referente a aspectos tales como la valoración positiva de la propia variedad lingüística; la resistencia a identificar a España como la única fuente de la norma lingüística de la lengua española; el rechazo a la unificación de la lengua y, por consiguiente, la defensa de la diversidad lingüística como portadora de riqueza cultural. Abstract This article analyzes the linguistic attitudes of native Spanish speakers from Buenos Aires City, towards Spanish spoken in Argentina and in the other Spanish-speaking countries. It is a result of the LIAS-Project (Linguistic Identity and Attitudes in Spanish-speaking Latin America), funded by The Research Council of Norway (RCN). The data were gathered in the capital of the country, interviewing a stratified sample of 400 respondents, based on the variables of age, sex and socioeconomic status. The analysis of the data rendered interesting results on most of the analyzed topics; especially important was the positive appraisal of Argentineans' own linguistic variety; the strong resistance against identifying Spain as the only source of the linguistic norm for the Spanish language; and the rejection of language unification, defending in this way linguistic diversity as an important conveyor of cultural richness.
The authors focus on how to segment semantic units in Chinese discourse and how to label relations among semantic units automatically. During the parsing process, several sequence labelling methods are compared for discourse segmentation, while a maximum entropy-based training and decoding algorithm is specially proposed. Experiments are done based on Tsinghua Chinese Treebank, which is annotated with logical and semantic relations at complex-sentence level. Experimental results show that F-score of discourse segmentation reaches 89.1%. When parsing discourses with no more than 6 relations included, the labeling F-score can achieve 63%.
For languages such as English, several constituent-to-dependency conversion schemes are pro-posed to construct corpora for dependency parsing. It is hard to determine which scheme is better because they reflect different views of dependency analysis. We usually obtain dependen-cy parsers of different schemes by training with the specific corpus separately. It neglects the correlations between these schemes, which can potentially benefit the parsers. In this paper, we study how these correlations influence final dependency parsing performances, by proposing a joint model which can make full use of the correlations between heterogeneous dependencies, and finally we can answer the following question: parsing heterogeneous dependencies jointly or separately, which is better? We conduct experiments with two different schemes on the Penn Treebank and the Chinese Penn Treebank respectively, arriving at the same conclusion that joint-ly parsing heterogeneous dependencies can give improved performances for both schemes over the individual models.
Part-of-speech (POS) taggers can be quite accurate, but for practical use, accuracy often has to be sacrificed for speed. For example, the maintainers of the Stanford tagger (Toutanova et al., 2003; Manning, 2011) recommend tagging with a model whose per tag error rate is 17% higher, relatively, than their most accurate model, to gain a factor of 10 or more in speed. In this paper, we treat POS tagging as a single-token independent multiclass classification task. We show that by using a rich feature set we can obtain high tagging accuracy within this framework, and by employing some novel feature-weight-combination and hypothesis-pruning techniques we can also get very fast tagging with this model. A prototype tagger implemented in Perl is tested and found to be at least 8 times faster than any publicly available tagger reported to have comparable accuracy on the standard Penn Treebank Wall Street Journal test set.
The web today is huge and enormous collection of data today and it goes on increasing day by day. Thus, searching for some particular data in this collection has a significant impact. Researches taking place give prominence to the relevancy and relatedness of the data that is found. Inspite of their relevance pages for any search topic, the results are still huge to be explored. Another important issue to be kept in mind is the users’ standpoint differs from time to time from topic to topic. Effective relevance prediction can help avoid downloading and visiting many irrelevant pages. The performance of a crawler depends mostly on the opulence of links in the specific topic being searched. This paper reviews the researches on web crawling algorithms used for searching. Keywords— Web Crawling Algorithms, Crawling Algorithm Survey, Search Algorithms, Lexical Database, Metadata, Semantic. __________________________________________________*****_________________________________________________
With growing interest in the creation and search of linguistic annotations that form general graphs (in contrast to formally simpler, rooted trees), there also is an increased need for infrastructures that support the exploration of such representations, for example logical-form meaning representations or semantic dependency graphs. In this work, we heavily lean on semantic technologies and in particular the data model of the Resource Description Framework (RDF) to represent, store, and efficiently query very large collections of text annotated with graph-structured representations of sentence meaning. Keywords:Semantic Dependency Graphs, Treebank Search, Resource Description Framework 1.
This paper introduces a new technique for phrase-structure parser analysis, catego-rizing possible treebank structures by inte-grating regular expressions into derivation trees. We analyze the performance of the Berkeley parser on OntoNotes WSJ and the English Web Treebank. This provides some insight into the evalb scores, and the problem of domain adaptation with the web data. We also analyze a “test-on-train ” dataset, showing a wide variance in how the parser is generalizing from differ-ent structures in the training material. 1
Natural language is a fundamental thing of human-society to communicate and interact with one another. In this globalization era, we interact with different regional people as per our interest in social, cultural, economical, educational and professional domain. There are thousands of natural languages exist in our earth. It is quite tough, rather impossible to know all the languages. So we need a computerized approach to convert one natural language to another as per our necessity. This computerized conversion among multiple languages is known as multilingual machine translation. But in this paper we work with a bilingual model, where we concern with two languages: English and Bengali. We use soft computational approach where fuzzy If-Then rule is applied to choose a lemma from prior knowledge; Penn TreeBank PoS tags and HMM tagger are used as lexical class marker to each word in corpora.
This paper investigates the recognition of unknown words in Chinese parsing. Two methods are proposed to handle this problem. One is the modification of a character-based model. We model the emission probability of an unknown word using the first and last characters in the word. It aims to reduce the POS tag ambiguities of unknown words to improve the parsing performance. In addition, a novel method, using graph-based semisupervised learning (SSL), is proposed to improve the syntax parsing of unknown words. Its goal is to discover additional lexical knowledge from a large amount of unlabeled data to help the syntax parsing. The method is mainly to propagate lexical emission probabilities to unknown words by building the similarity graphs over the words of labeled and unlabeled data. The derived distributions are incorporated into the parsing process. The proposed methods are effective in dealing with the unknown words to improve the parsing. Empirical results for Penn Chinese Treebank and TCT Treebank revealed its effectiveness.
Chunking or shallow syntactic parsing is proving to be a task of interest to many natural language processing applications. The problem gets worse for the Arabic language because of its specific features that make it quite different and even more ambiguous than other natural languages when processed. In this paper, we present a method for chunking Arabic texts based on supervised learning. We use the Conditional Random Fields algorithm and the Penn Arabic Treebank to train the model. For the experimentation, we use over than 10,100 sentences as training data and 2,524 sentences for the test. The evaluation of the method consists of the calculation of the generated model accuracy and the results are very encouraging.
Imagine saying to Hegel that his dialectic patterns occur in the behavior of magnets and rivers and trees, and that the self-understanding of spirit is an operation going on in brain tissue. That the descriptions in his Phenomenology of Spirit picture the behavior of social groups and changes in cultural norms and memes. Hegel's reaction would be complex but not hostile. You would find yourself in a discussion with him about different kinds and levels of categories, and the relation of physical science to his overall logic.Now imagine saying to Heidegger that his description of the care structure of Dasein is a sophisticated reworking of folk psychology, and that it is the result of what amounts to a software program. That his Fourfold is a description of the appearing of a world, based upon neurological and social processes. That his history of being is open to sociological analysis and historical and economic explanation. Heidegger would react to such claims more sharply. You would find yourself in a discussion with him about being caught in das Gestell, and the need to step back from attempts to absolutize one language and one revelation of beings.How would these different reactions play out in a discussion of the ontology of the self? This present essay approaches these major continental thinkers with a question from analytic philosophy, to see how they might respond.In standard mind/body discussions we find two rival descriptive languages with different sets of entities. One talks about ideas, purposes, awareness, meanings, concepts, norms, intentions, and so on. The other talks about the behavior of cells and electrical currents, brain activities, and entities described by physics.There are thinkers who argue that folk talk about thoughts and meanings could be in principle abandoned even in everyday life, replaced by descriptions that involved only physical scientific entities. Others argue that folk talk is not dispensable. Daniel Dennett and Wilfrid Sellars would say that we need to take up the intentional stance and avow social and linguistic norms in order to function in a meaningful social world. Kant had already argued that there is a practical necessity to view ourselves as free agents, no matter what our science may say. These considerations suggest that the tension between the modes of discussion cannot be easily wished away by eliminating one of them.DIFFERENT ONTOLOGIESOne could say that these two languages are using different ontologies. I am using ontology here in a sense derived from analytic philosophers who ask your ontology include... [relations, sets, second level properties, mereological wholes, abstract entities, Cartesian souls, etc.]? An ontology in this sense is the list of approved types of entities that are being affirmed as ultimately real. In most such discussions there is little talk about ontology in an older sense, namely, about the mode of being of those beings which analytic discussions often presuppose to be immediate factual presence. Much more active modes of being are affirmed in Whitehead or in Bergson or Deleuze, or in Aristotle's doctrine of potentiality. For none of these thinkers does being equal simple, positive presence. For them such presence is a result, and it is so independently of whatever processes bring it to presence for us in experience.But if we do ask about modes of being, beyond the factual lists, then other more traditional ontological questions arise: In Aristotle's terms, are the movements of Hegel's dialectic substantial changes or accidental, or relational, or what? Is Heidegger's Dasein or Hegel's spirit a substance? Or a set of emergent properties? How do we individuate Dasein(s)? Does Hegel's spirit have its own individuated self-consciousness? What kind of being do the components of the Fourfold have, and how do they relate to everyday objects and scientific entities?Ontology matters. When David Hume looks for his self, he can't find it. …
In this article, we propose the first work that investigates the feasibility of Arabic discourse segmentation into elementary discourse units within the segmented discourse representation theory framework. We first describe our annotation scheme that defines a set of principles to guide the segmentation process. Two corpora have been annotated according to this scheme: elementary school textbooks and newspaper documents extracted from the syntactically annotated Arabic Treebank. Then, we propose a multiclass supervised learning approach that predicts nested units. Our approach uses a combination of punctuation, morphological, lexical, and shallow syntactic features. We investigate how each feature contributes to the learning process. We show that an extensive morphological analysis is crucial to achieve good results in both corpora. In addition, we show that adding chunks does not boost the performance of our system.
We investigate how the granularity of POS tags influences POS tagging, and furthermore, how POS tagging performance relates to parsing results.For this, we use the standard "pipeline" approach, in which a parser builds its output on previously tagged input.The experiments are performed on two German treebanks, using three POS tagsets of different granularity, and six different POS taggers, together with the Berkeley parser.Our findings show that less granularity of the POS tagset leads to better tagging results.However, both too coarse-grained and too fine-grained distinctions on POS level decrease parsing performance.
The aim of this article is to measure the indexes of productivity of the prefix ful - and the suffix - ful in Old English adjective formation. This analysis is based on Baayen’s framework, which comprises different measures on productivity. The major sources of the analysis are The Dictionary of Old English Corpus and the lexical database of Old English Nerthus. This study of productivity allows for a diachronic perspective on the evolution of these affixes from the Old English period to the present. The main conclusion drawn from this analysis is that the suffix -ful is more productive than its prefixal counterpart, which implies that more productive patterns are still maintained in Present-day English in contradistinction to the less productive ones.
We investigate the usefulness of syntactic knowledge in estimating the quality of English-French translations. We find that dependency and constituency tree kernels perform well but the error rate can be further reduced when these are combined with hand-crafted syntactic features. Both types of syntactic features provide information which is complementary to tried-and-tested nonsyntactic features. We then compare source and target syntax and find that the use of parse trees of machine translated sentences does not affect the performance of quality estimation nor does the intrinsic accuracy of the parser itself. However, the relatively flat structure of the French Treebank does appear to have an adverse effect, and this is significantly improved by simple transformations of the French trees. Finally, we provide further evidence of the usefulness of these transformations by applying them in a separate task ‐ parser accuracy prediction.
The architecture of writing systems metaphor has special relevance for understanding the structural nature of the Japanese writing system, and, more specifically, for appreciating how the 2,136 kanji of the 常用漢字表 /jō-yō-kan-ji-hyō/* ‘List of characters for general use’ function as the core building blocks in the orthographic representation of a considerable proportion of the Japanese lexicon. In seeking to illuminate the multiple layers of internal structure within Japanese kanji, the Japanese lexicon, and the Japanese writing system, the paper draws on insights and observations gained from an ongoing project to construct a large-scale Japanese lexical database system. Reflecting structural distinctions within the database, the paper consists of three main sections addressing the different structural levels of kanji components, jōyō kanji, and the lexicon. Keywords: Japanese writing system; building blocks; jōyō kanji; components; orthographic structure; database