Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper presents preliminary investigations on the statistical parsing of French by bringing a complete evaluation on French data of the main probabilistic lexicalized and unlexicalized parsers first designed on the Penn Treebank. We adapted the parsers on the two existing treebanks of French (Abeillé et al., 2003; Schluter and van Genabith, 2007). To our knowledge, mostly all of the results reported here are state-of-the-art for the constituent parsing of French on every available treebank. Regarding the algorithms, the comparisons show that lexicalized parsing models are outperformed by the unlexicalized Berkeley parser. Regarding the treebanks, we observe that, depending on the parsing model, a tag set with specific features has direct influence over evaluation results. We show that the adapted lexicalized parsers do not share the same sensitivity towards the amount of lexical material used for training, thus questioning the relevance of using only one lexicalized model to study the usefulness of lexicalization for the parsing of French.
In his later work in the philosophy of language Davidson analyses communicative exchanges and arrives at the startling and weighty conclusion that linguistic norms and conventions are entirely inessential to linguistic meaning. I argue that this inference is flawed: if we place the account of communication against the backdrop of Davidson's own views about radical interpretation, then it becomes evident that linguistic norms are an essential feature of the Davidsonian picture.
One of the biggest challenges in compiling a dictionary of a minority language is managing the large quantity of lexical data. Decisions about the format and content of the dictionary or the orthography typically evolve over the years that such projects usually take. This results in inconsistencies between older and newer entries. Revising the data for publication as a dictionary introduces further inconsistencies as does having multiple contributors and/or editors. Proofreading a lexical database takes a great deal of time and the richer its structure the more this is the case. The tools described in this presentation significantly reduce this effort. Tools developed for checking the consistency of the lexical database in the Iu Mien—Chinese—English dictionary project have proven extremely helpful. Two basic approaches are used: 1) use of a program written to check for likely errors that scans the lexical database and produces an error report that is used by a lexicographer to make appropriate corrections. 2) outputting the lexical data in alternate forms that make it easier for the lexicographer to spot problem areas. These alternative forms include the reverse indexes and views structured according to semantic domains. The Iu Mien—Chinese—English dictionary project, like many minority language dictionary projects, uses SIL's Toolbox software. It is very flexible software but its capabilities to enforce consistency are quite limited. Some parts of the approach described here are specific to MDF (Multi-Dictionary Formatter) lexical databases in Toolbox but will be equally useful for other MDF databases. Other parts are specific to each of the three languages involved but will be useful for non-Toolbox lexical databases. Every dictionary is unique and this applies not only to content of the entries but also the decisions about how entries should be arranged to suit the languages involved. Other decisions about the structure are likely to be made differently even in other dictionaries of the same languages. It is the way that each dictionary combines themes that are found in many dictionaries that makes them unique, e.g. to be root based or not, to have include subentries. Therefore our approach is to use a toolkit based approach to curating lexical databases. This allows checking techniques to be mixed and matched to suit the unique aspects of a lexical project. The checking software is written in Python and relies on the toolbox module in NLTK (The Natural Language Toolkit http://nltk.sourceforge.net).
Olfactory perception was examined in deficit syndrome (DS) and nondeficit syndrome (ND) schizophrenia patients. Participants included 22 controls (CN) and 41 patients with schizophrenia who were divided into DS (n = 15) and ND (n = 26) subtypes using the Schedule for the Deficit Syndrome (SDS). Olfactory perception for pleasant and unpleasant odors was assessed using the Brief Smell Identification Test. Participants were instructed to identifying each smell as well as provide hedonic judgment ratings of each smell on a 7-point scale (1 = extremely pleasant, 4 = neutral, and 7 = extremely unpleasant). Results indicated that when compared with the ND patients, the DS patients rated pleasant smells as being significantly less pleasant, although no difference between the groups was present for unpleasant smells, and both ND and DS groups significantly differed from CN on rating and identifying pleasant and unpleasant items. Additionally, lower smell identification accuracy was negatively correlated with SDS symptom severity, and valence ratings for pleasant odors were positively correlated with SDS diminished emotional range. Findings suggest that the DS is characterized by a unique pattern of olfactory valence judgment that is characterized by abnormalities in processing positively valenced stimuli.
Parkinson's disease (PD) involves facial masking, which may impair social interaction. Older adult observers who viewed segments of videotaped interviews of individuals with PD expressed less interest in relationships with women with higher masking and judged them as less supportive. Masking did not affect ratings of men in these domains, possibly because higher masking violates gender norms for expressivity in women but not in men. Observers formed less accurate ratings of the social supportiveness and social strain of women than men, and higher masking decreased accuracy for ratings of strain. Results suggest that some of the problems with social relationships in PD may be due to inaccurate impressions and reduced desire to interact with individuals with higher masking, especially women.
The paper first briefly outlines the differences between a lexical data base as it typically results from a language documentation project and the kind of dictionaries the speech community wants for educational purposes with respect to the choice of head words, grammatical information, definitions of meaning and translations, encyclopaedic information and the choice of examples. The second part of the paper then explores how in spite of limited resources in terms of time, money and man power the speech community and the linguists can develop a method of dictionary making that both satisfies the needs of the community and the interests of linguists. Since it is impossible to create a comprehensive dictionary in a language documentation project, we opted for the thematic approach in which lexicographers work on particular semantic domains such as body parts, architecture or fishing, and try to cover all those lexemes of the respective domain that seem to be important for the intended dictionary users. In our project the headwords of the lexical database were classified according to their domains, and then filtered and exported from the lexical database in order to produce a mini dictionary for each selected domain. In the case of body parts, for instance, we did not only select the nouns that signify the body parts, but also verbs that express bodily actions like ‘sweat’ and ‘comb your hair’ as well as speech formulas like ‘have a heavy heart’. The indigenous lexicographers then checked this preliminary mini-dictionary for missing head words and fixed multi-word expressions,and revised the examples which often were so context dependent that they did not make much sense in isolation. Focussing on one particular semantic domain at a time helps to easily identify various kinds of lexical relations such as taxonomies, meronymies and metonymies as well as metaphorical usages, collocational restrictions and grammatical constructions. As one and the same lexeme can belong to more than one semantic domain, the mini-dictionaries must be accompanied by an index. The paper concludes with a discussion of how this kind of practical lexicography is related to frame semantics and how it can be used for semantic typology.
This paper describes the MulTra project, aiming at the development of an efficient multilingual translation technology based on an abstract and generic linguistic model as well as on object-oriented software design. In particular, we will address the issue of the rapid growth both of the transfer modules and of the bilingual databases. For the latter, we will show that a significant part of bilingual lexical databases can be derived automatically through transitivity, with corpus validation.
This study evaluated evidence for 2 forms of emotional abnormality in posttraumatic stress disorder (PTSD): numbing and heightened negative emotionality. Forty-nine male veterans with PTSD and 75 without the disorder rated their emotional responses to photographs that depicted scenes of Vietnam combat or were drawn from the International Affective Picture System (Lang et al., 2005). Images varied in their trauma-relatedness and affective qualities. A series of repeated measures ANOVAs revealed that Vietnam combat veterans with PTSD responded to unpleasant images with greater negative emotionality (i.e., enhanced arousal and lower valence ratings) than those without the disorder and this effect was modified by the trauma-relatedness of the image with stronger effects for trauma-related images. In contrast, the 2 groups showed equivalent patterns of responses to pleasant images. Findings raise questions about the sensitivity of the International Affective Picture System rating protocol for the assessment of PTSD-related emotional numbing.
BACKGROUND: Time-limited group cognitive behavioral treatments (GCBT) for obsessive-compulsive disorder have demonstrated improvement in target symptoms. One small sample study of GCBT specifically for hoarding problems also showed benefit. This study examines the efficacy of a specialized GCBT for compulsive hoarding on a larger sample. METHODS: Thirty-two clients diagnosed with hoarding participated in five groups. Four groups met once weekly for 2 hour over 16 weeks (n=27) and one group met for 20 weeks (n=5). All participants had two individual 90-min home sessions. Self-report assessments were completed at baseline, mid-treatment, and post-treatment about hoarding behavior and related symptoms (e.g., depression). The sample was predominantly female, White, highly educated, unemployed, and not partnered/married; mean age was 53. A majority was diagnosed with major depressive disorder and obsessive-compulsive personality disorder. RESULTS: Participants showed significant improvement from pre- to post-treatment on the Saving Inventory Revised, Saving Cognitions Inventory, Clutter Image Rating, and Clinical Global Severity. The most recent group (n=8) that used a more formalized treatment and research protocol improved significantly more than did earlier members. CONCLUSION: This study demonstrates the feasibility and modest success of GCBT methods in improving hoarding symptoms. Group treatment may be especially valuable because of its cost-effectiveness, greater client access to trained clinicians, and reduction in social isolation and stigma linked to this problem. Further research is needed to improve the efficacy of GCBT methods for hoarding and to examine durability of change, predictors of outcomes, and processes that influence change.
There is strong evidence that human sentence processing is incremental, i.e., that structures are built word by word.Recent experiments show that the processor also predicts upcoming linguistic material on the basis of previous input.We present a computational model of human parsing that is based on a variant of tree-adjoining grammar and includes an explicit mechanism for generating and verifying predictions, while respecting incrementality and connectedness.An algorithm for deriving a lexicon from a treebank, a fully implemented parser, and a probability model for this formalism are also presented.We devise a linking function that explains processing difficulty as a combination of prefix probability (surprisal) and verification cost.The resulting model captures locality effects such as the subject/object relative clause asymmetry, as well as surprisal effects such as prediction in either...or constructions.
In this paper, we focus on facial displays, eye gaze and head tilts to express social dominance. In particular, we are interested in the interaction of different non-verbal cues. We present a study which systematically varies eye gaze and head tilts for five basic emotions and a neutral state using our own graphics and animation engine. The resulting images are then presented to a large number of subjects via a web-based interface who are asked to attribute dominance values to the character shown in the images. First, we analyze how dominance ratings are influenced by the conveyed emotional facial expression. Further, we investigate how gaze direction and head pose influence dominance perception depending on the displayed emotional state.
In this paper, we propose a linear model-based general framework to combine k-best parse outputs from multiple parsers. The proposed framework leverages on the strengths of previous system combination and re-ranking techniques in parsing by integrating them into a linear model. As a result, it is able to fully utilize both the logarithm of the probability of each k-best parse tree from each individual parser and any additional useful features. For feature weight tuning, we compare the simulated-annealing algorithm and the perceptron algorithm. Our experiments are carried out on both the Chinese and English Penn Treebank syntactic parsing task by combining two state-of-the-art parsing models, a head-driven lexicalized model and a latent-annotation-based un-lexicalized model. Experimental results show that our F-Scores of 85.45 on Chinese and 92.62 on English outperform the previously best-reported systems by 1.21 and 0.52, respectively.
This article focuses on every day communication in New Media with special regards to private writing on Instant Messaging. After brief introductory thoughts about writings beyond the linguistic norm in New Media we compare the specific circumstances of "new" writing via internet and mobile phone with "traditional" offline writing that can be realized by the use of a computer, a type writer or by hand. How this new writing is judged by the public, whether it is considered to be "good" or "bad" and how experts position themselves in this discussion, is shown in section 3. Section 4 takes a look at which linguistic theories might apply to the analysis of typed dialogues in computer mediated communication. The main focus here is on the theory of Interactional Linguistics which formerly had been applied only to the analysis of oral communication. Finally, language critical and linguistic aspects of writing in the New Media are discussed in a brief synopsis.
In this paper, we describe and evaluate a bigram part-of-speech (POS) tagger that uses latent annotations and then investigate using additional genre-matched unlabeled data for self-training the tagger. The use of latent annotations substantially improves the performance of a baseline HMM bigram tagger, outperforming a trigram HMM tagger with sophisticated smoothing. The performance of the latent tagger is further enhanced by self-training with a large set of unlabeled data, even in situations where standard bigram or trigram taggers do not benefit from self-training when trained on greater amounts of labeled training data. Our best model obtains a state-of-the-art Chinese tagging accuracy of 94.78% when evaluated on a representative test set of the Penn Chinese Treebank 6.0.
Although gaze direction and face shape have each been shown to affect perceptions of the dominance of others, the question whether gaze direction and face shape have independent main effects on perceptions of dominance, and whether these effects interact, has not yet been studied. To investigate this issue, we compared dominance ratings of faces with masculinised shapes and direct gaze, masculinised shapes and averted gaze, feminised shapes and direct gaze, and feminised shapes and averted gaze. While faces with direct gaze were generally rated as more dominant than those with averted gaze, this effect of gaze direction was greater when judging faces with masculinised shapes than when judging faces with feminised shapes. Additionally, faces with masculinised shapes were rated as more dominant than those with feminised shapes when faces were presented with direct gaze, but not when faces were presented with averted gaze. Collectively, these findings reveal an interaction between the effects of gaze direction and sexually dimorphic facial cues on judgments of the dominance of others, presenting novel evidence for the existence of complex integrative processes that underpin social perception of faces. Integrating information from face shape and gaze cues may increase the efficiency with which we perceive the dominance of others.
Large scale efforts are underway to create dependency treebanks and parsers for Hindi and other Indian languages. Hindi, being a morphologically rich, flexible word order language, brings challenges such as handling non-projectivity in parsing. In this work, we look at non-projectivity in Hyderabad Dependency Treebank (HyDT) for Hindi. Non-projectivity has been analysed from two perspectives: graph properties that restrict non-projectivity and linguistic phenomenon behind non-projectivity in HyDT. Since Hindi has ample instances of non-projectivity (14% of all structures in HyDT are non-projective), it presents a case for an in depth study of this phenomenon for a better insight, from both of these perspectives.
The aim of the present study was to examine the association between familiarity of odors, cued and free odor identification performance and cognitive function in elderly adults. It was further investigated how age affects performance on the various odor tasks. A third aim was to investigate the role of familiarity in explaining performance on the free identification task. One hundred and thirty-six participants (aged 45-79 years) with normal olfactory sensitivity were assessed with the Scandinavian Odor Identification Test (SOIT) and standardized tests of cognitive function. Familiarity did not correlate with any measure of cognitive function, while verbal identification performance was associated with several cognitive measures, although correlations were modest. In this sample, free odor identification was affected by increasing age to a marginally larger extent than cued identification performance and familiarity ratings. The results suggest that the different olfactory tasks involve different levels of cognitive processing.
OBJECTIVES: To explore the effectiveness of acupressure and Montessori-based activities in decreasing the agitated behaviors of residents with dementia. DESIGN: A double-blinded, randomized (two treatments and one control; three time periods) cross-over design was used. SETTING: Six special care units for residents with dementia in long-term care facilities in Taiwan were the sites for the study. PARTICIPANTS: One hundred thirty-three institutionalized residents with dementia. INTERVENTION: Subjects were randomized into three treatment sequences: acupressure-presence-Montessori methods, Montessori methods-acupressure-presence and presence-Montessori methods-acupressure. All treatments were done once a day, 6 days per week, for a 4-week period. MEASUREMENT: The Cohen-Mansfield Agitation Inventory, Ease-of-Care, and the Apparent Affect Rating Scale. RESULTS: After receiving the intervention, the acupressure and Montessori-based-activities groups saw a significant decrease in agitated behaviors, aggressive behaviors, and physically nonaggressive behaviors than the presence group. Additionally, the ease-of-care ratings for the acupressure and Montessori-based-activities groups were significantly better than for the presence group. In terms of apparent affect, positive affect in the Montessori-based-activities group was significantly better than in the presence group. CONCLUSION: This study confirms that a blending of traditional Chinese medicine and a Western activities program would be useful in elderly care and that in-service training for formal caregivers in the use of these interventions would be beneficial for patients
We review lexical Association Measures (AMs) that have been employed by past work in extracting multiword expressions. Our work contributes to the understanding of these AMs by categorizing them into two groups and suggesting the use of rank equivalence to group AMs with the same ranking performance. We also examine how existing AMs can be adapted to better rank English verb particle constructions and light verb constructions. Specifically, we suggest normalizing (Pointwise) Mutual Information and using marginal frequencies to construct penalization terms. We empirically validate the effectiveness of these modified AMs in detection tasks in English, performed on the Penn Treebank, which shows significant improvement over the original AMs.
Patients suffering from major depressive disorder (MDD) have been shown to exhibit increased thresholds towards experimentally induced thermal pain applied to the skin. In contrast, the induction of sad mood can increase pain perception in healthy controls. Here, we aimed to test the hypothesis that heat pain thresholds are further increased after sad mood induction in depressed patients. Thermal pain thresholds were obtained from 25 female depressed patients and 25 controls before and after sad mood induction applying a modified Velten Mood Induction procedure (MIP). Valence and arousal ratings were obtained using the self-assessment manikin. The Montgomery Depression Rating Scale and the Beck Depression Inventory (BDI) were obtained at baseline from all participants. Pain thresholds at baseline did not significantly differ between groups. Pain thresholds and valence of mood significantly decreased both in patients and controls, while arousal showed an inverse time course between groups. Therefore, our hypothesis could not be confirmed. From these data, we propose that the depressed mood as seen in MDD patients influences pain experience differently as compared to the shorter-lasting mood change after MIP. A differential interaction of both affective states with brain areas of the pain matrix might be assumed. Eventually, the induction of sad mood might mirror the increased number of pain complaints in depressed patients and thus adds to the current concept of adjuvant antidepressant treatment both in depressed patients with pain complaints and in chronic pain patients.
We present parsing algorithms for various mildly non-projective dependency formalisms. In particular, algorithms are presented for: all well-nested structures of gap degree at most 1, with the same complexity as the best existing parsers for constituency formalisms of equivalent generative power; all well-nested structures with gap degree bounded by any constant k; and a new class of structures with gap degree up to k that includes some ill-nested structures. The third case includes all the gap degree k structures in a number of dependency treebanks.
This paper describes log-linear models for a general-purpose sentence realizer based on dependency structures. Unlike traditional realizers using grammar rules, our method realizes sentences by linearizing dependency relations directly in two steps. First, the relative order between head and each dependent is determined by their dependency relation. Then the best linearizations compatible with the relative order are selected by log-linear models. The log-linear models incorporate three types of feature functions, including dependency relations, surface words and headwords. Our approach to sentence realization provides simplicity, efficiency and competitive accuracy. Trained on 8,975 dependency structures of a Chinese Dependency Treebank, the realizer achieves a BLEU score of 0.8874.
In this paper, we present our in-progress research tasks for building lexical database of the verb valences in the Arabic Quran using FrameNet frames. We study the verbs in their context in the Quran, and compare that with matching frames and frame evoking verbs in the English FrameNet. We analyze the gaps and make appropriate amendments to the FrameNet by adding new frame elements and relations.
This paper reports results on grammatical induction for French. We investigate how to best train a parser on the French Treebank We compare, for French, a supervised lexicalized parsing algorithm with a semi-supervised unlexicalized algorithm We report the best results known to us on French statistical parsing, that we obtained with the semi-supervised learning algorithm. The reported experiments can give insights for the task of grammatical learning for a morphologically-rich language, with a relatively limited amount of training data, annotated with a rather flat structure.
Regardless of their name (dictionary, glossary, encyclopaedia, or even ‘leximat’, in the case of a new generation of online, semi-automated lexicographic tools), subject-field, purpose, or medium (paper or cyber), lexicographic reference works should be regarded as functional information tools that are solely designed to cater to the information needs of their users in different usage situations and that consequently help them solve specific communication (reading, writing, translation) or knowledge problems (acquiring new knowledge or verifying existing knowledge, learning a language or a subject field). In this article, we briefly outline the evolution of lexicographic reference works from stand-alone to multifunctional lexicographic tools, and we describe the theoretical principles and innovative functionalities of a new task and problem-oriented lexical database, the Base Lexicale du Français. In line with Tarp (2006), a tool that should be truly regarded as a ‘leximat’.
We present a framework for interfacing a PCFG parser with lexical information from an external resource following a different tagging scheme than the treebank. This is achieved by defining a stochastic mapping layer between the two resources. Lexical probabilities for rare events are estimated in a semi-supervised manner from a lexicon and large unannotated corpora. We show that this solution greatly enhances the performance of an unlexicalized Hebrew PCFG parser, resulting in state-of-the-art Hebrew parsing results both when a segmentation oracle is assumed, and in a real-word parsing scenario of parsing unsegmented tokens.
Treebanks play an increasing role in computational linguistics for training parsers. Some treebanking environments allow annotators to quickly navigate through the parse forest and identify the correct or incorrect or preferred analysis in the current context by selecting or rejecting discriminants. Although, these treebanking decisions are recorded in log files or databases, but, to our best knowledge, until now nobody has inspected potentiality of incorporating such fine-grained decisions made by human annotators for automatic parse disambiguation. This thesis examines this new potential research direction by developing a novel approach for extracting discriminative features using treebanking decisions. The thesis presents comparative analyses of the performance of discriminative disambiguation models built using the treebanking decision features and the state-of-the-art features which indicate features extracted using treebanking decisions are more efficient and informative compared to their traditional counterparts. We highlight how these different types of features scale when their corresponding models are tested on out-of-domain data. The result suggests that, treebanking decision features are more robust. Analyses from different perspectives such as impact of different types of decisions on the disambiguation model, or using the disambiguation model of the treebanking decisions feature as a re-ranker are also included. The study also develops a method to extract patterns of correlated discriminant from human decisions and use them for parse forest reduction. The empirical results indicate that, finding such patterns that yields substantial reduction of parse forest preserving the preferred analyses is not an easy task. The thesis argues that, the discriminative nature of the treebanking decisions allows them to be highly effective features to contribute to an efficient disambiguation model. This is demonstrated by a number of experiments that also reveal some open research questions for future works.
The aim of Evalita Parsing Task is at defining and extending the state of the art for parsing Italian by encouraging the application of existing models and approaches. Therefore, as in the first edition, the Task includes two tracks, i.e. dependency and constituency. This second track is based on a development set in a format, which is an adaptation for Italian of the Penn Treebank format, and has been applied by conversion to an existing dependency Italian treebank. The paper describes the constituency track and the data for development and testing of the participant systems. Moreover it presents and discusses the results, which positively compare with those obtained for constituency parsing in the Evalita’07 Parsing Task.
Abstract. We present a simplified Data-Oriented Parsing (DOP) formalism for learning the constituency structure of Italian sentences. In our approach we try to simplify the original DOP methodology by constraining the number and type of fragments we extract from the training corpus. We provide some examples of the types of constructions that occur more often in the treebank, and quantify the performance of our grammar on the constituency parsing task. Keywords: Data-Oriented Parsing, Tree substitution grammar, statistical model, fragments, kernel methods.
In this paper, we present a document visualization technique for data analysis based on the semantic representation of text in the form of a directed graph, referred to as semantic graph. It is derived using natural language processing as follows. Firstly subject– verb – object triplets are automatically extracted from the Penn Treebank parse tree obtained for each sentence in the document. Secondly, the triplets are further enhanced by linking them to their corresponding co-referenced named entity, by resolving pronominal anaphors as well as attaching the associated WordNet synset. Starting from the document's semantic graph and the list of extracted triplets we automatically generate the document summary, for which we also derive the semantic representation.
Linguistic database summaries in the sense of Yager (1982), further extended to an implementable form by Kacprzyk & Yager (2001) and Kacprzyk, Yager & Zadrozny (2000), are extremely simple natural language like statements exemplified by, for a personnel database, “most employees are young and well paid” (with some degree of truth). They have been implemented in business contexts (cf. Kacprzyk & Zadrozny,????, Kacprzyk, Wilbik and Zadrozny, 2006–2008). An effective and efficient way of their generation was proposed by Kacprzyk & Zadrozny (????) by using an interactive procedure based on Kacprzyk & Zadrozny's (????) fuzzy database queries with linguistic quantifiers. Moreover, in Kacprzyk & Zadrozny (???) the role of Zadeh's (???) protoform was shown and their use advocated. Though linguistic database summaries have a strong resemblance to natural language generation (NLG), this issue was never considered. In this paper we indicate some important issues that are common to linguistic database summarization and natural language generation, and propose some possible research directions.
The existence of graded structure in fruit and flower odour categories and its stability in different cultures is examined. Groups of students from France, the United States, and Vietnam performed a typicality rating task, a similarity judgment task, a membership verification task, a recognition memory task, a familiarity rating task, and a free identification task using a set of 40 odorants (20 fruit odorants and 20 flower odorants). Overall, our results demonstrate that fruit and flower odour categories possess graded structure. Moreover, principal component analyses of the data revealed the implication of typicality in a variety of cognitive tasks where typical odours receive a preferential processing compared to atypical ones. Finally, our results suggest that typicality can be predicted to a certain extent by experiential knowledge but that other determinants play a role in odour category structure. Altogether, this study confirms that graded structure is a universal property of categories and suggests that universals and cultural specifics can both constrain the emergence of odour category structures.
Abstract We investigated how viewing positive and negative emotional stimuli influences the functional field of view. Two types of emotional pictures were used, representing either a negative emotion such as disgust or fear, or a positive emotion such vigour or excitement. The participants' task was to detect and identify a digit presented in any of the four corners of a picture on the display, while discriminating a letter in the centre of the display. We used two stimulus onset asynchronies (SOAs), of 500 and 3000 ms, between the picture and the digit. Performance was poorer in the negative condition than in the positive and non-emotional conditions. Performance in the 500 ms SOA condition was poorer than in the 3000 ms SOA condition. These results suggest that the functional field of view (FFOV) becomes narrower when people view negative emotional stimuli, whereas it does not change when viewing positive or neutral emotional stimuli. Keywords: Emotional stimuliThe functional field of viewSOA Acknowledgements This work was supported by a grant to the first author from the Research Fellowships of the Japan Society for the Promotion of Science for Young Scientists and Grant no. 15330156 from the Japan Society for the Promotion of the Science to the second author. The authors thank Sachio Nakamizo for his helpful comments on an earlier draft of this manuscript. Notes 1The IAPS slide numbers were as follows: positive, 4608, 4660, 5470, 8030, 8170, 8179, 8185, 8186, 8200, and 8490; negative, 1050, 1525, 2811, 3215, 3500, 3550, 6021, 6022, 6300, 6312, 6350, and 9921; neutral, 2102, 2393, 2396, 2579, 2850, 5530, 5731, 7034, 7179, 7205, 7490, and 7710. An ANOVA was performed on the pleasant ratings, arousal ratings, and the number of bytes of the compressed image file sizes with one factor of emotion (positive, negative, and neutral). In the pleasant ratings, the main effect of emotion was significant, F(2, 33) = 560.63, MSE=0.13, p<.001. A post hoc multiple comparison test using Ryan's method indicated that the positive picture (M=7.28) was higher than the neutral (M=5.25) or the negative (M=2.54) picture (p<.05), and the neutral picture was higher than the negative picture (p<.05). In the arousal ratings, the main effect of emotion was significant, F(2, 33) = 257.17, MSE=0.19, p<.001. A post hoc multiple-comparison test using Ryan's method indicated that the positive (M=6.64) or the negative picture (M=6.47) was higher than the neutral (M=3.04) picture (p<.05). In the number of bytes of the compressed image file sizes, a main effect of emotion was significant, F(2, 33) = 8.74, MSE=3641.56, p<.001. A post hoc multiple comparison test using Ryan's method indicated that the negative picture (M=112.83) was lower than the positive (M=187.00) or the neutral (M=211.33) pictures (p<.05).
German genitive attributes are usually tagged as such in treebanks. However, it is well known that this information is not sufficient for determining the type of relation between head nouns and attributes, as genitive attributes can express many different semantic relations. Various linguistic classifications have been worked out, but to my knowledge, nobody has so far proposed to apply this linguistic knowledge to a corpus. The challenge here is to come up with a classification that is both easy to verify and sufficiently fine-grained. Using earlier linguistic approaches as guidelines, I propose in this paper a detailed annotation scheme for German genitive attributes based on readily identifiable noun features. First insights from its application to the Smultron Treebank show that it is easy to distinguish between the proposed classes and that my classification of genitive attributes can be related to a more general semantic annotation level.
Dysfunctional emotional processing affects social functioning in patients with schizophrenia. However, the relationship between emotional perception and response in social interaction has not been elucidated. Twenty-seven patients with schizophrenia and 27 normal controls performed a virtual reality social encounter task in which they introduced themselves to avatars expressing happy, neutral, or angry emotions while verbal response duration and onset time were measured and perception of emotional valence and arousal, and state anxiety were rated afterwards. Self-reported trait-affective scale scores and the Positive and Negative Syndrome Scale (PANSS) ratings were also obtained. Patient group significantly underestimated the valence and arousal of angry emotions expressed by an avatar. While valence and arousal ratings of happy avatars were comparable between groups, patient group reported significantly higher state anxiety in response to happy avatars. State anxiety ratings significantly decreased from encounters with neutral to happy avatars in normal controls while no significant decrease was observed in the patient group. The Social Anhedonia Scale and PANSS negative symptom subscale scores (blunted affect, emotional withdrawal, and passive/ apathetic social withdrawal items) were significantly correlated with state anxiety ratings of the encounters with happy avatars. These results suggest that patients with schizophrenia have interference with the experience of pleasure in social interactions which may be associated with negative symptoms.
In this paper we describe our participation at the EVALITA 2009 Con- stituency Parsing Task. We used the Berkeley Parser, obtaining the best F1, that is 78:73. This result corresponds to an increment of 15.85% with respect to the best result obtained at EVALITA 2007 by the Bikel's parser (F1 = 67:96). A further important advantage of the Berkeley parser is that it does not require any language adaptation in addition to the need of retraining it on the new treebank. For comparison, we also report the results obtained by the Bikel's parser on the 2009 treebank.
Alternative paths to linguistic annotation, such as those utilizing games or exploiting the web users, are becoming popular in recent times owing to their very high benefit-to-cost ratios. In this paper, however, we report a case study on POS annotation for Bangla and Hindi, where we observe that reliable linguistic annotation requires not only expert annotators, but also a great deal of supervision. For our hierarchical POS annotation scheme, we find that close supervision and training is necessary at every level of the hierarchy, or equivalently, complexity of the tagset. Nevertheless, an intelligent annotation tool can significantly accelerate the annotation process and increase the inter-annotator agreement for both expert and non-expert annotators. These findings lead us to believe that reliable annotation requiring deep linguistic knowledge (e.g., POS, chunking, Treebank, semantic role labeling) requires expertise and supervision. The focus, therefore, should be on design and development of appropriate annotation tools equipped with machine learning based predictive modules that can significantly boost the productivity of the annotators.
Most of text mining techniques are based on word and/or phrase analysis of the text. The statistical analysis of a term (word or phrase) frequency captures the importance of the term within a document. However, to achieve a more accurate analysis, the underlying mining technique should indicate terms that capture the semantics of the text from which the importance of a term in a sentence and in the document can be derived. Incorporating semantic features from the WordNet lexical database is one of many approaches that have been tried to improve the accuracy of text clustering techniques. A new semantic-based model that analyzes documents based on their meaning is introduced. The proposed model analyzes terms and their corresponding synonyms and/or hypernyms on the sentence and document levels. In this model, if two documents contain different words and these words are semantically related, the proposed model can measure the semantic-based similarity between the two documents. The similarity between documents relies on a new semantic-based similarity measure which is applied to the matching concepts between documents. Experiments using the proposed semantic-based model in text clustering are conducted. Experimental results demonstrate that the newly developed semantic-based model enhances the clustering quality of sets of documents substantially.
Automatic syllabification of words is challenging, not least because the syllable is not easy to define precisely. Consequently, no accepted standard algorithm for automatic syllabification exists. There are two broad approaches: rule-based and data-driven. The rule-based method effectively embodies some theoretical position regarding the syllable, whereas the data-driven paradigm tries to infer "new" syllabifications from examples assumed to be correctly syllabified already. This article compares the performance of several variants of the two basic approaches. Given the problems of definition, it is difficult to determine a correct syllabification in all cases and so to establish the quality of the "gold standard" corpus used either to evaluate quantitatively the output of an automatic algorithm or as the example-set on which data-driven methods crucially depend. Thus, we look for consensus in the entries in multiple lexical databases of pre-syllabified words. In this work, we have used two independent lexicons, and extracted from them the same 18,016 words with their corresponding (possibly different) syllabifications. We have also created a third lexicon corresponding to the 13,594 words that share the same syllabifications in these two sources. As well as two rule-based approaches (Hammond's and Fisher's implementation of Kahn's), three data-driven techniques are evaluated: a look-up procedure, an exemplar-based generalization technique, and syllabification by analogy (SbA). The results on the three databases show consistent and robust patterns. First, the data-driven techniques outperform the rule-based systems in word and juncture accuracies by a very significant margin but require training data and are slower. Second, syllabification in the pronunciation domain is easier than in the spelling domain. Finally, best results are consistently obtained with SbA.
This paper presents a structural statistical machine translation (SSMT) model to deal with the data sparseness problem that occurs as a result of the necessarily small corpus to translate Chinese into Taiwanese Sign Language (TSL). A parallel bilingual corpus was developed, and linguistic information from the Sinica Treebank is adopted for Chinese sentence analysis. The synchronous context free grammar (SCFG) was adopted to convert a Chinese structure to the corresponding TSL structure and then extract a translation memory which comprises the thematic relations between the grammar rules of both structures. In structural translation, the statistical MT (SMT) approach was used to align the thematic roles in the grammar rules and the translation memory provides the reference templates for TSL structure translation. Finally, the agreement information for TSL verbs was labeled for enriching the expressiveness of the translated TSL sequence. Several experiments were conducted to evaluate the translation performance and the communication effectiveness for the deaf. The evaluation results demonstrate that the proposed approach outperforms a baseline statistical MT system using the same small corpus, especially for the translation of long sentences.
The systematic presentation of collocations is increasingly recognized as a very useful addition to specialized reference works. However, few dictionaries or terminological databases actually include this kind of data. More surprisingly still, no method has been designed yet to allow efficient access to and retrieval of specific specialized collocations from electronic reference tools. This article presents two new search paths for accessing and extracting collocations from an English-French specialized lexical database. The paths have been designed according to two specific user-defined situations: (1) translation from L1 to L2; and (2) text production in L2. We exploit a formal semantic encoding of collocations based on Lexical Functions (LFs). LFs allow us to establish an equivalence relationship between collocations that convey the same meaning in different languages without having to link the collocations formally. They also allow us to extract sets of collocations associated with specific meanings.
We adapt a semantic role parser to the domain of goal-directed speech by creating an artificial treebank from an existing text tree-bank. We use a three-component model that includes distributional models from both target and source domains. We show that we improve the parser's performance on utterances collected from human-machine dialogues by training on the artificially created data without loss of performance on the text treebank.
Purpose The purpose of this paper is to investigate the characteristics of social network comments to give a broad overview to serve as a baseline for future research. Design/methodology/approach English comments from a representative sample of public MySpace profiles were examined with a collection of exploratory analyses, using automatic data processing, quantitative techniques and content analyses. Findings Comments were normally for general friendship maintenance and were typically short, with 95 per cent having 57 or fewer words. They contained a combination of standard spelling, apparently accidental mistakes, slang, sentence fragments, “typographic slang” and interjections. Several new creative spelling variants derived from previous forms of computer‐mediated communication have become extremely common, including u, ur,:), haha and lol. The vast majority of comments (97 per cent) contained at least one non‐standard language feature, suggesting that members almost universally recognise the informal nature of this kind of messaging. Research limitations/implications The investigation only covered MySpace and only analysed English comments. Practical implications MySpace comments should not be written in, or judged by, standard linguistic norms and may cause special problems for information retrieval. Originality/value This is the first large‐scale study of language in social network comments.