Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
OBJECTIVE: The neuropeptide oxytocin is implicated in social processing, and recent research has begun to explore how gender relates to the reported effects. This study examined the effects of oxytocin on social affective perception and learning. METHODS: Forty-seven male and female participants made judgments of faces during two different tasks, after being randomized to either double-blinded intranasal oxytocin or placebo. In the first task, "unseen" affective stimuli were presented in a continuous flash suppression paradigm, and participants evaluated faces paired with these stimuli on dimensions of competence, trustworthiness, and warmth. In the second task, participants learned affective associations between neutral faces and affective acts through a gossip learning procedure and later made affective ratings of the faces. RESULTS: In both tasks, we found that gender moderated the effect of oxytocin, such that male participants in the oxytocin condition rated faces more negatively, compared with placebo. The opposite pattern of findings emerged for female participants: they rated faces more positively in the oxytocin condition, compared with placebo. CONCLUSIONS: These findings contribute to a small but growing body of research demonstrating differential effects of oxytocin in men and women.
The objective of this paper is to provide an overview of the CDT annotation design with special emphasis on the modelling of the interface between the syntactic level and two other linguistic levels, viz. morphology and discourse. In connection with the description of NP annotation we present the fundamentals of how CDT is marked up with semantic relations in accordance with the dependency principles governing the annotation on the other levels of CDT. Specifically, focus will be on how Generative Lexicon (GL) theory has been incorporated into the unitary theoretical dependency framework of CDT. An annotation scheme for lexical semantics has been designed so as to account for the lexico-semantic structure of complex NPs, and the four GL qualia also appear in some of the CDT discourse relation labels as a description of parallel semantic relations at this level.
We present a novel toolkit that implements the long short-term memory (LSTM) neural network concept for language modeling. The main goal is to provide a software which is easy to use, and which allows fast training of standard recurrent and LSTM neural network language models. The toolkit obtains state-of-the-art performance on the standard Treebank corpus. To reduce the training time, BLAS and related libraries are supported, and it is possible to evaluate multiple word sequences in parallel. In addition, arbitrary word classes can be used to speed up the computation in case of large vocabulary sizes. Finally, the software allows easy integration with SRILM, and it supports direct decoding and rescoring of HTK lattices. The toolkit is available for download under an open source license.
Methylphenidate mainly enhances dopamine neurotransmission whereas 3,4-methylenedioxymethamphetamine (MDMA, "ecstasy") mainly enhances serotonin neurotransmission. However, both drugs also induce a weaker increase of cerebral noradrenaline exerting sympathomimetic properties. Dopaminergic psychostimulants are reported to increase sexual drive, while serotonergic drugs typically impair sexual arousal and functions. Additionally, serotonin has also been shown to modulate cognitive perception of romantic relationships. Whether methylphenidate or MDMA alter sexual arousal or cognitive appraisal of intimate relationships is not known. Thus, we evaluated effects of methylphenidate (40 mg) and MDMA (75 mg) on subjective sexual arousal by viewing erotic pictures and on perception of romantic relationships of unknown couples in a double-blind, randomized, placebo-controlled, crossover study in 30 healthy adults. Methylphenidate, but not MDMA, increased ratings of sexual arousal for explicit sexual stimuli. The participants also sought to increase the presentation time of implicit sexual stimuli by button press after methylphenidate treatment compared with placebo. Plasma levels of testosterone, estrogen, and progesterone were not associated with sexual arousal ratings. Neither MDMA nor methylphenidate altered appraisal of romantic relationships of others. The findings indicate that pharmacological stimulation of dopaminergic but not of serotonergic neurotransmission enhances sexual drive. Whether sexual perception is altered in subjects misusing methylphenidate e.g., for cognitive enhancement or as treatment for attention deficit hyperactivity disorder is of high interest and warrants further investigation.
It is well established that categorising the emotional content of facial expressions may differ depending on contextual information. Whether this malleability is observed in the auditory domain and in genuine emotion expressions is poorly explored. We examined the perception of authentic laughter and crying in the context of happy, neutral and sad facial expressions. Participants rated the vocalisations on separate unipolar scales of happiness and sadness and on arousal. Although they were instructed to focus exclusively on the vocalisations, consistent context effects were found: For both laughter and crying, emotion judgements were shifted towards the information expressed by the face. These modulations were independent of response latencies and were larger for more emotionally ambiguous vocalisations. No effects of context were found for arousal ratings. These findings suggest that the automatic encoding of contextual information during emotion perception generalises across modalities, to purely non-verbal vocalisations, and is not confined to acted expressions.
Measuring the similarity of words is important in accurately representing and comparing documents, and thus improves the results of many natural language processing (NLP) tasks. The NLP community has proposed various measurements based on WordNet, a lexical database that contains relationships between many pairs of words. Recently, a number of techniques have been proposed to address software engineering issues such as code search and fault localization that require understanding natural language documents, and a measure of word similarity could improve their results. However, WordNet only contains information about words senses in general-purpose conversation, which often differ from word senses in a software-engineering context, and the software-specific word similarity resources that have been developed rely on data sources containing only a limited range of words and word uses.
BACKGROUND: Patients with schizophrenia often experience problems regulating their emotions. Non-affected relatives show similar difficulties, although to a lesser extent, and the neural basis of such difficulties remains to be elucidated. In the current paper we investigated whether schizophrenia patients, non-affected siblings and healthy controls (HC) exhibit differences in brain activation during emotion regulation. METHODS: All subjects (n = 20 per group) performed an emotion regulation task while they were in an fMRI scanner. The task contained two experimental conditions for the down-regulation of emotions (reappraise and suppress), in which IAPS pictures were used to generate a negative affect. We also assessed whether the groups differed in emotion regulation strategies used in daily life by means of the emotion regulation questionnaire (ERQ). RESULTS: Though the overall negative affect was higher for patients as well as for siblings compared to HC for all conditions, all groups reported decreased negative affect after both regulation conditions. Nonetheless, neuroimaging results showed hypoactivation relative to HC in VLPFC, insula, middle temporal gyrus, caudate and thalamus for patients when reappraising negative pictures. In siblings, the same pattern was evident as in patients, but only in cortical areas. CONCLUSIONS: Given that all groups performed similarly on the emotion regulation task, but differed in overall negative affect ratings and brain activation, our findings suggest reduced levels of emotion regulation processing in neural circuits in patients with schizophrenia. Notably, this also holds for siblings, albeit to a lesser extent, indicating that it may be part and parcel of a vulnerability for psychosis.
After a period when the focus was essentially on mental architecture, the cognitive sciences are increasingly integrating the social dimension. The rise of a cognitive sociolinguistics is part of this trend. The article argues that this process requires a re-evaluation of some entrenched positions in linguistics: those that see linguistic norms as antithetical to a descriptive and variational linguistics. Once such a re-evaluation has taken place, however, the social recontextualization of cognition will enable linguistics (including sociolinguistics as an integral part), to eliminate the cracks in the foundations that were the result of suppressing the sociocultural underpinnings of linguistic facts. Structuralism, cognitivism and social constructionism introduced new and necessary distinctions, but in their strong forms they all turned into unnecessary divides. The article tries to show that an evolutionary account can reintegrate the opposed fragments into a whole picture that puts each of them in their ‘ecological position’ with respect to each other. Empirical usage facts should be seen in the context of operational norms in relation to which actual linguistic choices represent adaptations. Variational patterns should be seen in the context of structural categories without which there would be only ‘differences’ rather than variation. And emergence, individual choice, and flux should be seen in the context of the individual’s dependence on lineages of community practice sustained by collective norms.
International audience
OBJECTIVE: This study investigated how lexical effects account for word recognition in monolinguals versus bilinguals. DESIGN: Listener-specific error rate and familiarity rating of 200 NU-6 words were obtained. Lexical data (normative familiarity, frequency of occurrence, neighborhood density, and frequency of neighborhood competitors) for these words were obtained from the Hoosier mental lexicon. STUDY SAMPLE: Participants included 10 monolinguals and three groups of 10 bilinguals differing mainly in age of acquisition and length of schooling/working in English. RESULTS: Lexical effects were minimal for monolinguals' word recognition. Listener-specific familiarity rating correlated to error rate better than the Hoosier normative rating. Frequency of occurrence was the most significant lexical variable in accounting for bilinguals' measures and its effect was the greatest on bilinguals foreign born and educated. Age of English acquisition tended to affect familiarity rating, whereas length of schooling/working in English tended to affect error rate. CONCLUSIONS: Frequency of word occurrence significantly affects bilinguals' familiarity rating and error rate of the NU-6 words. Listener-specific familiarity rating should be obtained to best predict error rate on the test.
ABSTRACT. I show a consistent line of anti-Cartesian and anti-Kantian philosophies established by Herder, Humboldt, Hamann, and Vossler transformed the sign-character of human expressions into an expression of style by concentrating on the German term 'Art' as a sort of habitus. Especially Humboldt concentrates on the non-natural and non-functional character of language in defining it as a cultural product. Herder, Humboldt, Hamann, and Vossler recognize the insignificance of grammar as long as it is regarded as an abstract foundation for linguistics. Style is of utmost importance for all of them since they undermine the representation-theory by means of an open rhetoric using flexible rules and defining a sort of creative structural analysis. Die Art or the how of language gives linguistic expression the status of something spiritual and non-static. All four thinkers formulated their opposition philosophical abstractions in terms of a philosophy of life and attempted read signs as expressions of a historical and cultural reason.Keywords: Wilhelm von Humboldt; Gottfried Herder; Johan Georg Hamann; Karl Vossler; structuralist stylistics; representation theory1. IntroductionIn this article I show a consistent line of anti-Cartesian and anti-Kantian philosophies established by Herder, Humboldt, Hamann, and Vossler attempted transform the sign-character of human expressions into an expression of style by concentrating on the German term 'Art' as a sort of habitus. First, this is opposed Kant's view, which dismisses language in favor of purity of reason. However, it is also different from contemporary structuralist stylistics. In the twentieth century, structuralists concentrate on language as a cultural product, but they do so only in order find a kind of pure reason in language. The paradoxical character of this constellation will be made clear through an examination of style, habitus, and the German term Art. Charles Taylor held the main work of linguistic understanding is to see the place representation has in culture (1985: 291). Taylor, the skepticism of the above three philosophers of language (whom he calls the HHH) towards any autonomy of linguistic representation corresponds exactly the skepticism modem philosophers have towards (linguistic) norms. As a matter of fact, the essential critical part of the HHH-philosophy consists in undermining the representation-theory by means of an open rhetoric using flexible rules and defining a sort of creative structural analysis. On the other hand, the French linguist Oswald Ducrot is able point Humboldt's linguistics as an initial movement of a structuralist variation of stylistics relying on the theory of representation or sign theory when writing: For Humboldt (...) a language's organizational mode, even though arbitrary, is a means fulfill its representational function. It is, in a way, the style and the way any people choses express the most universal spiritual power. Ducrot concludes that nineteenth century linguistics had a concept of structure or a system (these two words are central for linguists of this epoque) (Ducrot 1973: 32).This is surprising if one considers Humboldt's dynamic innere Form has indeed been used as a weapon against static methods of French linguistic theories focusing on external structures (cf. Hutton 1998: 19). Most commonly, Humboldt's anti-sign linguistics has been opposed Saussure's structuralist linguistics (Trabant 2001) because inner form implies the word is not a copy of the object but a mental image produced by the speaker, which links language ways of thinking. The paradox is the innere Form remains a form; there is nothing in Humboldt suggesting a romantic unity of language and people producing mystical, intimate connections between both or even the formulation of a national soul as advanced by Fichte and Schelling later (cf. …
Manually tagged dependency treebank, analytical layer according to the PDT formalism adapted for Croatian
Although cognitive regulation of emotion has been extensively examined, there is a lack of studies assessing cognitive regulation in stressful achievement situations. This study used functional magnetic resonance imaging in 23 females and 20 males to investigate cognitive downregulation of negative, stressful sensations during a frequently used psychosocial stress task. Additionally, subjective responses, cognitive regulation strategies, salivary cortisol, and skin conductance response were assessed. Subjective response supported the experimental manipulation by showing higher anger and negative affect ratings after stress regulation than after the mere exposure to stress. On a neural level, right middle frontal gyrus (MFG) and right superior temporal gyrus (STG) were more strongly activated during regulation than nonregulation, whereas the hippocampus was less activated during regulation. Sex differences were evident: after regulation females expressed higher subjective stress ratings than males, and these ratings were associated with right hippocampal activation. In the nonregulation block, females showed greater activation of the left amygdala and the right STG during stress than males while males recruited the putamen more robustly in this condition. Thus, cognitive regulation of stressful achievement situations seems to induce additional stress, to recruit regions implicated in attention integration and working memory and to deactivate memory retrieval. Stress itself is associated with greater activation of limbic as well as attention areas in females than males. Additionally, activation of the memory system during cognitive regulation of stress is associated with greater perceived stress in females. Sex differences in cognitive regulation strategies merit further investigation that can guide sex sensitive interventions for stress-associated disorders.
Many event models are available for event extraction. These models have event information for agent, time and place. Effort to collect information for similar events to construct a lexical database is yet to be carried out. This research intends to achieve such objective by proposing an automated event extraction approach based on semantic role label (SRL) predicate argument structure (PAS). Events which are indicated by predicates with similar sense are gathered to collect event information for lexical database. This is a generic approach which does not need any event model. This paper has described and illustrated how SRL PAS is mapped to event information to create the lexical database.
Stanford Dependencies (SD) represent nowadays a de facto standard as far as dependency annotation is concerned. The goal of this paper is to explore pros and cons of different strategies for generating SD annotated Italian texts to enrich the existing Italian Stanford Dependency Treebank (ISDT). This is done by comparing the performance of a statistical parser (DeSR) trained on a simpler resource (the augmented version of the Merged Italian Dependency Treebank or MIDT+) and whose output was automatically converted to SD, with the results of the parser directly trained on ISDT. Experiments carried out to test reliability and effectiveness of the two strategies show that the performance of a parser trained on the reduced dependencies repertoire, whose output can be easily converted to SD, is slightly higher than the performance of a parser directly trained on ISDT. A non-negligible advantage of the first strategy for generating SD annotated texts is that semi-automatic extensions of the training resource are more easily and consistently carried out with respect to a reduced dependency tag set. Preliminary experiments carried out for generating the collapsed and propagated SD representation are also reported.
Previous work by Lin et al. (2011) demonstrated the effectiveness of using discourse relations for evaluating text coherence. However, their work was based on discourse relations annotated in accordance with the Penn Discourse Treebank (PDTB) (Prasad et al., 2008), which encodes only very shallow discourse structures; therefore, they cannot capture long-distance discourse dependencies. In this paper, we study the impact of deep discourse structures for the task of co-herence evaluation, using two approaches: (1) We compare a model with features derived from discourse relations in the style of Rhetorical Structure Theory (RST) (Mann and Thompson, 1988), which annotate the full hierarchical discourse structure, against our re-implementation of Lin et al.’s model; (2) We compare a model encoded using only shallow RST-style discourse relations, against the one encoded using the complete set of RST-style discourse relations. With an evaluation on two tasks, we show that deep discourse structures are truly useful for better dif-ferentiation of text coherence, and in general, RST-style encoding is more powerful than PDTB-style encoding in these settings. 1
The neuropeptide oxytocin enhances in-group favoritism and ethnocentrism in males. However, whether such effects also occur in women and extend to national symbols and companies/consumer products is unclear. In a between-subject, double-blind placebo controlled experiment we have investigated the effect of intranasal oxytocin on likeability and arousal ratings given by 51 adult Chinese males and females for pictures depicting people or national symbols/consumer products from both strong and weak in-groups (China and Taiwan) and corresponding out-groups (Japan and South Korea). To assess duration of treatment effects subjects were also re-tested after 1 week. Results showed that although oxytocin selectively increased the bias for overall liking for Chinese social stimuli and the national flag, it had no effect on the similar bias toward other Chinese cultural symbols, companies, and consumer products. This enhanced bias was maintained 1 week after treatment. No overall oxytocin effects were found for Taiwanese, Japanese, or South Korean pictures. Our findings show for the first time that oxytocin increases liking for a nation's society and flag in both men and women, but not that for other cultural symbols or companies/consumer products.
Studies comparing memory and future event simulation find that future events are more positive, and more often depend on life script events (e.g., culturally normative landmark events) than past events. Previous research does not address the link between this positivity bias and the life stage of college-age participants or their reliance on these scripted events. To examine this positivity bias, narratives of past and anticipated future events were elicited from participants aged 18-74 years, and were examined for reliance on the life script and valence ratings. Results showed that, across age groups, future events were rated as more positive than past events, and that life script events were common in the distant future. Notably, whereas younger adult age groups wrote primarily about their own life script events, older participants more commonly wrote about attending the life script events of significant others, such as children and grandchildren. These findings suggest that simulated future events play a valuable role in self-enhancement across the lifespan. Furthermore, the life script can be viewed as a useful search mechanism when one is missing the episodic details that are more available in memories; however, it is not the source of positivity bias for future events.
We present a user-centered approach for defining the dependency syntactic specification for a treebank. We show that by collecting information on syntactic interpretations from the future users of the treebank, we can model so far dependency-syntactically undefined syntactic structures in a way that corresponds to the users’ intuition. By consulting the users at the grammar definition phase we aim at better usage of the treebank in the future. We focus on two complex syntactic phenomena: elliptical comparative clauses and participial NPs or NPs with a verb-derived noun as their head. We show how the phenomena can be interpreted in several ways and ask for the users’ intuitive way of modeling them. The results aid in constructing the syntactic specification for the treebank.
Social media applications such as Twitter provide a powerful medium through which users can communicate their observations with friends and with the world at large. We have witnessed live reporting of many events, from soccer games in Johannesburg to revolutions in Cairo and Tunis, and these reports have in many ways rivaled the content provided by the official media. Tapping into this valuable resource is a challenge, due to the heterogeneity and noise inherent in realtime text, diversity of languages, and fast-evolving linguistic norms. In this paper we seek to analyze a tweet stream to automatically discover points in time when an important event happens, and to classify such events based on the type of the sentiments they evoke, using only non-textual features of the tweeting pattern. This results not only in a robust way of analyzing tweet streams independent of the languages used; it also provides insights about how users behave on social media websites. For example, we observe that users often react to an exciting external event by decreasing the volume of communication with other users. We explain this effect through a model of how users switch between producing information or sentiments and sharing others’ news or sentiments. We develop and evaluate our models and algorithms using several Twitter data sets, focusing in particular on the tweets sent during the soccer World Cup of 2010. This data set has the feature that the underlying ground truth is welldefined and known whereby goals serve as events.
Abstract This study was conducted to understand the relationship between familiarity and cross‐cultural acceptance for an ethnic sweet treat ( Y ackwa; K orean traditional cookie) by K orean, J apanese and F rench consumers. Descriptive analysis and consumer testing were performed on six Y ackwa samples. Overall, the samples received favorable responses from the foreign consumers. K orean consumers liked samples with a soft and cohesive texture, whereas J apanese and F rench consumers liked flaky and crispy texture. French consumers rated stronger sweetness to be more appropriate for Y ackwa compared to K orean and J apanese consumers. Texture liking was strongly correlated with familiarity rating in all three countries, indicating that the consumers' previous experience with similar products might affect their preference for certain textural attributes. Familiarity was correlated with all hedonic ratings by K orean consumers, who are most familiar with Y ackwa, but with overall and texture liking by J apanese consumers and flavor and texture liking by French consumers. These results suggest that familiarity partly contributes to a foreign consumers' hedonic rating. Practical Applications Globalization and cultural diversity have increased interest in ethnic foods. This trend is motivating food industries to expand into the ethnic food market sector. In this study, the sensory attributes and the cross‐cultural acceptability of Y ackwa ( K orean traditional cookie) were evaluated and the potential role of familiarity in determining consumer acceptance was measured. The outcome of this study will help food exporters, R&D scientists and food marketers in ethnic food market to optimize an ethnic food for other cultural communities by educating them to consider familiarity as an important factor for product development and promotion.
BACKGROUND: Depression is frequently characterized by patterns of inflexible, maladaptive, and ruminative thinking styles, which are thought to result from a combination of decreased attentional control, decreased executive functioning, and increased negative affect. Cognitive Control Training (CCT) uses computer-based behavioral exercises with the aim of strengthening cognitive and emotional functions. A previous study found that severely depressed participants who received CCT exhibited reduced negative affect and rumination as well as improved concentration. AIMS: The present study aimed to extend this line of research by employing a more stringent control group and testing the efficacy of three sessions of CCT over a 2-week period in a community population with depressed mood. METHOD: Forty-eight participants with high Beck Depression Inventory (BDI-II) scores were randomized to CCT or a comparison condition (Peripheral Vision Training; PVT). RESULTS: Significant large effect sizes favoring CCT over PVT were found on the BDI-II (d = 0.73, p <.05) indicating CCT was effective in reducing negative mood. Additionally, correlations showed significant relationships between CCT performance (indicating ability to focus attention on CCT) and state affect ratings. CONCLUSIONS: Our results suggest that CCT is effective in altering depressed mood, although it may be specific to select mood dimensions.
We present a framework for identifying the most representative sentence patterns from semantically and syntactically-annotated corpora via a Semantic Frame Generation (SFG). One of the difficulties to find out similar concepts from a text is because of the variations in linguistic expressions. SFG uses linguistic units as backbones to generate the most prominent patterns from various Chinese DE phrases.
Recursive neural models have achieved promising results in many natural language processing tasks. The main difference among these models lies in the composition function, i.e., how to obtain the vector representation for a phrase or sentence using the representations of words it contains. This paper introduces a novel Adaptive Multi-Compositionality (AdaMC) layer to recursive neural models. The basic idea is to use more than one composition functions and adaptively select them depending on the input vectors. We present a general framework to model each semantic composition as a distribution over these composition functions. The composition functions and parameters used for adaptive selection are learned jointly from data. We integrate AdaMC into existing recursive neural models and conduct extensive experiments on the Stanford Sentiment Treebank. The results illustrate that AdaMC significantly outperforms state-of-the-art sentiment classification methods. It helps push the best accuracy of sentence-level negative/positive classification from 85.4% up to 88.5%.
This paper presents the first results on parsing the Penn Parsed Corpus of Modern British English (PPCMBE), a millionword historical treebank with an annotation style similar to that of the Penn Treebank (PTB). We describe key features of the PPCMBE annotation style that differ from the PTB, and present some experiments with tree transformations to better compare the results to the PTB. First steps in parser analysis focus on problematic structures created by the parser.
The article focuses on the hypothesis that the structural complexity of languages is variable and historically changeable. By means of a quantitative statistical analysis of naturalistic corpus data, the question is raised as to what role language contact and adult second language acquisition play in the simplification and complexification of language varieties. The results confirm that there is a significant correlation between intensity of contact and linguistic complexity, while at the same time showing that there is a need to consider other social factors, and, in particular, the attitude of a speech community toward linguistic norms. *
The extinction of conditioned fear depends on an efficient interplay between the amygdala and the medial prefrontal cortex (mPFC). In rats, high-frequency electrical mPFC stimulation has been shown to improve extinction by means of a reduction of amygdala activity. However, so far it is unclear whether stimulation of homologues regions in humans might have similar beneficial effects. Healthy volunteers received one session of either active or sham repetitive transcranial magnetic stimulation (rTMS) covering the mPFC while undergoing a 2-day fear conditioning and extinction paradigm. Repetitive TMS was applied offline after fear acquisition in which one of two faces (CS+ but not CS-) was associated with an aversive scream (UCS). Immediate extinction learning (day 1) and extinction recall (day 2) were conducted without UCS delivery. Conditioned responses (CR) were assessed in a multimodal approach using fear-potentiated startle (FPS), skin conductance responses (SCR), functional near-infrared spectroscopy (fNIRS), and self-report scales. Consistent with the hypothesis of a modulated processing of conditioned fear after high-frequency rTMS, the active group showed a reduced CS+/CS- discrimination during extinction learning as evident in FPS as well as in SCR and arousal ratings. FPS responses to CS+ further showed a linear decrement throughout both extinction sessions. This study describes the first experimental approach of influencing conditioned fear by using rTMS and can thus be a basis for future studies investigating a complementation of mPFC stimulation to cognitive behavioral therapy (CBT).
Dictionaries are designed as huge texts made up of a collection of much smaller texts, i.e. lexicographic articles. To put it differently, dictionaries are two-dimensional textual models of natural language lexicons. Lexicographers, however, are well aware of the fact that their task is to account for a truly multidimensional entity: a gigantic graph of lexical units connected by various paradigmatic and syntagmatic relations. The most significant advance that computer science will bring to the future of lexicography is therefore not the ability to better store, search and manipulate textual lexicographic data; it will be to allow lexicographers to bypass the text as a formal representation of lexicons and to directly work on lexical networks. Such networks are more suitable to the lexicographic endeavour because they are better formal metaphors of the “natural” structure we are trying to account for. This paper presents lexical systems as graph models of lexicons and introduces the corresponding lexicography of virtual dictionaries. It is based on extensive lexicographic work that is being conducted on the French Lexical Network, a lexical database built according to theoretical and methodological principles borrowed from Explanatory Combinatorial Lexicology.
In pursuing machine understanding of human language, highly accurate syntactic analysis is a crucial step. In this work, we focus on dependency grammar, which models syntax by encoding transparent predicate-argument structures. Recent advances in dependency parsing have shown that employing higher-order subtree structures in graph-based parsers can substantially improve the parsing accuracy. However, the inefficiency of this approach increases with the order of the subtrees. This work explores a new reranking approach for dependency parsing that can utilize complex subtree representations by applying efficient subtree selection methods. We demonstrate the effectiveness of the approach in experiments conducted on the Penn Treebank and the Chinese Treebank. Our system achieves the best performance among known supervised systems evaluated on these datasets, improving the baseline accuracy from 91.88% to 93.42% for English, and from 87.39% to 89.25% for Chinese.
Many languages, including Modern Stan-dard Arabic (MSA), insert resumptive pro-nouns in relative clauses, whereas many others, such as English, do not, using empty categories instead. This discrep-ancy is a source of difficulty when trans-lating between these languages because there are words in one language that cor-respond to empty categories in the other, and these words must either be inserted or deleted—depending on translation di-rection. In this paper, we first examine challenges presented by resumptive pro-nouns in MSA-English translations and re-view resumptive pronoun translations gen-erated by a popular online MSA-English MT engine. We then present what is, to the best of our knowledge, the first system for automatic identification of resumptive pronouns. The system achieves 91.9 F1 and 77.8 F1 on Arabic Treebank data when using gold standard parses and automatic parses, respectively. 1
Discourse relations bind smaller linguistic elements into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked sentences. A more subtle challenge is that it is not enough to represent the meaning of each sentence of a discourse relation, because the relation may depend on links between lower-level elements, such as entity mentions. Our solution computes distributional meaning representations by composition up the syntactic parse tree. A key difference from previous work on compositional distributional semantics is that we also compute representations for entity mentions, using a novel downward compositional pass. Discourse relations are predicted not only from the distributional representations of the sentences, but also of their coreferent entity mentions. The resulting system obtains substantial improvements over the previous state-of-the-art in predicting implicit discourse relations in the Penn Discourse Treebank.
In this paper we present a system for experimenting with combinations of dependency parsers. The system supports initial training of different parsing models, creation of parsebank(s) with these models, and different strategies for the construction of ensemble models aimed at improving the output of the individual models by voting. The system employs two algorithms for construction of dependency trees from several parses of the same sentence and several ways for ranking of the arcs in the resulting trees. We have performed experiments with state-of-the-art dependency parsers including MaltParser (Nivre et al., 2006), MSTParser (McDonald, 2006), TurboParser (Martins et al., 2010), and MATEParser (Bohnet, 2010), on the data from the Bulgarian treebank – BulTreeBank. Our best result from these experiments is slightly better then the best result reported in the literature for this language (Martins et al.,
Incorporating knowledge for training a parser has been shown to remedy the weaknesses of probabilistic context-free grammar. Previous parsing systems have exploited content words semantic resource and word-formation knowledge. However, they are limited in that they do not take into account conjunction category refinement, which stands out to be helpful in predicting the syntactic structure and syntactic label in Chinese. We define a conjunction taxonomy representing intrinsic syntactic constraints, and show that refined categories in the taxonomy for conjunctions contribute to improved parsing performance. The taxonomy is used to supervise the splitting of these refined tags, and the automatic hierarchical state-split approach is employ to compensate the limitation in the scope and refinement degree of the taxonomy. The experiments are carried out on Penn Chinese Treebank, which show that our method can improve parsing performance significantly.
Prior research suggested the possibility of establishing systematic linkages between some intrinsic features of a presupposition and textual and pragmatic functions that it can carry out with greater probability. This study aims, firstly, to provide an organic view of semantics of presupposition triggers, thanks to a lexical database comprising 19,500 entries. Secondly, the database was used to investigate a corpus of chat conversations including about 200,000 tokens. The results show that triggers occur mainly as non-informative, maintaining an information already known by all participants of the communication; but, depending on their different features, some of them are systematically associated to a function of anaphora and textual cohesion; others to strengthen social conventions and stereotypes. The informative function, although in a minority proportion, is quantitatively significant only in correspondence to a single class of presupposition triggers.
Over the past 15 years, one of the basic functions of the Linguistic Landscape (LL) has been to measure and assess the visibility of multiple languages in public spaces. In research carried out around the world, English has recurrently appeared as one of the most important features of the LL, though its classification, and the extent to which its presence renders a space multilingual, remains controversial. In the light of Seargeant’s (2011) and Yano’s (2011) recent claims that English may be considered a feature of the local landscape rather than a language in its own right, this paper explores code classification in the LL of Toulouse. It challenges the concept of ‘global’ languages (Crystal, 1997), and illustrates the ways in which various codes become affected, consumed, and enveloped by the dominant language, French. Not only is this the case for English, but also the regional tongue, Occitan, whose visibility, despite appearing autonomous, may be more accurately considered a part of the French landscape. This illustrates the shortcomings in classifying languages according to fixed, subjective (and often ad hoc) criteria, because even the presentation of so-called ‘foreign’ languages in Toulouse is influenced by, and subject to, local linguistic norms. The implications of this are profound, as we see that French is the driving force behind every language in the city. Moreover, this calls into question the very definition of multilingualism, as LL authors exhibit a preference for the national code even when writing in others. As such, this research will have a meaningful impact on language coding in the LL, and contribute to the advancement of the methodologies by which we measure and analyse multilingualism, language beliefs, and language practices in public spaces.
The French Lexical Network (fr-LN) is a global model of the French lexicon presently under construction. The fr-LN accounts for lexical knowledge as a lexical network structured by paradigmatic and syntagmatic relations holding between lexical units. This paper describes how morphological knowledge is presently being introduced into the fr-LN through the implementation and lexicographic exploitation of a dynamic morphological model. Section 1 presents theoretical and practical justifications for the approach which we believe allows for a cognitively sound description of morphological data within semantically-oriented lexical databases. Section 2 gives an overview of the structure of the dynamic morphological model, which is constructed through two complementary processes: a Morphological Process--section 3--and a Lexicographic Process--section 4.
In the context of processing Bengali words through a computer, there may arise several issues that are directly linked with surface structure of words. These issues may create problems in manual and computer-based counting of number of words in a corpus. They can also create problems in morphological processing of words. These issues come up because there is hardly any consistency in orthographic representation of words in written Bengali texts. The high irregularities in writing of inflected words, proper names, adjectival forms, adverbial forms, compound words, reduplicated words, onomatopoeic words, hyphenated words, etc. present a daunting task before an investigator in normalizing the surface forms of words for generating a lexical database as well as developing a word processing system for the works of language technology.
In this study, a dictionary-based method is used to extract expressive concepts from documents. So far, there have been many studies concerning concept mining in English, but this area of study for Turkish, an agglutinative language, is still immature. We used dictionary instead of WordNet, a lexical database grouping words into synsets that is widely used for concept extraction. The dictionaries are rarely used in the domain of concept mining, but taking into account that dictionary entries have synonyms, hypernyms, hyponyms and other relationships in their meaning texts, the success rate has been high for determining concepts. This concept extraction method is implemented on documents, that are collected from different corpora.
The purpose of our work is to explore the possibility of using sentence diagrams produced by schoolchildren as training data for automatic syntactic analysis. We have implemented a sentence diagram editor that schoolchildren can use to practice morphology and syntax. We collect their diagrams, combine them into a single diagram for each sentence and transform them into a form suitable for training a particular syntactic parser. In this study, the object language is Czech, where sentence diagrams are part of elementary school curriculum, and the target format is the annotation scheme of the Prague Dependency Treebank. We mainly focus on the evaluation of individual diagrams and on their combination into a merged better version.
We suggest a new annotation scheme for unlexicalized PCFGs that is inspired by formal language theory and only depends on the structure of the parse trees. We evaluate this scheme on the TüBa-D/Z treebank w.r.t. several metrics and show that it improves both parsing accuracy and parsing speed considerably. We also show that our strategy can be fruitfully com-bined with known ones like parent annota-tion to achieve accuracies of over 90 % la-beled F1 and leaf-ancestor score. Despite increasing the size of the grammar, our annotation allows for parsing more than twice as fast as the PCFG baseline. 1
This paper reports on a corpus-based quantitative study of the use of nominalizations across China English and British English in two comparable media corpora. In contrast to previous corpus-based studies of nominalizations, we start by using a syntactic approach and proceed with some methodological innovations incorporating large lexical databases and syntactically annotated corpora. The data show that there are significant differences in the use of nominalizations across these two English varieties. It is hoped that this research will offer useful insights on variations in nominalization across different English varieties and also on the understanding of the two English varieties in question. 1
This study presents a comparison of the themes associated with the phenomena of birth and death in Pakistani and British English fictions i.e. PEF and BEF in the vast framework of several varieties of Englishes around the world. As culture of Pakistan differs entirely from that of Britain, so are the themes which are associated with birth and death in their fictions. Different words have different associative meaning in different societies (e.g. meaning and usage of ‘dear’ or ‘clever’ is quite different in Pakistan and in England. Truly representative corpus of English fiction, comprising various genres from both the varieties of English has been analyzed and processed through AntConc 3.2.2w (windows) 2008. The study has investigated PEF as well as BEF thoroughly on the basis of Dixon’s Semantic approach to English grammar (2005) and elaborated that the adjectives used with birth and death in both fictions are entirely different besides the universally common psychological incidents of birth and death in all human beings. The study reveals that the rituals and customs associated with birth and death are entirely different in BEF and PEF. It establishes that Pakistani variety of English uses entirely distinctive linguistic norms as compared to British variety of English. Keywords: PEF, BEF, themes, birth, death, associative meaning.
This paper showcases development on a digital learning environment built for users interested in ancient languages. The environment adopts a two-pronged approach: first, it assesses user progress during the language learning phase by using exercises such as treebanking and alignment; then, it offers users the opportunity to contribute their own annotated data. This paper provides a description of user testing conducted by the eLearning team in order to understand how these methods affect user engagement and persistence. The results of user testing showed higher user engagement and retention rates when compared to a more traditional quizzing method.
We propose Diverse Embedding Neural Network (DENN), a novel architecture for language models (LMs). A DENNLM projects the input word history vector onto multiple diverse low-dimensional sub-spaces instead of a single higher-dimensional sub-space as in conventional feed-forward neural network LMs. We encourage these sub-spaces to be diverse during network training through an augmented loss function. Our language modeling experiments on the Penn Treebank data set show the performance benefit of using a DENNLM.