Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Online content analysis employs algorithmic methods to identify entities in unstructured text. Both machine learning and knowledge-base approaches lie at the foundation of contemporary named entities extraction systems. However, the progress in deploying these approaches on web-scale has been been hampered by the computational cost of NLP over massive text corpora. We present SpeedRead (SR), a named entity recognition pipeline that runs at least 10 times faster than Stanford NLP pipeline. This pipeline consists of a high performance Penn Treebank- compliant tokenizer, close to state-of-art part-of-speech (POS) tagger and knowledge-based named entity recognizer.
This article considers dictionaries as lexical information / knowledge sources to be derived from a deeper, underlying, lexical database. These dictionary-tokens or -instantiations are inter alia specified by the users' needs. As a case in point of such a derivation meeting the needs of a multilingual society, a bidirectional bilingual learner dictionary is presented. Specific tools, such as editors with reversal function, and models, such as the hub-and-spoke model, are discussed as means to function within the lexicographical infrastructure of a multilingual society.
El objetivo del trabajo consiste en reutilizar el Treebank de dependencias EPEC-DEP (BDT) para construir el gold standard de la sintaxis superficial del euskera. El paso basico consiste en el estudio comparativo de los dos formalismos aplicados sobre el mismo corpus: el formalismo de la Gramatica de Restricciones (Constraint Grammar, CG) y la Gramatica de Dependencias (Dependency Grammar, DP). Como resultado de dicho estudio hemos establecido los criterios linguisticos necesarios para derivar la funciones sintacticas en estilo CG. Dichos criterios han sido implementados y evaluados, asi en el 75% de los casos somos capaces de derivar automaticamente las funciones sintacticas para construir el gold standard.
Four different patterns of biased ratings of facial expressions of emotions have been found in socially anxious participants: higher negative ratings of (1) negative, (2) neutral, and (3) positive facial expressions than nonanxious controls. As a fourth pattern, some studies have found no group differences in ratings of facial expressions of emotion. However, these studies usually employed valence and arousal ratings that arguably may be less able to reflect processing of social information. We examined the relationship between social anxiety and face ratings for perceived trustworthiness given that trustworthiness is an inherently socially relevant construct. Improving on earlier analytical strategies, we evaluated the four previously found result patterns using a Bayesian approach. Ninety-eight undergraduates rated 198 face stimuli on perceived trustworthiness. Subsequently, participants completed social anxiety questionnaires to assess the severity of social fears. Bayesian modeling indicated that the probability that social anxiety did not influence judgments of trustworthiness had at least three times more empirical support in our sample than assuming any kind of negative interpretation bias in social anxiety. We concluded that the deviant interpretation of facial trustworthiness is not a relevant aspect in social anxiety.
Part-of-speech (POS) tagging is a fundamental task in natural language processing (NLP). It provides useful information for many other NLP tasks, including word sense disambiguation, text chunking, named entity recognition, syntactic parsing, semantic role labeling, and semantic parsing. In this paper, we present a new method for Vietnamese POS tagging using dual decomposition. We show how dual decomposition can be used to integrate a word-based model and a syllable-based model to yield a more powerful model for tagging Vietnamese sentences. We also describe experiments on the Viet Treebank corpus, a large annotated corpus for Vietnamese POS tagging. Experimental results show that our model using dual decomposition outperforms both word-based and syllable-based models.
We present a classification model that predicts the presence or omission of a lexical connective between two clauses, based upon linguistic features of the clauses and the type of discourse relation holding between them.The model is trained on a set of high frequency relations extracted from the Penn Discourse Treebank and achieves an accuracy of 86.6%.Analysis of the results reveals that the most informative features relate to the discourse dependencies between sequences of coherence relations in the text.We also present results of an experiment that provides insight into the nature and difficulty of the task.
This paper investigates the impact on French dependency parsing of lexical generalization methods beyond lemmatization and morphological analysis. A distributional thesaurus is created from a large text corpus and used for distributional clustering and WordNet automatic sense ranking. The standard approach for lexical generalization in parsing is to map a word to a single generalized class, either replacing the word with the class or adding a new feature for the class. We use a richer framework that allows for probabilistic generalization, with a word represented as a probability distribution over a space of generalized classes: lemmas, clusters, or synsets. Probabilistic lexical information is introduced into parser feature vectors by modifying the weights of lexical features. We obtain improvements in parsing accuracy with some lexical generalization configurations in experiments run on the French Treebank and two out-of-domain treebanks, with slightly better performance for the probabilistic lexical generalization approach compared to the standard single-mapping approach. 1
The state is the primary provider of education in Singapore. The rise of the global knowledge economy, however, has generated an education industry that is worth about USD2.2 trillion per year globally. The Government of Singapore therefore revamped its higher-education sector so that Higher Private Education Organisations (HPEOs) can offer more flexible transnational education to attract fee-paying, overseas students and increase the education industry’s contribution to the national GDP. The private-education industry is currently facing competition from more established markets such as the United States, United Kingdom and Australia. Locally, HPEOs need to upgrade their service standard in order to meet the stringent registration requirements imposed by Singapore’s regulators and, at the same time, compete with each other in a high-cost environment. HPEOs therefore need to gain a better understanding of the relations among variables such as service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth. A better understanding will allow HPEOs to improve their marketing efforts in order to obtain competitive advantages. Prior research only focused on the correlation of variables and failed to take into account the inter-relationship among two or more variables. The measurement of these variables is often dependent on geographical factors, types of services and types of stakeholders. It is therefore the objective of this study to examine the measurement and relationships among service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth in the context of HPEOs in Singapore. A survey involving 554 participants from local HPEOs was conducted over a period of three years. Analysis of the data shows that attributes such as image rating and service quality are unique and HPEOs have to customize these attributes to meet the needs and wants of their students. This study found that service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth are all positively correlated. It was also found that some factors (e.g., overall satisfaction and repurchase intention) are mediators of the relationships between other variables.
Although disgust propensity (DP) has been implicated in the development of some anxiety disorders, the mechanism that may account for this association has not been fully elucidated. The present study examined the extent to which the potentiation of learned aversion might be one such mechanism. Participants (n = 103) were randomized to one of two evaluative conditioning (EC) paradigms consisting of 12 reinforced conditioned stimulus (CS+) pairings of the word "part" (condition one) or "some" (condition two) with 12 aversive unconditioned stimulus (US) images, and 12 pairings of the CS- word "cylinder" with 12 neutral images. Participants then completed measures of DP and trait anxiety and provided subjective affective ratings for the aversive US. The findings revealed that participants experienced significantly more disgust, anxiety, anger, sadness, and less happiness toward their respective CS+. In contrast, participants experienced significantly more happiness toward the CS-. Examination of the magnitude of evaluative change to the CS+ revealed the strongest effect for disgust. DP, but not trait anxiety, also predicted a greater increase in disgust, anger, and anxiety in response to the CS+ relative to the CS-. Furthermore, the association between DP and greater disgust, anger, and anxiety in response to the CS+ was mediated by more intense negative affective responding to the US among those higher in DP. The implication of these findings for better understanding how DP may confer risk for anxiety-related psychopathology is discussed.
The aim of the work is to profit the existing dependency Treebank EPEC-DEP (BDT) in order to build the gold standard for the surface syntax of Basque. As basic step, we make a comparative study of both formalisms, the Constraint Grammar formalism (CG) and the Dependency Grammar (DP) that have been applied on the corpus. As a result, we establish some criteria that will serve us to derive automatically the CG style syntactic function tags. Those criteria were implemented and evaluated; as a result, in the 75 % of the cases we are able to derive the CG style syntactic function tags for building the gold standard.
Extracting the users expected information from a large text collection based on some query is the aim of a Information Retrieval (IR) system. Now a days Assamese Digital documents are increasing at a huge rate and to collect the information efficiently from them we are in need of an Assamese IR system for retrieving documents. Comparing query and document term on lexical level and the shorter length query implies the problem like word mismatch. Adding additional term with user's query means expanding the query can improve the IR system's performance. The electronic lexical database, WordNet can help identifying the synonymous expressions and linguistic entities that are semantically similar with the input query. Here we present an Assamese IR system based on vector space model and show our result by considering the query vector as the original user's query and the query is extended using the Assamese WordNet synsets.
This article uses semi-supervised Expectation Maximization (EM) to learn lexico-syntactic dependencies, i.e. associations between words and the structures that occur with them. Due to Zipfian distributions in language, such dependencies are extremely sparse in labelled data, and unlabelled data are the only source for learning them. Specifically, we learn sparse lexical parameters of a generative parsing model (a Probabilistic Context-Free Grammar, PCFG) that is initially estimated over the Penn Treebank. Our lexical parameters are similar to supertags—they are fine-grained, and encode complex structural information at the pre-terminal level. Our goal is to use unlabelled data to learn these for words that are rare or unseen in the labelled data. We get large error reductions (up to 17.5%) in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies, resulting in a statistically significant improvement in labelled bracketing score of the treebank PCFG. Our semi-supervised method incorporates structural and lexical priors from the labelled data to guide estimation from unlabelled data, and is the first successful use of semi-supervised EM to improve a generative structured model already trained over large labelled data. The method scales well to larger amounts of unlabelled data, and also gives substantial error reductions (up to 11.5%) for models trained on smaller amounts of labelled data, making it relevant to low-resource languages with small treebanks as well.
The Digital World encounters rapid development nowadays, especially through the proliferation of social media in Indonesia. Twitter has become one of social media with expanded users within every sectors of society. There are so many part both individual as well as organization/enterprise which utilize twitter as tool for communication, business, customer relation, and other activities. Through the twitter's ever-expanding users with those particular purposes, the precise method to effectively and efficiently analyzing opinion-contained sentences become crucially needed. Therefore this research made for method analyzing through lexical based and model based approaches by machine learning to classify opinion-contained tweets using those 2 methods. The tested machine learning method are Support Vector Machine (SVM), Maximum Entropy (ME), Multinomial Naive Bayes (MNB), and k-Nearest Neighbor (k-NN). Based on the test outcome, lexical based approach highly depended on lexical database which became opinion classification matrix. Whilst machine learning approach can produce better accuracy due to its capability in new training data modeling based on outcome model. However, machine learning model based approach depends on various factors in analyzing sentiment.
Abstract In this article we present some statistical data on the distribution of parts of speech and dependency relations in a large manually annotated Hungarian Treebank, the Szeged Dependency Treebank. We hypothesize that the domain of the text influences the distribution of the above elements, thus we pay special attention to differences between domains. We present the characteristic rank-frequency distributions of parts of speech and dependency relations in Hungarian and analyse the domain similarities and differences among sub-corpora as regards the above distributions. Our results reveal that the computer and newspaper texts are most similar to each other while the domains literature and compositions also exhibit some similarities. On the other hand, the business news and the law sub-corpora are unique, both having their own characteristics.
Body image disturbances are core symptoms of eating disorders (EDs). Recent evidence suggests that changes in body image may occur prior to ED onset and are not restricted to in-vivo exposure (e.g. mirror image), but also evident during presentation of abstract cues such as body shape and weight-related words. In the present study startle modulation, heart rate and subjective evaluations were examined during reading of body words and neutral words in 41 student female volunteers screened for risk of EDs. The aim was to determine if responses to body words are attributable to a general negativity bias regardless of ED risk or if activated, ED relevant negative body schemas facilitate priming of defensive responses. Heart rate and word ratings differed between body words and neutral words in the whole female sample, supporting a general processing bias for body weight and shape-related concepts in young women regardless of ED risk. Startle modulation was specifically related to eating disorder symptoms, as was indicated by significant positive correlations with self-reported body dissatisfaction. These results emphasize the relevance of examining body schema representations as a function of ED risk across different levels of responding. Peripheral-physiological measures such as the startle reflex could possibly be used as predictors of females' risk for developing EDs in the future.
Syntactic parsing is an important technique in the natural language processing, yet Latvian is still lacking an efficient general coverage syntax parser. This paper reports on the first experiments on statistical syntactic parsing for Latvian — a highly inflective Indo-European language with a relatively free word order. We have induced a statistical parser from a small, non-balanced Latvian Treebank using the MaltParser toolkit and measured the unlabeled attachment score (UAS). As MaltParser is based on the dependency grammar approach, we have also developed a convertor from the hybrid dependency-based annotation model used in the Latvian Treebank to the pure dependency annotation model. We have obtained a promising 74.63 % UAS in 10-fold cross-validation using only ~2500 sentences. The results revealed that best results can be achieved using non-projective stack parsing algorithm with lazy arc adding strategy, but comparably good results can be achieved using projective parsing algorithms combined with appropriate projectiviziation preprocessing.
This chapter deals with the main methodological issues underlying the building of the SciE-Lex lexical database and discusses and justifies the information included. SciE-Lex was initially conceived as a response to the lack of reference tools that can help scientists write scientific papers in phraseologically competent and native-like English. While there are a number of specialised dictionaries that include specific terminological information, there is a shortage of writing aids that provide information about the use of non-technical terms in scientific genres. SciE-Lex aims at filling this gap by focusing on the description of general terms in scientific English. This article describes the two stages in the building of the database, the first one including morphosyntactic and collocational information, and the second one focusing on phraseological information.
Turkish is an agglutinative language with rich morphology-syntax interactions. As an extension of this property, the Turkish Treebank is designed to represent sublexical dependencies, which brings extra challenges to parsing raw text. In this work, we use a joint POS tagging and parsing approach to parse Turkish raw text, and we show it outperforms a pipeline approach. Then we experiment with incorporating morphological feature prediction into the joint system. Our results show statistically significant improvements with the joint systems and achieve the state-ofthe-art accuracy for Turkish dependency parsing.
Exploiting data from a parallel treebank recently developed for Italian, English and French, the paper discusses issues related to the development of a dependency-based alignment system.We focus on the alignment of linguistic expressions and constructions which are structurally different in the languages that have to be aligned, and on how to deal with them using dependency rather than constituency.In order to analyze in particular the shifts related to syntactic structure, we present a selection of cases where a dependencybased and a constituency-based representation has been applied and compared.
The ways in which literacy in English is taught in school generally subscribe to and perpetuate the notion of a homogenous, unvaried set of writing conventions associated with the language they represent, especially in relation to spelling and punctuation as well as grammar. Such teaching also perpetuates the myth that there is one correct way of language use which is fixed and invariant, and that any deviation is at best incorrect or illiterate and at worst, a threat to social stability. It is also very clear that the linguistic norms associated with standard English are predicated upon and replicate white, cultural hegemony. Yet, at the same time, there are plenty of literary and creative works written by authors from all kinds of different cultural, ethnic and linguistic backgrounds, including canonical ones, where spelling and punctuation are varied and championed as a sign of creativity. In the world beyond school, pupils are also surrounded by variational use of written language, especially in public displays such as shop signs, writing on mugs and t-shirts, posters, graffiti and so on, which link language to place. Equally, the voices we hear in entertainment and public broadcasting, far from being homogenous, celebrate diversity in Englishes. The homes and backgrounds of pupils in our schools, including their linguistic backgrounds, may also be very different either in terms of a different variation of English or languages spoken other than English. Since the emphasis is usually upon correct and fixed ways of teaching writing in English, it has often been difficult for teachers and pupils to reconcile the kind of English taught in school as the correct way and thus, by definition, all others as incorrect. However, narrow definitions of linguistic correctness are becoming increasingly difficult to uphold given that the public spaces with which we are surrounded are peppered by examples of variational use in writing. Recent sociolinguistic research into variation points to an increasing fluidity of linguistic use, especially when it comes to public displays of writing, particularly in media such as newspapers, websites, shop signs, TV channel logos and so on. Linguistic variability can thus be seen as a resource in creating unique voices and marking allegiance to, for example, a particular place and culture. Such research is indicative of the fact that variational use of English, far from being incorrect or illiterate, is increasingly being drawn upon creatively to mark a place identity. It also points to a shift in our conceptual thinking about language(s) and varieties from being perceived as static, fixed, totalised and immobile to being thought of as dynamic, fragmented and mobile, with the focus upon mobile resources rather than immobile languages. At the same time, the teaching of literacy centres upon the teaching of linguistic norms of spelling and grammar as fixed. There is a tension then, between creative expression of linguistic use often linked to place and those linked to standard English. This article explores those tensions and discusses the implications and possibilities for the teaching of English and literacy. © 2013.
In emotional speech research, it has been suggested that loudness, along with other prosodic features, may be an important cue in communicating high activation affects. In earlier studies, we found different voice quality stimuli to be consistently associated with certain affective states. In these stimuli, as in typical human productions, the different voice qualities entailed differences in loudness. To examine the extent to which the loudness differences among these voice qualities might influence the affective coloring they impart, two experiments were conducted with the synthesized stimuli, in which loudness was systematically manipulated. Experiment 1 used stimuli with distinct voice quality features including intrinsic loudness variations and stimuli where voice quality (modal voice) was kept constant, but loudness was modified to match the non-modal qualities. If loudness is the principal determinant in affect cueing for different voice qualities, there should be little or no difference in the responses to the two sets of stimuli. In Experiment 2, the stimuli included distinct voice quality features but all had equal loudness to test the hypothesis that equalizing the perceived loudness of different voice quality stimuli will have relatively little impact on affective ratings. The results suggest that loudness variation on its own is relatively ineffective whereas variation in voice quality is essential to the expression of affect. In Experiment 1, stimuli incorporating distinct voice quality features consistently obtained higher ratings than the modal voice stimuli with varied loudness. In Experiment 2, non-modal voice quality stimuli proved potent in affect cueing even with loudness differences equalized. Although loudness per se does not seem to be the major determinant of perceived affect, it can contribute positively to affect cueing: when combined with a tense or modal voice quality, increased loudness can enhance signaling of high activation states.
This paper describes our approaches to Na-tive Language Identification (NLI) for the NLI shared task 2013. NLI as a sub area of au-thor profiling focuses on identifying the first language of an author given a text in his sec-ond language. Researchers have reported sev-eral sets of features that have achieved rel-atively good performance in this task. The type of features used in such works are: lex-ical, syntactic and stylistic features, depen-dency parsers, psycholinguistic features and grammatical errors. In our approaches, we se-lected lexical and syntactic features based on n-grams of characters, words, Penn TreeBank (PTB) and Universal Parts Of Speech (POS) tagsets, and perplexity values of character of n-grams to build four different models. We also combine all the four models using an en-semble based approach to get the final result. We evaluated our approach over a set of 11 na-tive languages reaching 75 % accuracy. 1
At present, discourse parsing is an important research topic. Rhetorical Structure Theory (RST) is one of the most popular approaches in this field. In general, discourse parsing includes three stages: discourse segmentation, discourse relations detection and building up rhetorical trees. Different strategies are used when developing discourse parsers. One of the strategies to detect discourse relations is based on symbolic rules that take into account linguistic clues, such as discourse markers. Nevertheless, some discourse markers are ambiguous, that is, they can indicate more than one discourse relation. This fact constitutes a problem when assigning discourse relations automatically. In this paper, a symbolic approach to detect and solve discourse markers ambiguity in Spanish is developed. First, we detect ambiguous discourse markers, using the training corpus of the RST Spanish Treebank. Second, we extract linguistic contexts for these markers. Third, we design linguistic rules to solve the ambiguity of discourse markers. Fourth, we evaluate the rules, using the test corpus of the RST Spanish Treebank. Our approach outperforms the baseline created following the methodology of the state of the art. Therefore, we consider that the results obtained in our experiments are representative and constitute the first step towards the disambiguation of discourse markers senses in Spanish. However, there is room for improvement and the main limitations of the approach are presented. In the future, the rules will be integrated in a discourse parser for Spanish, and several related applications will be developed (automatic summarization and information extraction, among others).
The Penn Discourse Treebank (PDTB) was released to the public in 2008 and remains the largest corpus of manually annotated discourse relations — both relations that are signaled explicitly (e.g., by a coordinating or subordinating conjunction, or by a discourse adverbial or other construction) and ones that otherwise appear implicit. The Penn Discourse TreeBank also diverges from other discourse-annotated corpora in permitting more than one discourse relation to be annotated as holding concurrently. Annotators could indicate this by assigning multiple sense labels to an explicit connective. Or, in those cases where adjacent sentences had no explicit connective, annotators could indicate concurrent discourse relations by either annotating a single implicit connective that concurrently conveyed multiple senses or annotating multiple implicit connectives, each conveying one of the concurrent relation(s). Subsequent experiments carried out using Mechanical Turk showed that, when a discourse adverbial explicitly signalled a discourse relation, there was often a separate concurrent relation that could be associated with an implicit coordinating or subordinating conjunction. There are different circumstances in which different sets of concurrent discourse relations are taken to hold. I will go through these, and conclude with what I take the implications of this to be for various language technologies, including statistical machine translation. Bonnie Webber. 2013. Concurrent Discourse Relations. In Proceedings of Australasian Language Technology Association Workshop, page 3.
For years observational techniques along with other methods have sought to explore the relationships of couple interactional exchanges to marital quality and longevity. However, many of the previous methodological procedures used might be inadequate at capturing the influential micro-dimensional nuances of interpartner couple affective stability and reciprocity. This study explored the dyadic patterns in 23 married couples' continuous affect ratings during two communication episodes. Multilevel modeling was used to assess the structure in the stability of one's own affect and the influence of partner affect over 3-, 6-, and 9-second time lags. Implications regarding the use of nested models to explore patterns of actor and partner effects are discussed.
We present a comparative study of transition-, graph- and PCFG-based models aimed at illuminating more precisely the likely contribution of CFGs in improving Chinese dependency parsing accuracy, especially by combining heterogeneous models. Inspired by the impact of a constituency grammar on dependency parsing, we propose several strategies to acquire pseudo CFGs only from dependency annotations. Compared to linguistic grammars learned from rich phrase-structure treebanks, well designed pseudo grammars achieve similar parsing accuracy and have equivalent contributions to parser ensemble. Moreover, pseudo grammars increase the diversity of base models; therefore, together with all other models, further improve system combination. Based on automatic POS tagging, our final model achieves a UAS of 87.23%, resulting in a significant improvement of the state of the art.
This paper describes our submission for SemEval2013 Task 2: Sentiment Analysis in Twitter. For the limited data condition we use a lexicon-based model. The model uses an affective lexicon automaticallygeneratedfrom a very large corpus of raw web data. Statistics are calculated over the word and bigram affective ratings and used as features of a Naive Bayes tree model. For the unconstrained data scenario we combine the lexicon-based model with a classifier built on maximum entropy language models and trained on a large external dataset. The two models are fused at the posterior level to produce a final output. The approach proved successful, reaching rankings of 9th and 4th in the twitter sentiment analysis constrained and unconstrained scenario respectively, despite using only lexical features.
Morphology is the study of internal structure of words and is an essential early step in many NLP applications such as parsing and machine translation. Researchers working in Hindi NLP have either used the widely popular paradigm based analyzer (PBA) or extensions of it. In this work, we undertook a comprehensive evaluation of PBA using the data from the Hindi Treebank (HTB) and presented a new morphological analyzer trained on the HTB. Our morphological analyzer has better coverage and accuracy when compared to the existing analyzers for Hindi. An oracle system that takes the best values from the PBA’s output achieves only 63.41% for lemma, gender, number, person and case. Our statistical analyzer has an accuracy of 84.16% for these morphological attributes when evaluated on the test section of the Hindi Treebank.
Large-scale linguistically annotated cor-pora have played a crucial role in advanc-ing the state of the art of key natural lan-guage technologies such as syntactic, se-mantic and discourse analyzers, and they serve as training data as well as evaluation benchmarks. Up till now, however, most of the evaluation has been done on mono-lithic corpora such as the Penn Treebank, the Proposition Bank. As a result, it is still unclear how the state-of-the-art analyzers perform in general on data from a vari-ety of genres or domains. The completion of the OntoNotes corpus, a large-scale, multi-genre, multilingual corpus manually annotated with syntactic, semantic and discourse information, makes it possible to perform such an evaluation. This paper presents an analysis of the performance of publicly available, state-of-the-art tools on all layers and languages in the OntoNotes v5.0 corpus. This should set the bench-mark for future development of various NLP components in syntax and semantics, and possibly encourage research towards an integrated system that makes use of the various layers jointly to improve overall performance. 1
This paper presents an effective algorithm of annotation adaptation for constituency treebanks, which transforms a treebank from one annotation guideline to another with an iterative optimization procedure, thus to build a much larger treebank to train an enhanced parser without increasing model complexity. Experiments show that the transformed Tsinghua Chinese Treebank as additional training data brings significant improvement over the baseline trained on Penn Chinese Treebank only. 1
Abstract. Recent research and development have created the necessary ingredients for a major push in web-scale language understanding: large repositories of structured knowledge (DBpedia, the Google knowledge graph, Freebase, YAGO) progress in language processing (parsing, information extraction, computational semantics), linguistic knowledge resources (Treebanks, WordNet, BabelNet, UWN) and new powerful techniques for machine learning. A major goal is the automatic aggregation of knowledge from textual data. A central component of this endeavor is relation extraction (RE). In this paper, we will outline a new approach to connecting repositories of world knowledge with linguistic knowledge (syntactic and lexical semantics) via web-scale relation extraction technologies.
The aim of the article is to identify the exponent for the semantic prime TOUCH in Old English. Therefore, this research contributes to the frame of the Natural Semantic Metalanguage Research Programme (NSMRP) by applying it to the study of a historical language. Throughout such an application several descriptive and methodological questions arise. On the descriptive side, it is necessary to propose a cluster of semantic, morphological, textual and syntactic criteria that allow for the identification of the prime at stake, given that the nature of the object of study is not compatible with the translation into the native language generally adopted by the NSMRP. The analysis focuses on the category actions, events, movement and contact, and relies on data retrieved from the Historical Thesaurus of the Oxford English Dictionary, the Dictionary of Old English Corpus and the lexical database of Old English Nerthus. Although the cluster of criteria evinces a clear candidate for semantic prime it also raises the methodological issue of the distinction between the semantic prime and the hyperonym because some of the criteria used in the search for the former also play a role in the process of identification of the latter. The conclusion is reached that the verb hrīnan is the main exponent for the semantic prime TOUCH in Old English because it satisfies the criteria of meaning, word-formation, textual frequency and syntactic complementation.
Today, more than ten years after the resolution of the language controversy on a state level, it is far from resolved in the scholarly or the public sphere, where representatives of the two countries continue to debate the question of how distinct Macedonian is from Bulgarian. This chapter explains why this question is so hotly debated and so politicized. It deals with the process of codification of the contemporary Macedonian linguistic norm and with the conflicts between Bulgarians and Macedonians about the definition both of the Slavic vernacular dialects in geographic Macedonia and of the Macedonian norm itself. The chapter argues that the codification of the contemporary Macedonian idiom cannot be understood without examining a larger international context. Finally, it shows how the codification of a separate Macedonian norm has shaped Bulgarian nationalist representations-especially in the field of linguistics. Keywords:Bulgarian; Macedonian linguistic norm; Slavic vernacular dialects
In a large scale study on 843 transcripts of Technology, Entertainment and Design (TED) talks, the authors address the relation between word usage and categorical affective ratings of lectures by a large group of internet users. Users rated the lectures by assigning one or more predefined tags which relate to the affective state evoked in the audience (e. g., ‘fascinating’, ‘funny’, ‘courageous’, ‘unconvincing’ or ‘long-winded’). By automatic classification experiments, they demonstrate the usefulness of linguistic features for predicting these subjective ratings. Extensive test runs are conducted to assess the influence of the classifier and feature selection, and individual linguistic features are evaluated with respect to their discriminative power. In the result, classification whether the frequency of a given tag is higher than on average can be performed most robustly for tags associated with positive valence, reaching up to 80.7% accuracy on unseen test data.
Tree substitution grammar (TSG) is a generalization of context-free grammar (CFG) that permits non-terminals to rewrite as fragments of arbitrary size, instead of just depth-one productions. We discuss connections between the TSG framework and the larger family of usage-based approaches to language, showing how TSG allows us to make some of the claims of these approaches sufficiently concrete for computational modeling. A fundamental difficulty in defining a TSG is to determine the set of fragments for the grammar, because the set of possible fragments is exponential in the size of the parse trees from which TSGs are typically learned. We describe a model-based approach that learns a TSG using Gibbs sampling with a non-parametric prior to control fragment size, yielding grammars that contain mostly small fragments but that include larger ones. as the data permits. We evaluate these grammars on two tasks (parsing accuracy and grammaticality classification), and find that these Bayesian TSGs achieve excellent performance on two tasks relative to a set of heuristically extracted TSGs spanning the spectrum of representations, from a standard depth-one context-free Treebank grammar to explicit approximations of the Data-Oriented Parsing model.
A major computational burden, while performing document clustering, is the calculation of similarity measure between a pair of documents. Similarity measure is a function that assigns a real number between 0 and 1 to a pair of documents, depending upon the degree of similarity between them. A value of zero means that the documents are completely dissimilar whereas a value of one indicates that the documents are practically identical. Traditionally, vector-based models have been used for computing the document similarity. The vector-based models represent several features present in documents. These approaches to similarity measures, in general, cannot account for the semantics of the document. Documents written in human languages contain contexts and the words used to describe these contexts are generally semantically related. Motivated by this fact, many researchers have proposed seman-tic-based similarity measures by utilizing text annotation through external thesauruses like WordNet (a lexical database). In this paper, we define a semantic similarity measure based on documents represented in topic maps. Topic maps are rapidly becoming an industrial standard for knowledge representation with a focus for later search and extraction. The documents are transformed into a topic map based coded knowledge and the similarity between a pair of documents is represented as a correlation between the common patterns (sub-trees). The experimental studies on the text mining datasets reveal that this new similarity measure is more effective as compared to commonly used similarity measures in text clustering.
We explored the influence of implicit motives and activity inhibition (AI) on subjectively experienced affect in response to the presentation of six different facial expressions of emotion (FEEs; anger, disgust, fear, happiness, sadness, and surprise) and neutral faces from the NimStim set of facial expressions (Tottenham et al., 2009). Implicit motives and AI were assessed using a Picture Story Exercise (PSE) (Schultheiss et al., 2009b). Ratings of subjectively experienced affect (arousal and valence) were assessed using Self-Assessment Manikins (SAM) (Bradley and Lang, 1994) in a sample of 84 participants. We found that people with either a strong implicit power or achievement motive experienced stronger arousal, while people with a strong affiliation motive experienced less arousal and less pleasurable affect across emotions. Additionally, we obtained significant power motive × AI interactions for arousal ratings in response to FEEs and neutral faces. Participants with a strong power motive and weak AI experienced stronger arousal after the presentation of neutral faces but no additional increase in arousal after the presentation of FEEs. Participants with a strong power motive and strong AI (inhibited power motive) did not feel aroused by neutral faces. However, their arousal increased in response to all FEEs with the exception of happy faces, for which their subjective arousal decreased. These differentiated reaction patterns of individuals with an inhibited power motive suggest that they engage in a more socially adaptive manner of responding to different FEEs. Our findings extend established links between implicit motives and affective processes found at the procedural level to declarative reactions to FEEs. Implications are discussed with respect to dual-process models of motivation and research in motive congruence.
This paper analyzes the subjects/modules which presumably contribute to develop the communicative competence in Level C1 from the students of Primary Education Degree. So, taking as a starting point their course descriptions, we will check to what extend they reflect the communicative competence concept set by the CEFR and the teaching of linguistic norm.
It was repeatedly demonstrated that a negative emotional context enhances memory for central details while impairing memory for peripheral information. This trade-off effect is assumed to result from attentional processes: a negative context seems to narrow attention to central information at the expense of more peripheral details, thus causing the differential effects in memory. However, this explanation has rarely been tested and previous findings were partly inconclusive. For the present experiment 13 negative and 13 neutral naturalistic, thematically driven picture stories were constructed to test the trade-off effect in an ecologically more valid setting as compared to previous studies. During an incidental encoding phase, eye movements were recorded as an index of overt attention. In a subsequent recognition phase, memory for central and peripheral details occurring in the picture stories was tested. Explicit affective ratings and autonomic responses validated the induction of emotion during encoding. Consistent with the emotional trade-off effect on memory, encoding context differentially affected recognition of central and peripheral details. However, contrary to the common assumption, the emotional trade-off effect on memory was not mediated by attentional processes. By contrast, results suggest that the relevance of attentional processing for later recognition memory depends on the centrality of information and the emotional context but not their interaction. Thus, central information was remembered well even when fixated very briefly whereas memory for peripheral information depended more on overt attention at encoding. Moreover, the influence of overt attention on memory for central and peripheral details seems to be much lower for an arousing as compared to a neutral context.