Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This study undertook a critical appraisal of the correlation between the intractable social conflicts like the Boko-Haram and the Niger Delta crises, where youths are the key players, on the international image of Nigeria and tourism development in the country. It is motivated by the avalanche of media reports that the country’s image is being seriously battered abroad by these internal social problems. The specific objectives sought were to: ascertain the correlation between the Boko-Haram crisis and the nation’s image ratings abroad; the Niger Delta crisis and the nation’s image ratings abroad and their impacts on tourism development in the country. Survey design was adopted in the study, where electronic questionnaires (E-questionnaire) via the Internet were used to gather the primary data. The data so sourced were statistically presented/analyzed with Likert’s 5-points scale, Spearman’s correlation coefficient and Friedman chi-square. Results obtained show that both the Boko Haram crisis and the Niger Delta crisis have adverse impact on the country’s international image and tourism development, consequently on youths’ unemployment rate. It was then recommended that proactive public relations crisis management strategies should be used in nipping such crisis in their buds in future. Keywords: Boko Haram crisis, Niger Delta crisis, National Image, Tourism Development.
Conventional statistics-based methods for joint Chinese word segmentation and part-of-speech tagging (S&T) have generalization ability to recognize new words that do not appear in the training data. An undesirable side effect is that a number of meaningless words will be incorrectly created. We propose an effective and efficient framework for S&T that introduces features to significantly reduce meaningless words generation. A general lexicon, Wikepedia and a large-scale raw corpus of 200 billion characters are used to generate word-based features for the wordhood. The word-lattice based framework consists of a character-based model and a word-based model in order to employ our word-based features. Experiments on Penn Chinese treebank 5 show that this method has a 62.9% reduction of meaningless word generation in comparison with the baseline. As a result, the F1 measure for segmentation is increased to 0.984.
This paper builds six dependence syntactic networks based on six treebanks of different styles and gives a comparative analysis of overall characteristics of the networks, including the number of edges, the number of the nodes, the average degree, the clustering coefficient, the average path length, the centralization, the diameter, and the index of power-law, coefficient of determination. After that, the paper uses the Euclideanthe shortest distancemethod, with characteristics as variables, to do clustering analysis of these networks. The results show that using some main parameters of networks, namely the number of the nodes, the clustering coefficient, the average path length, the centralization and the index of power-law, can do cluster analysis on texts. Compared with the traditional text clustering, the results are easier to explain in linguistic angle.
Recent research has found it useful to distinguish between the form and meaning of sounds. To investigate the relevance of meaning, naïve students and professional drivers listened to four levels of meaning neutralisation and four levels of spectral slope of recorded truck sound. Self-assessment of emotional reactions showed that professional drivers did not vary much in activation and rated over all lower activation than naïve participants whose affect ratings moved more or less along the annoyance correlation line in the upper left quadrant of the affect map. This gives some information about the importance of the source being recognisable and of previous user experience for product sound quality. It is further supported by that the overall difference between naïve participants’ and professional drivers’ ratings decreased with increasing meaning neutralisation. The methodology applied in the current study may be adopted to form homogenous panels of experts for sound evaluation.
bilabial stop (although any literate adult might believe that this word’s orthography perfectly reflects its pronunciation). The following chapter continues with the closely related topic of determining word boundaries in spoken language and teaching students to write correctly. This issue naturally raises questions of what constitutes a “word,” particularly in relation to compound nouns such as sac à dos, which students might attempt to spell as “sacados.” Pothier recommends helping students sort through their mental lexicons, trying to identify different phrases or expressions in which a word may appear (for example, sac-poubelle, sac de couchage). From these two initial chapters and their clear and concrete examples and implications, the book ventures into more abstract territory, discussing the arbitrary relationship between signifiant and signifié (to illustrate this point, Pothier presents the numerical systems of a handful of foreign languages); synonymy and polysemy; language functions (including communicative and cognitive, but also metalinguistic and poetic); and prescriptive linguistic norms. Pothier’s presentation of these concepts is clear and interesting, but their pedagogical implications seem a bit less clear. The final chapter presents a hodgepodge of well-established trucs concerning the teaching of French, including discussions of mute and aspirated h, invariable words, two-verb constructions (the second verb always appearing in its infinitival form, except for in compound tenses), and differences between the conditional and future forms. Although this book is primarily intended for teachers of French as a first language, the notion that language teachers who understand the inner workings of the language will be better equipped to anticipate and correct student errors certainly also holds for second language acquisition. Writing for her specific audience of French (language) school teachers, however, Pothier offers little in terms of practical advice for second language pedagogues. The issue of nonnative-speaking pupils is directly addressed in the discussion of phonetics/phonology and their relationship with orthography, but this issue is not pursued any further in subsequent chapters. With more advice specifically directed toward teachers of second language learners, Pothier could have greatly expanded her target audience. As it stands now, this book constitutes a fairly good introduction to some of the basic concepts of (French) linguistics but falls a bit short of making clear connections (beyond the most obvious ones) between the study of linguistics and language pedagogy. Indiana University-Purdue University, Indianapolis A. Kate Miller SCHNEDECKER, CATHERINE, et CONSTANZE ARMBRECHT, éd. La quantification et ses domaines: actes du colloque de Strasbourg, 19–21 octobre 2006. Paris: Champion, 2012. ISBN 978-2-7453-2443-6. Pp. 657. 100 a. Conference proceedings, like journals, fall into two categories: refereed and non-refereed. Those that are non-refereed, and thus include all papers that were presented at the conference, occasionally display a wide range of quality, both in terms of the research conducted and the presentation of that research. La quantification is a pleasant surprise in this regard, however. Despite being non-refereed and apparently all-inclusive, this volume attests to the care that was given to the choice of conference presenters; there is not a single article that is extraneous to the topic of quantification, and many complement one another in ways that suggest Reviews 1237 that their authors routinely interact at conferences and through other scholarly exchanges. The forty-nine articles included here deal with the semantic issue of quantification from a wide variety of perspectives (too many to summarize in a review), examining a variety of source languages including French (primarily), Portuguese, Italian, Russian, Spanish, Korean and Romanian. Approximately 10% of the articles are written in English; the others are in French. The volume is subdivided into thematic units, including quantification, restriction, focalisation, totalisation; quantification via la prédication; quantification et négation. This organization allows the reader to find articles of interest among the many included. This is not a volume that one would read linearly, from cover to cover. There are, however, a plethora of articles, each different in its own way, which would interest a semanticist dealing with quantification. To the researcher working in this field, this volume offers a great deal. Because it contains conference proceedings, however, the articles are often short and at times somewhat undeveloped. Papers average...
This paper presents a novel method using graph-based semi-supervised learning (SSL) to improve the syntax parsing of unknown words. Different from conventional approaches that uses hand-crafted rules, rich morphological features, or a character-based model to handle unknown words, this method is based on a graph-based label propagation technique. It gives greater improvement on grammars trained on a smaller amount of labeled data and a large amount of unlabeled one. A transductiv 1 graph-based SSL method is employed to propagate POS and derive the emission distributions from labeled data to unlabeled one. The derived distributions are incorporated into the parsing process. The proposed method effectively augments the original supervised parsing model by contributing 2.28 % and 1.72 % absolute improvement on the accuracy of POS tagging and syntax parsing for Penn Chinese Treebank respectively. 1
This paper presents a new method of analysis by which structural similarities between brain data and linguistic data can be assessed at the semantic level. It shows how to measure the strength of these structural similarities and so determine the relatively better fit of the brain data with one semantic model over another. The first model is derived from WordNet, a lexical database of English compiled by language experts. The second is given by the corpus-based statistical technique of latent semantic analysis (LSA), which detects relations between words that are latent or hidden in text. The brain data are drawn from experiments in which statements about the geography of Europe were presented auditorily to participants who were asked to determine their truth or falsity while electroencephalographic (EEG) recordings were made. The theoretical framework for the analysis of the brain and semantic data derives from axiomatizations of theories such as the theory of differences in utility preference. Using brain-data samples from individual trials time-locked to the presentation of each word, ordinal relations of similarity differences are computed for the brain data and for the linguistic data. In each case those relations that are invariant with respect to the brain and linguistic data, and are correlated with sufficient statistical strength, amount to structural similarities between the brain and linguistic data. Results show that many more statistically significant structural similarities can be found between the brain data and the WordNet-derived data than the LSA-derived data. The work reported here is placed within the context of other recent studies of semantics and the brain. The main contribution of this paper is the new method it presents for the study of semantics and the brain and the focus it permits on networks of relations detected in brain data and represented by a semantic model.
Second language learners’ (L2ers’) perception and production of consonant clusters is influenced by the syllable structure of the native language (L1). This study investigates whether the perception of epenthetic vowels is partially responsible for why Spanish speakers have difficulty producing /s/ + Consonant (“sC”) clusters in English, and whether it affects word recognition in continuous speech. Spanish, German L2ers of English, and native English speakers completed: (i) an AXB task with (/ə/)sC-initial nonce words (e.g., [əsman]-[sman]); (ii) a word monitoring task with (/ə/)sC-initial words in semantically ambiguous sentences (e.g., I have lived in that (e)state for a long time); and (iii) a production task with the same sentences as in (i). L2ers also took a word-familiarity rating task and a cloze test to assess their proficiency. For (i) and (ii), accuracy rates were recorded, and response times were measured from target onset. For (iii), acoustic analyses showed whether the L2ers’ productions of sC-initial words contained an epenthetic vowel. Preliminary results suggest that perception difficulties may be partially responsible for Spanish speakers’ production and word-recognition difficulties with sC-clusters in English, but production data suggest that articulatory problems may also play an important role. Proficiency does not seem to help overcome this difficulty.
Discourse structure and discourse relations are an important ingredient in systems for the analysis of text that go beyond the boundary of single clauses. Discourse relations often indicate important additional information about the connection between two clauses, such as causality, and are widely believed to have an influence on aspects of reference resolution.In this article, we first present the general design choices that are to be made in the design of an annotation scheme for discourse structure and discourse relations. In a second part, we present the scheme used in our annotation of selected articles from the TüBa-D/Z treebank of German (Telljohann et al., 2009). The scheme used in the annotation is theory-neutral, but informed by more detailed linguistic knowledge in the way of linguistic tests that can help disambiguate between several plausible relations.
We present an automatic animacy classier for Dutch that can determine the animacy status of nouns | how alive the noun’s referent is (human, inanimate, etc.). Animacy is a semantic property that has been shown to play a role in human sentence processing, felicity and grammaticality. Although animacy is not marked explicitly in Dutch, we expect knowledge about animacy to be helpful for parsing, translation and other NLP tasks. Only a few animacy classiers and animacyannotated corpora exist internationally. For Dutch, animacy information is only available in the Cornetto lexical-semantic database. We augment this lexical information with context information from the Dutch Lassy Large treebank, to create training data for an animacy classier that uses a novel kind of context features. We use the k-nearest neighbour algorithm with distributional lexical features, e.g. how frequently the noun occurs as a subject of the verb ‘to think’ in a corpus, to decide on the (predominant) animacy class. The size of the Lassy Large corpus makes this possible, and the high level of detail these word association features provide, results in accurate Dutch-language animacy classication.
Traditional information retrieval systems rely on keywords to index documents and queries. In such systems, documents are retrieved based on the number of shared keywords with the query. This lexicalfocused retrieval leads to inaccurate and incomplete results when different keywords are used to describe the documents and queries. Semantic-focused retrieval approaches attempt to overcome this problem by relying on concepts rather than on keywords to indexing and retrieval. The goal is to retrieve documents that are semantically relevant to a given user query. This paper addresses this issue by proposing a solution at the indexing level. More precisely, we propose a novel approach for semantic indexing based on concepts identified from a linguistic resource. In particular, our approach relies on the joint use of WordNet and WordNetDomains lexical databases for concept identification. Furthermore, we propose a semantic-based concept weighting scheme that relies on a novel definition of concept centrality. The resulting system is evaluated on the TIME test collection. Experimental results show the effectiveness of our proposition over traditional IR approaches.
In this paper, we introduce experiment results of a Vietnamese sentence parser which is built by using the Chomsky's subcategorization theory and PDCG (Probabilistic Definite Clause Grammar). The efficiency of this subcategorized PDCG parser has been proved by experiments, in which, we have built by hand a Treebank with 1000 syntactic structures of Vietnamese training sentences, and used different testing datasets to evaluate the results. As a result, the precisions, recalls and F-measures of these experiments are over 98%.
This article analyses the structure of Yoruba numerals and their derivation. Data are collected from the compilation of Yoruba numerals and observation of its use coupled with the researcher's intuitive knowledge of the language. The work dwells on the existing literature on numerals too. The author adopts a descriptive method in analysing the data. The work looks at the roles of affixes in realising odd numbers, multiples of 20, centenary, bicentenary, and so on in their order of increase. It is discovered that the direction of counting in Yoruba is largely progressive. Besides, the language adopts base 5, decimal (base 10) and vigesimal (base 20) systems of counting. It is equally discovered that the choice of either of the two variations is largely dependent on the articulatory parameter of the first vowel (V1) of the root word. It is noted that the Yoruba numeral system offers a suitable linguistic database for both the theoretical and empirical domains of linguistic study especially documentary linguistics. The current study has general pedagogic implications for the teaching and learning of Yoruba numerals.
Abstract—Syntactic parsers are designed to detect the complete syntactic structure of grammatically correct sentences. In this paper, we introduce the concept of n-gram parsing, which corresponds to generating the constituency parse tree of n consecutive words in a sentence. We create a stand-alone n-gram parser derived from a baseline full discriminative constituency parser and analyze the characteristics of the generated n-gram trees for various values of n. Since the produced n-gram trees are in general smaller and less complex compared to full parse trees, it is likely that n-gram parsers are more robust compared to full parsers. Therefore, we use n-gram parsing to boost the accuracy of a full discriminative constituency parser in a hierarchical joint learning setup. Our results show that the full parser jointly trained with an n-gram parser performs statistically significantly better than our baseline full parser on the English Penn Treebank test corpus. Index Terms—Constituency parsing, n-gram parsing, discriminative learning, hierarchical joint learning. I.
A new operationalization was used to model a schema-based approach to moral judgment, as well as compare it to predictions based on the Social Intuitionist Model. Judgments were made about the moral wrongness of killing different animals. At Time 1, only moral judgments were made. At Time 2 judgments were made again, with questions and scales relating to attributing morally relevant cognitive capacities also included; further, two randomized conditions varied the presentation order of the scales. Differences between Time 1 and 2 indicated a reversed perspective-taking effect, with animals of lower capacities rated less empathically at Time 2. Affective ratings and attributed capacities were compared as different predictors, showing attributed capacities being more powerful. A group comparison was also made between active animal rights proponents and non-proponents, showing differences on several factors. These and other findings are discussed with relation to the Social Intuitionist Model and a schema-based account of morality.
In this paper, we consider the problem of cross-formalism transfer in parsing. We are interested in parsing constituencybased grammars such as HPSG and CCG using a small amount of data specific for the target formalism, and a large quantity of coarse CFG annotations from the Penn Treebank. While all of the target formalisms share a similar basic syntactic structure with Penn Treebank CFG, they also encode additional constraints and semantic features. To handle this apparent discrepancy, we design a probabilistic model that jointly generates CFG and target formalism parses. The model includes features of both parses, allowing transfer between the formalisms, while preserving parsing efficiency. We evaluate our approach on three constituency-based grammars — CCG, HPSG, and LFG, augmented with the Penn Treebank-1. Our experiments show that across all three formalisms, the target parsers significantly benefit from the coarse annotations. 1 1
This work is based on examples taken from tape-recordings made for a project entitled ‘Projeto da linguagem dos idosos velhos (LIV)’ (The speech of the old-old), comprising some fifty studies involving the interaction between young and old speakers in a wide variety of contextsat home, in old people’s homes and in rest homes. We have also used, exceptionally, some recordings from the ‘Projeto de estudo da norma lingüística urbana culta de São Paulo (NURC/SP)’ (Study of the educated urban linguistic norm of São Paulo, Brazil).
Using neural networks to estimate the probabilities of word sequences has shown significant promise for statistical language modeling. Typical modeling methods include multi-layer neural networks, log-bilinear networks and recurrent neural networks, etc. In this paper, we propose the temporal kernel neural network language model, a variant of models mentioned above. This model explicitly captures long-term dependencies of words with exponential kernel, where the memory of history is decayed exponentially. Additionally, several sentences with variable lengths as a mini-batch are efficiently implemented for speeding up. Experimental results show that the proposed model is very competitive to the recurrent neural network language model and obtains the lower perplexity of 111.6 (more than 10% reduction) than the state-of-the-art results reported in the standard Penn Treebank Corpus. We further apply this model to Wall Street Journal speech recognition task, and observe significant improvements in word error rate.
It has been observed that the inclusion of morphosyntactic information in dependency treebanks is crucial to obtain high results in dependency parsing for some languages. In this paper we explore in depth to what extent it is useful to include morphological features, and the impact of diverse morphosyntactic annotations on statistical dependency parsing of Spanish. For this, we give a detailed analysis of the results of over 80 experiments performed with MaltParser through the application of MaltOptimizer. Our goal is to isolate configurations of morphosyntactic features which would allow for optimizing the parsing of Spanish texts, and to evaluate the impact that each feature has, independently and in combination with others. 1
Compared to well-resourced languages such as English and Dutch, NLP tools for linguistic analysis in Afrikaans are still not abundant. In order to facilitate corpus-based linguistic research for Afrikaans, we are creating a treebank based on the Taalkommissie corpus. We adapted a tokenizer and a shallow parser, while using a TnT tagger to do part-of-speech annotation. A first linguistic phenomenon we are investigating is the occurrence of infinitivus pro participio (IPP) in Afrikaans. IPP refers to constructions with a perfect auxiliary, in which an infinitive appears instead of the expected past participle. The phenomenon has been studied extensively in Dutch and German, but studies on Afrikaans IPP triggers are sparse. In contrast to the former two languages, it is often mentioned in the literature that in Afrikaans, IPP occurs optionally. We want to check this statement doing a corpus analysis.
Depressive symptomatology is associated with impaired recognition of emotion. Previous investigations have predominantly focused on emotion recognition of static facial expressions neglecting the influence of social interaction and critical contextual factors. In the current study, we investigated how youth and maternal symptoms of depression may be associated with emotion recognition biases during familial interactions across distinct contextual settings. Further, we explored if an individual's current emotional state may account for youth and maternal emotion recognition biases. Mother-adolescent dyads (N = 128) completed measures of depressive symptomatology and participated in three family interactions, each designed to elicit distinct emotions. Mothers and youth completed state affect ratings pertaining to self and other at the conclusion of each interaction task. Using multiple regression, depressive symptoms in both mothers and adolescents were associated with biased recognition of both positive affect (i.e., happy, excited) and negative affect (i.e., sadness, anger, frustration); however, this bias emerged primarily in contexts with a less strong emotional signal. Using actor-partner interdependence models, results suggested that youth's own state affect accounted for depression-related biases in their recognition of maternal affect. State affect did not function similarly in explaining depression-related biases for maternal recognition of adolescent emotion. Together these findings suggest a similar negative bias in emotion recognition associated with depressive symptoms in both adolescents and mothers in real-life situations, albeit potentially driven by different mechanisms.
It is generally accepted that a discourse connective expresses a semantic and/or pragmatic relation between its matrix sentence or clause and something in the previous discourse. Usually the sense of this relation is expressed as a label, often within a hierarchy of sense labels. But the meaning of these labels may vary from system to system, and the same connective may be assigned different labels in different systems. Given this, we might learn more and make better predictions if (i) sense labels were associated with (some of) their entailments and (ii) connectives were characterized in terms of both their formal properties and their use conditions. I’ll give examples of both. The above-mentioned predictions tie in with an interesting property of Penn Discourse TreeBank annotation. Annotators were allowed to assign multiple sense labels to a single connective, to imply that all the senses held simultaneously. For those cases where adjacent sentences lacked an intervening connective, annotators were instructed to try to insert one or more connectives that (together) expressed the relation(s) between the sentences. Here too, in many cases, annotators inserted a single connective to which they assigned multiple meanings, Other times they inserted multiple connectives to convey the relation(s) they took as being expressed. Some of this will be shown to make more sense in terms of the entailments and formal properties of the connectives than in terms of any sense labels. I’ll close by trying to distinguish discourse connectives that are associated with coordinating or subordinating relations between sentences or clauses, which is an feature of discourse structure, from those connectives that simply convey additional relevant semantic or pragmatic content.
Abstract This short reply seeks to clarify the concept of linguistic norm circles and to correct some misunderstandings of it implicit in Sealey & Carter's response. It also reinforces some doubts over their version of the linguistic system. Norm circles, it argues, provide an important part of the explanation for linguistic practices, but always in conjunction with other interacting causal powers.
It has recently been shown that different NLP models can be effectively combined using dual decomposition.In this paper we demonstrate that PCFG-LA parsing models are suitable for combination in this way.We experiment with the different models which result from alternative methods of extracting a grammar from a treebank (retaining or discarding function labels, left binarization versus right binarization) and achieve a labeled Parseval F-score of 92.4 on Wall Street Journal Section 23 -this represents an absolute improvement of 0.7 and an error reduction rate of 7% over a strong PCFG-LA product-model baseline.Although we experiment only with binarization and function labels in this study, there is much scope for applying this approach to other grammar extraction strategies.
Tree Substitution Grammar rules form a large and expressive class of features capable of representing syntactic and lexical patterns that provide evidence of an author’s native language. However, this class of features can be applied to any general constituent based model of grammar and previous work has done little to explore these options, relying primarily on the common Penn Treebank annotation standard. In this work we contrast the performance of syntactic features for Native Language Indentification using five different formalisms. The use of different formalisms captures complementary information from second language data, and can be used in combination to yield classification performance superior to any formalism taken on its own. 1
This article considers dictionaries as lexical information / knowledge sources to be derived from a deeper, underlying, lexical database. These dictionary-tokens or -instantiations are inter alia specified by the users' needs. As a case in point of such a derivation meeting the needs of a multilingual society, a bidirectional bilingual learner dictionary is presented. Specific tools, such as editors with reversal function, and models, such as the hub-and-spoke model, are discussed as means to function within the lexicographical infrastructure of a multilingual society.
The state is the primary provider of education in Singapore. The rise of the global knowledge economy, however, has generated an education industry that is worth about USD2.2 trillion per year globally. The Government of Singapore therefore revamped its higher-education sector so that Higher Private Education Organisations (HPEOs) can offer more flexible transnational education to attract fee-paying, overseas students and increase the education industry’s contribution to the national GDP. The private-education industry is currently facing competition from more established markets such as the United States, United Kingdom and Australia. Locally, HPEOs need to upgrade their service standard in order to meet the stringent registration requirements imposed by Singapore’s regulators and, at the same time, compete with each other in a high-cost environment. HPEOs therefore need to gain a better understanding of the relations among variables such as service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth. A better understanding will allow HPEOs to improve their marketing efforts in order to obtain competitive advantages. Prior research only focused on the correlation of variables and failed to take into account the inter-relationship among two or more variables. The measurement of these variables is often dependent on geographical factors, types of services and types of stakeholders. It is therefore the objective of this study to examine the measurement and relationships among service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth in the context of HPEOs in Singapore. A survey involving 554 participants from local HPEOs was conducted over a period of three years. Analysis of the data shows that attributes such as image rating and service quality are unique and HPEOs have to customize these attributes to meet the needs and wants of their students. This study found that service quality, price satisfaction, image rating, overall satisfaction, repurchase intention and positive word of mouth are all positively correlated. It was also found that some factors (e.g., overall satisfaction and repurchase intention) are mediators of the relationships between other variables.
Although disgust propensity (DP) has been implicated in the development of some anxiety disorders, the mechanism that may account for this association has not been fully elucidated. The present study examined the extent to which the potentiation of learned aversion might be one such mechanism. Participants (n = 103) were randomized to one of two evaluative conditioning (EC) paradigms consisting of 12 reinforced conditioned stimulus (CS+) pairings of the word "part" (condition one) or "some" (condition two) with 12 aversive unconditioned stimulus (US) images, and 12 pairings of the CS- word "cylinder" with 12 neutral images. Participants then completed measures of DP and trait anxiety and provided subjective affective ratings for the aversive US. The findings revealed that participants experienced significantly more disgust, anxiety, anger, sadness, and less happiness toward their respective CS+. In contrast, participants experienced significantly more happiness toward the CS-. Examination of the magnitude of evaluative change to the CS+ revealed the strongest effect for disgust. DP, but not trait anxiety, also predicted a greater increase in disgust, anger, and anxiety in response to the CS+ relative to the CS-. Furthermore, the association between DP and greater disgust, anger, and anxiety in response to the CS+ was mediated by more intense negative affective responding to the US among those higher in DP. The implication of these findings for better understanding how DP may confer risk for anxiety-related psychopathology is discussed.
This article uses semi-supervised Expectation Maximization (EM) to learn lexico-syntactic dependencies, i.e. associations between words and the structures that occur with them. Due to Zipfian distributions in language, such dependencies are extremely sparse in labelled data, and unlabelled data are the only source for learning them. Specifically, we learn sparse lexical parameters of a generative parsing model (a Probabilistic Context-Free Grammar, PCFG) that is initially estimated over the Penn Treebank. Our lexical parameters are similar to supertags—they are fine-grained, and encode complex structural information at the pre-terminal level. Our goal is to use unlabelled data to learn these for words that are rare or unseen in the labelled data. We get large error reductions (up to 17.5%) in parsing ambiguous structures associated with unseen verbs, the most important case of learning lexico-structural dependencies, resulting in a statistically significant improvement in labelled bracketing score of the treebank PCFG. Our semi-supervised method incorporates structural and lexical priors from the labelled data to guide estimation from unlabelled data, and is the first successful use of semi-supervised EM to improve a generative structured model already trained over large labelled data. The method scales well to larger amounts of unlabelled data, and also gives substantial error reductions (up to 11.5%) for models trained on smaller amounts of labelled data, making it relevant to low-resource languages with small treebanks as well.
The Digital World encounters rapid development nowadays, especially through the proliferation of social media in Indonesia. Twitter has become one of social media with expanded users within every sectors of society. There are so many part both individual as well as organization/enterprise which utilize twitter as tool for communication, business, customer relation, and other activities. Through the twitter's ever-expanding users with those particular purposes, the precise method to effectively and efficiently analyzing opinion-contained sentences become crucially needed. Therefore this research made for method analyzing through lexical based and model based approaches by machine learning to classify opinion-contained tweets using those 2 methods. The tested machine learning method are Support Vector Machine (SVM), Maximum Entropy (ME), Multinomial Naive Bayes (MNB), and k-Nearest Neighbor (k-NN). Based on the test outcome, lexical based approach highly depended on lexical database which became opinion classification matrix. Whilst machine learning approach can produce better accuracy due to its capability in new training data modeling based on outcome model. However, machine learning model based approach depends on various factors in analyzing sentiment.
Abstract In this article we present some statistical data on the distribution of parts of speech and dependency relations in a large manually annotated Hungarian Treebank, the Szeged Dependency Treebank. We hypothesize that the domain of the text influences the distribution of the above elements, thus we pay special attention to differences between domains. We present the characteristic rank-frequency distributions of parts of speech and dependency relations in Hungarian and analyse the domain similarities and differences among sub-corpora as regards the above distributions. Our results reveal that the computer and newspaper texts are most similar to each other while the domains literature and compositions also exhibit some similarities. On the other hand, the business news and the law sub-corpora are unique, both having their own characteristics.
Body image disturbances are core symptoms of eating disorders (EDs). Recent evidence suggests that changes in body image may occur prior to ED onset and are not restricted to in-vivo exposure (e.g. mirror image), but also evident during presentation of abstract cues such as body shape and weight-related words. In the present study startle modulation, heart rate and subjective evaluations were examined during reading of body words and neutral words in 41 student female volunteers screened for risk of EDs. The aim was to determine if responses to body words are attributable to a general negativity bias regardless of ED risk or if activated, ED relevant negative body schemas facilitate priming of defensive responses. Heart rate and word ratings differed between body words and neutral words in the whole female sample, supporting a general processing bias for body weight and shape-related concepts in young women regardless of ED risk. Startle modulation was specifically related to eating disorder symptoms, as was indicated by significant positive correlations with self-reported body dissatisfaction. These results emphasize the relevance of examining body schema representations as a function of ED risk across different levels of responding. Peripheral-physiological measures such as the startle reflex could possibly be used as predictors of females' risk for developing EDs in the future.
Syntactic parsing is an important technique in the natural language processing, yet Latvian is still lacking an efficient general coverage syntax parser. This paper reports on the first experiments on statistical syntactic parsing for Latvian — a highly inflective Indo-European language with a relatively free word order. We have induced a statistical parser from a small, non-balanced Latvian Treebank using the MaltParser toolkit and measured the unlabeled attachment score (UAS). As MaltParser is based on the dependency grammar approach, we have also developed a convertor from the hybrid dependency-based annotation model used in the Latvian Treebank to the pure dependency annotation model. We have obtained a promising 74.63 % UAS in 10-fold cross-validation using only ~2500 sentences. The results revealed that best results can be achieved using non-projective stack parsing algorithm with lazy arc adding strategy, but comparably good results can be achieved using projective parsing algorithms combined with appropriate projectiviziation preprocessing.
This chapter deals with the main methodological issues underlying the building of the SciE-Lex lexical database and discusses and justifies the information included. SciE-Lex was initially conceived as a response to the lack of reference tools that can help scientists write scientific papers in phraseologically competent and native-like English. While there are a number of specialised dictionaries that include specific terminological information, there is a shortage of writing aids that provide information about the use of non-technical terms in scientific genres. SciE-Lex aims at filling this gap by focusing on the description of general terms in scientific English. This article describes the two stages in the building of the database, the first one including morphosyntactic and collocational information, and the second one focusing on phraseological information.
Turkish is an agglutinative language with rich morphology-syntax interactions. As an extension of this property, the Turkish Treebank is designed to represent sublexical dependencies, which brings extra challenges to parsing raw text. In this work, we use a joint POS tagging and parsing approach to parse Turkish raw text, and we show it outperforms a pipeline approach. Then we experiment with incorporating morphological feature prediction into the joint system. Our results show statistically significant improvements with the joint systems and achieve the state-ofthe-art accuracy for Turkish dependency parsing.
Today, more than ten years after the resolution of the language controversy on a state level, it is far from resolved in the scholarly or the public sphere, where representatives of the two countries continue to debate the question of how distinct Macedonian is from Bulgarian. This chapter explains why this question is so hotly debated and so politicized. It deals with the process of codification of the contemporary Macedonian linguistic norm and with the conflicts between Bulgarians and Macedonians about the definition both of the Slavic vernacular dialects in geographic Macedonia and of the Macedonian norm itself. The chapter argues that the codification of the contemporary Macedonian idiom cannot be understood without examining a larger international context. Finally, it shows how the codification of a separate Macedonian norm has shaped Bulgarian nationalist representations-especially in the field of linguistics. Keywords:Bulgarian; Macedonian linguistic norm; Slavic vernacular dialects
In a large scale study on 843 transcripts of Technology, Entertainment and Design (TED) talks, the authors address the relation between word usage and categorical affective ratings of lectures by a large group of internet users. Users rated the lectures by assigning one or more predefined tags which relate to the affective state evoked in the audience (e. g., ‘fascinating’, ‘funny’, ‘courageous’, ‘unconvincing’ or ‘long-winded’). By automatic classification experiments, they demonstrate the usefulness of linguistic features for predicting these subjective ratings. Extensive test runs are conducted to assess the influence of the classifier and feature selection, and individual linguistic features are evaluated with respect to their discriminative power. In the result, classification whether the frequency of a given tag is higher than on average can be performed most robustly for tags associated with positive valence, reaching up to 80.7% accuracy on unseen test data.
Tree substitution grammar (TSG) is a generalization of context-free grammar (CFG) that permits non-terminals to rewrite as fragments of arbitrary size, instead of just depth-one productions. We discuss connections between the TSG framework and the larger family of usage-based approaches to language, showing how TSG allows us to make some of the claims of these approaches sufficiently concrete for computational modeling. A fundamental difficulty in defining a TSG is to determine the set of fragments for the grammar, because the set of possible fragments is exponential in the size of the parse trees from which TSGs are typically learned. We describe a model-based approach that learns a TSG using Gibbs sampling with a non-parametric prior to control fragment size, yielding grammars that contain mostly small fragments but that include larger ones. as the data permits. We evaluate these grammars on two tasks (parsing accuracy and grammaticality classification), and find that these Bayesian TSGs achieve excellent performance on two tasks relative to a set of heuristically extracted TSGs spanning the spectrum of representations, from a standard depth-one context-free Treebank grammar to explicit approximations of the Data-Oriented Parsing model.
A major computational burden, while performing document clustering, is the calculation of similarity measure between a pair of documents. Similarity measure is a function that assigns a real number between 0 and 1 to a pair of documents, depending upon the degree of similarity between them. A value of zero means that the documents are completely dissimilar whereas a value of one indicates that the documents are practically identical. Traditionally, vector-based models have been used for computing the document similarity. The vector-based models represent several features present in documents. These approaches to similarity measures, in general, cannot account for the semantics of the document. Documents written in human languages contain contexts and the words used to describe these contexts are generally semantically related. Motivated by this fact, many researchers have proposed seman-tic-based similarity measures by utilizing text annotation through external thesauruses like WordNet (a lexical database). In this paper, we define a semantic similarity measure based on documents represented in topic maps. Topic maps are rapidly becoming an industrial standard for knowledge representation with a focus for later search and extraction. The documents are transformed into a topic map based coded knowledge and the similarity between a pair of documents is represented as a correlation between the common patterns (sub-trees). The experimental studies on the text mining datasets reveal that this new similarity measure is more effective as compared to commonly used similarity measures in text clustering.