Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Chunking or shallow syntactic parsing is proving to be a task of interest to many natural language processing applications. The problem gets worse for the Arabic language because of its specific features that make it quite different and even more ambiguous than other natural languages when processed. In this paper, we present a method for chunking Arabic texts based on supervised learning. We use the Conditional Random Fields algorithm and the Penn Arabic Treebank to train the model. For the experimentation, we use over than 10,100 sentences as training data and 2,524 sentences for the test. The evaluation of the method consists of the calculation of the generated model accuracy and the results are very encouraging.
This paper investigates the recognition of unknown words in Chinese parsing. Two methods are proposed to handle this problem. One is the modification of a character-based model. We model the emission probability of an unknown word using the first and last characters in the word. It aims to reduce the POS tag ambiguities of unknown words to improve the parsing performance. In addition, a novel method, using graph-based semisupervised learning (SSL), is proposed to improve the syntax parsing of unknown words. Its goal is to discover additional lexical knowledge from a large amount of unlabeled data to help the syntax parsing. The method is mainly to propagate lexical emission probabilities to unknown words by building the similarity graphs over the words of labeled and unlabeled data. The derived distributions are incorporated into the parsing process. The proposed methods are effective in dealing with the unknown words to improve the parsing. Empirical results for Penn Chinese Treebank and TCT Treebank revealed its effectiveness.
Natural language is a fundamental thing of human-society to communicate and interact with one another. In this globalization era, we interact with different regional people as per our interest in social, cultural, economical, educational and professional domain. There are thousands of natural languages exist in our earth. It is quite tough, rather impossible to know all the languages. So we need a computerized approach to convert one natural language to another as per our necessity. This computerized conversion among multiple languages is known as multilingual machine translation. But in this paper we work with a bilingual model, where we concern with two languages: English and Bengali. We use soft computational approach where fuzzy If-Then rule is applied to choose a lemma from prior knowledge; Penn TreeBank PoS tags and HMM tagger are used as lexical class marker to each word in corpora.
This paper introduces a new technique for phrase-structure parser analysis, catego-rizing possible treebank structures by inte-grating regular expressions into derivation trees. We analyze the performance of the Berkeley parser on OntoNotes WSJ and the English Web Treebank. This provides some insight into the evalb scores, and the problem of domain adaptation with the web data. We also analyze a “test-on-train ” dataset, showing a wide variance in how the parser is generalizing from differ-ent structures in the training material. 1
With growing interest in the creation and search of linguistic annotations that form general graphs (in contrast to formally simpler, rooted trees), there also is an increased need for infrastructures that support the exploration of such representations, for example logical-form meaning representations or semantic dependency graphs. In this work, we heavily lean on semantic technologies and in particular the data model of the Resource Description Framework (RDF) to represent, store, and efficiently query very large collections of text annotated with graph-structured representations of sentence meaning. Keywords:Semantic Dependency Graphs, Treebank Search, Resource Description Framework 1.
The web today is huge and enormous collection of data today and it goes on increasing day by day. Thus, searching for some particular data in this collection has a significant impact. Researches taking place give prominence to the relevancy and relatedness of the data that is found. Inspite of their relevance pages for any search topic, the results are still huge to be explored. Another important issue to be kept in mind is the users’ standpoint differs from time to time from topic to topic. Effective relevance prediction can help avoid downloading and visiting many irrelevant pages. The performance of a crawler depends mostly on the opulence of links in the specific topic being searched. This paper reviews the researches on web crawling algorithms used for searching. Keywords— Web Crawling Algorithms, Crawling Algorithm Survey, Search Algorithms, Lexical Database, Metadata, Semantic. __________________________________________________*****_________________________________________________
The sentiment mining approaches can typically be divided into lexicon and machine learning approaches. Recently there are an increasing number of approaches which combine both to improve the performance when used separately. However, this still lacks contextual understanding which led to the introduction of deep learning approaches which allows for semantic compositionality over a sentiment treebank. This paper enhances the deep learning approach with semantic lexicon so that scores can be computed in-stead merely nominal classification. Besides, neutral classification is also improved. Results suggest that the approach outperforms its original.
BACKGROUND: Careful observation of the longitudinal course of bipolar disorders is pivotal to finding optimal treatments and improving outcome. A useful tool is the daily prospective Life-Chart Method, developed by the National Institute of Mental Health. However, it remains unclear whether the patient version is as valid as the clinician version. METHODS: We compared the patient-rated version of the Lifechart (LC-self) with the Young-Mania-Rating Scale (YMRS), Inventory of Depressive Symptoms-Clinician version (IDS-C), and Clinical Global Impression-Bipolar version (CGI-BP) in 108 bipolar I and II patients who participated in the Naturalistic Follow-up Study (NFS) of the German centres of the Bipolar Collaborative Network (BCN; formerly Stanley Foundation Bipolar Network). For statistical evaluation, levels of severity of mood states on the Lifechart were transformed numerically and comparison with affective scales was performed using chi-square and t tests. For testing correlations Pearson´s coefficient was calculated. RESULTS: Ratings for depression of LC-self and total scores of IDS-C were found to be highly correlated (Pearson coefficient r = -.718; p <.001), whilst the correlation of ratings for mania with YMRS compared to LC-self were slightly less robust (Pearson coefficient r =.491; p =.001). These results were confirmed by good correlations between the CGI-BP IA (mania), IB (depression) and IC (overall mood state) and the LC-self ratings (Pearson coefficient r =.488, r =.721 and r =.65, respectively; all p <.001). CONCLUSIONS: The LC-self shows a significant correlation and good concordance with standard cross sectional affective rating scales, suggesting that the LC-self is a valid and time and money saving alternative to the clinician-rated version which should be incorporated in future clinical research in bipolar disorder. Generalizability of the results is limited by the selection of highly motivated patients in specialized bipolar centres and by the open design of the study.
This paper describes experiments for statistical dependency parsing using two different parsers trained on a recently extended dependency treebank for Greek, a language with a moderately rich morphology. We show how scores obtained by the two parsers are influenced by morphology and dependency types as well as sentence and arc length. The best LAS obtained in these experiments was 80.16 on a test set with manually validated POS tags and lemmas. 1
OBJECTIVE: Premenstrual dysphoric disorder (PMDD) is associated with increased pain, but there has been a lack of well-controlled research assessing pain responsivity, sex hormones, and their relationships in this group. This study was designed to address this gap in the literature. MATERIALS AND METHODS: Healthy, regularly cycling participants (14 PMDD, 14 non-PMDD) attended pain testing sessions during the mid-follicular, ovulatory, and late-luteal phases of the menstrual cycle (order counterbalanced) and salivary estradiol, progesterone, and testosterone were assessed at each testing session. Pain sensitivity was measured from electrocutaneous threshold/tolerance, ischemic threshold/tolerance, sensory and affective ratings of electrocutaneous and ischemic stimuli, and the nociceptive flexion reflex threshold (NFR, a measure of spinal nociception). RESULTS: Women with PMDD had higher sensory pain ratings of electrocutaneous stimuli and trends for lower ischemic thresholds and higher affective pain ratings of electrocutaneous stimuli. However, there were no group differences observed in NFR threshold. Testosterone levels were also lower during the mid-follicular and ovulatory phases in PMDD. Correlations between pain outcomes and estradiol and testosterone indicated that these hormones are hypoalgesic, with estradiol having a greater hypoalgesic effect within the PMDD group. DISCUSSION: Overall, women with PMDD may have a phase-independent hyperalgesia, with pain amplification likely occurring at the supraspinal level rather than the spinal level, given the lack of group differences in NFR threshold. Because testosterone was hypoalgesic and lower in women with PMDD, and there were strong associations between pain and estradiol in PMDD, sex hormones may play a role in PMDD-related hyperalgesia.
Linguistic norms emerge in human communities because people imitate each other. A shared linguistic system provides people with the benefits of shared knowledge and coordinated planning. Once norms are in place, why would they ever change? This question, echoing broad questions in the theory of social dynamics, has particular force in relation to language. By definition, an innovator is in the minority when the innovation first occurs. In some areas of social dynamics, important minorities can strongly influence the majority through their power, fame, or use of broadcast media. But most linguistic changes are grassroots developments that originate with ordinary people. Here, we develop a novel model of communicative behavior in communities, and identify a mechanism for arbitrary innovations by ordinary people to have a good chance of being widely adopted. To imitate each other, people must form a mental representation of what other people do. Each time they speak, they must also decide which form to produce themselves. We introduce a new decision function that enables us to smoothly explore the space between two types of behavior: probability matching (matching the probabilities of incoming experience) and regularization (producing some forms disproportionately often). Using Monte Carlo methods, we explore the interactions amongst the degree of regularization, the distribution of biases in a network, and the network position of the innovator. We identify two regimes for the widespread adoption of arbritrary innovations, viewed as informational cascades in the network. With moderate regularization of experienced input, average people (not well-connected people) are the most likely source of successful innovations. Our results shed light on a major outstanding puzzle in the theory of language change. The framework also holds promise for understanding the dynamics of other social norms.
The increasing diversity of languages used on the web introduces a new level of complexity to Information Retrieval (IR) systems. We can no longer assume that textual content is written in one language or even the same language family. In this paper, we demonstrate how to build massive multilingual annotators with minimal human expertise and intervention. We describe a system that builds Named Entity Recognition (NER) annotators for 40 major languages using Wikipedia and Freebase. Our approach does not require NER human annotated datasets or language specific resources like treebanks, parallel corpora, and orthographic rules. The novelty of approach lies therein - using only language agnostic techniques, while achieving competitive performance. Our method learns distributed word representations (word embeddings) which encode semantic and syntactic features of words in each language. Then, we automatically generate datasets from Wikipedia link structure and Freebase attributes. Finally, we apply two preprocessing stages (oversampling and exact surface form matching) which do not require any linguistic expertise. Our evaluation is two fold: First, we demonstrate the system performance on human annotated datasets. Second, for languages where no gold-standard benchmarks are available, we propose a new method, distant evaluation, based on statistical machine translation.
Processing unpleasant affective cues induces elevated momentary symptom reports, especially in persons with high levels of symptom reporting in daily life. The present study aimed to examine whether applying an emotion regulation strategy, i.e. affect labeling, can inhibit these emotion influences on symptom reporting. Student participants (N = 61) with varying levels of habitual symptom reporting completed six picture viewing trials of homogeneous valence (three pleasant, three unpleasant) under three conditions: merely viewing, emotional labeling, or content (non-emotional) labeling. Affect ratings and symptom reports were collected after each trial. Participants completed a motor inhibition task and self-control questionnaires as indices of their inhibitory capacities. Heart rate variability was also measured. Labeling, either emotional or non-emotional, significantly reduced experienced affect, as well as the elevated symptoms reports observed after unpleasant picture viewing. These labeling effects became more pronounced with increasing levels of habitual symptom reporting, suggesting a moderating role of the latter variable, but did not correlate with any index of general inhibitory capacity. Our findings suggest that using an emotion regulation strategy, such as labeling emotional stimuli, can reverse the effects of unpleasant stimuli on symptom reporting and that such strategies can be especially beneficial for individuals suffering from medically unexplained physical symptoms.
Recent resting-state functional magnetic resonance imaging (fMRI) studies using graph theory metrics have revealed that the functional network of the human brain possesses small-world characteristics and comprises several functional hub regions. However, it is unclear how the affective functional network is organized in the brain during the processing of affective information. In this study, the fMRI data were collected from 25 healthy college students as they viewed a total of 81 positive, neutral, and negative pictures. The results indicated that affective functional networks exhibit weaker small-worldness properties with higher local efficiency, implying that local connections increase during viewing affective pictures. Moreover, positive and negative emotional processing exhibit dissociable functional hubs, emerging mainly in task-positive regions. These functional hubs, which are the centers of information processing, have nodal betweenness centrality values that are at least 1.5 times larger than the average betweenness centrality of the network. Positive affect scores correlated with the betweenness values of the right orbital frontal cortex (OFC) and the right putamen in the positive emotional network; negative affect scores correlated with the betweenness values of the left OFC and the left amygdala in the negative emotional network. The local efficiencies in the left superior and inferior parietal lobe correlated with subsequent arousal ratings of positive and negative pictures, respectively. These observations provide important evidence for the organizational principles of the human brain functional connectome during the processing of affective information.
The chapter argues that language, which rests on the sharing of linguistic norms, honest information, and moral norms, evolved through a co-evolutionary process with a pivotal role for intersubjectivity. Mainstream evolutionary models, based only on individual-level and gene-level selection, are argued to be incapable to account for such sharing of care, values and information, thus implying the need to evoke multi-level selection, including (cultural) group selection. Four of the most influential current theories of the evolution of human-scale sociality, those of Dunbar, Deacon, Tomasello and Hrdy, are compared and evaluated on the basis of their answers to five questions: (1) Why we and not others? (2) How: by what mechanisms? (3) When? (4) In what kind of social settings? (5) What are the implications for ontogeny? The conclusions are that the theories are to a large degree complementary, and that they all assume, explicitly or not, a role for group selection. Hrdy’s theory, focusing on the evolution of alloparenting, is argued to provide the best explanation for the onset of the evolution of human intersubjectivity, and can furthermore offer a Darwinian framework for Tomasello’s theory of shared intentionality. Deacon’s theory deals rather with the evolution of morality and its co-evolution with “symbolic reference”, but these are necessarily antecedent to the primary evolution of human intersubjectivity. Dunbar’s theory on the transition from “musical” vocal-grooming to vocal “gossip” can be seen as providing a partial explanation for evolution of spoken language, most likely with Homo heidelbergensis 0.5 MYA, but presupposes the capacities accounted for by the other models.
The paper tries to contribute to the general discussion on discourse connectives, concretely to the question whether it is meaningful to distinguish two separate groups of connectives -i.e."classical" connectives limited to few predefined classes like conjunctions or adverbs (e.g.but) vs. alternative lexicalizations of connectives (i.e.unrestricted expressions and phrases like the reason is, he added, the condition was etc.).In this respect, the paper focuses on one group of these broader connectives in Czech -the selected verbs of saying doplnit/doplňovat (to complement), upřesnit/upřesňovat (to specify), dodat/dodávat (to add), pokračovat (to continue) -and analyses their occurrence and function in texts from the Prague Discourse Treebank.The paper demonstrates that these verbs of saying have a special place within the other connectives, as they contain two items -e.g. he added means and he said so the verb to add contains an information about the relation to the previous context (and) plus the verb of saying (to say).This information led us to a more general observation, i.e. discourse connectives in broader sense do not necessarily connect two pieces of a text but some of them carry the second argument right in their semantics, which "classical" connectives can never do.
Sentiment analysis of short texts such as single sentences and Twitter messages is challenging because of the limited contextual information that they normally contain. Effectively solving this task requires strategies that combine the small text content with prior knowledge and use more than just bag-of-words. In this work we propose a new deep convolutional neural network that exploits from characterto sentence-level information to perform sentiment analysis of short texts. We apply our approach for two corpora of two different domains: the Stanford Sentiment Treebank (SSTb), which contains sentences from movie reviews; and the Stanford Twitter Sentiment corpus (STS), which contains Twitter messages. For the SSTb corpus, our approach achieves state-of-the-art results for single sentence sentiment prediction in both binary positive/negative classification, with 85.7% accuracy, and fine-grained classification, with 48.3% accuracy. For the STS corpus, our approach achieves a sentiment prediction accuracy of 86.4%.
The conceptions of a linguistic norm for non-linguists form the focus of this article. These conceptions manifest themselves as mental values in the speakers’ minds and can be viewed as implicit reference values of linguistic judgements. There will hereby be an attempt to reconstruct the conceptual contours of people without any academic background in linguistics and to illustrate which criteria play a role in the assessments given and which structural domains can even be assessed. With respect to the method employed here, the study builds on the qualitative content analysis used for extracting and interpreting data. The result of which forms a categorical system that has been constructed deductively and rechecked inductively by the material. The system of categories constructed, which is empirically based on 56 qualitative interviews, also forms the data’s interpretative framework. With the aid of the acquired results, it is shown that non-linguists certainly have a clear picture of what constitutes a good language or how a good language should be; it is closely oriented on written language, invariant, understandable, and has a high communicative scope. In this way, a complex and multilayered conception of linguistic norms can be established. The contours of which will be sketched in this article.
Languages that have no explicit word de-limiters often have to be segmented for sta-tistical machine translation (SMT). This is commonly performed by automated seg-menters trained on manually annotated corpora. However, the word segmentation (WS) schemes of these annotated corpora are handcrafted for general usage, and may not be suitable for SMT. An analysis was performed to test this hypothesis us-ing a manually annotated word alignment (WA) corpus for Chinese-English SMT. An analysis revealed that 74.60 % of the sentences in the WA corpus if segmented using an automated segmenter trained on the Penn Chinese Treebank (CTB) will contain conflicts with the gold WA an-notations. We formulated an approach based on word splitting with reference to the annotated WA to alleviate these con-flicts. Experimental results show that the refined WS reduced word alignment error rate by 6.82 % and achieved the highest BLEU improvement (0.63 on average) on the Chinese-English open machine trans-lation (OpenMT) corpora compared to re-lated work. 1
We present a new dependency parsing method for Korean applying cross-lingual transfer learning and domain adaptation techniques. Unlike existing transfer learning methods relying on aligned corpora or bilingual lexicons, we propose a feature transfer learning method with minimal supervision, which adapts an existing parser to the target language by transferring the features for the source language to the target language. Specifically, we utilize the Triplet/Quadruplet Model, a hybrid parsing algorithm for Japanese, and apply a delexicalized feature transfer for Korean. Experiments with Penn Korean Treebank show that even using only the transferred features from Japanese achieves a high accuracy (81.6%) for Korean dependency parsing. Further improvements were obtained when a small annotated Korean corpus was combined with the Japanese training corpus, confirming that efficient crosslingual transfer learning can be achieved without expensive linguistic resources.
The global spread of English and the advent of a need for English as an International Language has become one of the hotly-debated issues in recent years. This owes much to the fact that English speakers today are more likely to be non-native speakers of English than native speakers, and most likely to use English in communication with other non-native speakers of English than native speakers. A significant number of scholars (e.g., Honna, 2003; Widdowson, 2003) even believe that English is no longer the sole property of its native speakers. Nevertheless, majority of English language teaching coursebooks are still being published by major Anglo-American publishers and are based on the linguistic norms and cultures of native English speaking countries, mainly the USA and the UK. Inevitably, criticism regarding an accurate presentation of cultural information and images about a variety of norms and cultures beyond the Anglo-Saxon and European world has risen. In fact, the English presented in these coursebooks has been seen as mainly representing the linguistic norms and culture of its native speakers, thereby offering ‘English of Specific Cultures’. The current discussions on the English language teaching and culture axis, however, make possible an understanding of an English language that has become first international and then global, thereby creating possibilities of portrayal of linguistic norms and cultures of Outer and Expanding circle countries especially through ELT coursebooks. Commissioned as such, then, English can be regarded as a language through which access to Englishes and cultures of the world accompanies its pedagogy, hence ‘ English for Specific Cultures’ (Yano, 2009). Discussing at length the role of English as an International Language and its cultural implications, this article investigates the varieties of Englishes in a series of EIL-based coursebooks, inquiring whether they are based on English of Specific Cultures or English for Specific Cultures.
What factors contribute to subjective experiences of familiarity, and are these subject to unconscious selection? We investigated the circumstances under which judgments of familiarity are sensitive to task-irrelevant sources using the artificial grammar learning paradigm, a task known to be heavily reliant on familiarity-based responding. In 2 experiments, we manipulated ‘free-floating feelings of familiarity’ by subliminally priming participants with either a subjectively familiar stimulus (their surname) or unfamiliar stimulus (a random letter string). In Experiment 1, after training on an artificial grammar, participants were required to rate the familiarity of a new set of grammar strings where the subliminal priming manipulation preceded each rating. Under these instructions the manipulation significantly altered ratings of familiarity. In Experiment 2, the training, the request for familiarity ratings, and the subliminal manipulation were all unchanged. In addition, however, participants were informed about the presence of rules dictating the structure of the training strings and were required to judge both whether each test-string conformed to those rules and to report the basis for their judgment. This broader decision context eliminated the effect of subliminal primes on ratings of familiarity even when participants’ reported basis for their judgments revealed no conscious knowledge of the rule structure. These results demonstrate that unconscious sources of familiarity can be selected or excluded according to conscious task contexts. The findings are incompatible with theories that equate familiarity with automaticity and those that state people must always be aware of the structural antecedents of metacognition.
Part-of-speech (POS) taggers can be quite accurate, but for practical use, accuracy often has to be sacrificed for speed. For example, the maintainers of the Stanford tagger (Toutanova et al., 2003; Manning, 2011) recommend tagging with a model whose per tag error rate is 17% higher, relatively, than their most accurate model, to gain a factor of 10 or more in speed. In this paper, we treat POS tagging as a single-token independent multiclass classification task. We show that by using a rich feature set we can obtain high tagging accuracy within this framework, and by employing some novel feature-weight-combination and hypothesis-pruning techniques we can also get very fast tagging with this model. A prototype tagger implemented in Perl is tested and found to be at least 8 times faster than any publicly available tagger reported to have comparable accuracy on the standard Penn Treebank Wall Street Journal test set.
For languages such as English, several constituent-to-dependency conversion schemes are pro-posed to construct corpora for dependency parsing. It is hard to determine which scheme is better because they reflect different views of dependency analysis. We usually obtain dependen-cy parsers of different schemes by training with the specific corpus separately. It neglects the correlations between these schemes, which can potentially benefit the parsers. In this paper, we study how these correlations influence final dependency parsing performances, by proposing a joint model which can make full use of the correlations between heterogeneous dependencies, and finally we can answer the following question: parsing heterogeneous dependencies jointly or separately, which is better? We conduct experiments with two different schemes on the Penn Treebank and the Chinese Penn Treebank respectively, arriving at the same conclusion that joint-ly parsing heterogeneous dependencies can give improved performances for both schemes over the individual models.
The authors focus on how to segment semantic units in Chinese discourse and how to label relations among semantic units automatically. During the parsing process, several sequence labelling methods are compared for discourse segmentation, while a maximum entropy-based training and decoding algorithm is specially proposed. Experiments are done based on Tsinghua Chinese Treebank, which is annotated with logical and semantic relations at complex-sentence level. Experimental results show that F-score of discourse segmentation reaches 89.1%. When parsing discourses with no more than 6 relations included, the labeling F-score can achieve 63%.
Resumen Este artículo analiza las actitudes lingüísticas de hablantes nativos de español de la Ciudad Autónoma de Buenos Aires, hacia al español de la Argentina y el español de los otros países hispanohablantes. El artículo es parte de los resultados del Proyecto LIAS (Linguistic Identity and Attitudes in Spanish-speaking Latin America), financiado por El Consejo Noruego de Investigación (RCN). La recolección de los datos se realizó en la capital del país, entrevistando a una muestra de 400 informantes previamente estratificada con las variables de edad, sexo y nivel socioeconómico. El procesamiento estadístico de los datos de campo recolectados arrojó resultados de interés en torno a la mayoría de los tópicos analizados y especialmente en lo referente a aspectos tales como la valoración positiva de la propia variedad lingüística; la resistencia a identificar a España como la única fuente de la norma lingüística de la lengua española; el rechazo a la unificación de la lengua y, por consiguiente, la defensa de la diversidad lingüística como portadora de riqueza cultural. Abstract This article analyzes the linguistic attitudes of native Spanish speakers from Buenos Aires City, towards Spanish spoken in Argentina and in the other Spanish-speaking countries. It is a result of the LIAS-Project (Linguistic Identity and Attitudes in Spanish-speaking Latin America), funded by The Research Council of Norway (RCN). The data were gathered in the capital of the country, interviewing a stratified sample of 400 respondents, based on the variables of age, sex and socioeconomic status. The analysis of the data rendered interesting results on most of the analyzed topics; especially important was the positive appraisal of Argentineans' own linguistic variety; the strong resistance against identifying Spain as the only source of the linguistic norm for the Spanish language; and the rejection of language unification, defending in this way linguistic diversity as an important conveyor of cultural richness.
We propose a novel approach for learning image representation based on qualitative assessments of visual aesthetics. It relies on a multi-node multi-state model that represents image attributes and their relations. The model is learnt from pair wise image preferences provided by annotators. To demonstrate the effectiveness we apply our approach to fashion image rating, i.e., comparative assessment of aesthetic qualities. Bag-of-features object recognition is used for the classification of visual attributes such as clothing and body shape in an image. The attributes and their relations are then assigned learnt potentials which are used to rate the images. Evaluation of the representation model has demonstrated a high performance rate in ranking fashion images.
The well-established memory bias for arousing-negative stimuli seems to be enhanced in high trait-anxious persons and persons suffering from anxiety disorders. We monitored the emergence and development of such a bias during and after learning, in high and low trait anxious participants. A word-learning paradigm was applied, consisting of spoken pseudowords paired either with arousing-negative or neutral pictures. Learning performance during training evidenced a short-lived advantage for arousing-negative associated words, which was not present at the end of training. Cued recall and valence ratings revealed a memory bias for pseudowords that had been paired with arousing-negative pictures, immediately after learning and two weeks later. This held even for items that were not explicitly remembered. High anxious individuals evidenced a stronger memory bias in the cued-recall test, and their ratings were also more negative overall compared to low anxious persons. Both effects were evident, even when explicit recall was controlled for. Regarding the memory bias in anxiety prone persons, explicit memory seems to play a more crucial role than implicit memory. The study stresses the need for several time points of bias measurement during the course of learning and retrieval, as well as the employment of different measures for learning success.
Search engines have become the main way for people to get expected information, most of them are based on keyword search. However, keyword search is based on computing the similarity of letters of the keywords, instead of semantic meaning, therefore the searching results often include irrelevant information to user intention. This paper aims to find a way on improving keyword search efficiency. Using Wikipedia, which is the largest online encyclopedia, this paper explores the relations of terms through computing the semantic relatedness between words, and presents an algorithm called WLA in the light of link structure and text message in Wikipedia. What is more, we design a terms query platform through which users will be able to get all the meanings about the concepts. By making a comparison with lexical database WordNet, it has demonstrated the feasibility on our methods.
Automatically acquiring semantic verb classes from corpora is a challenging task, especially with no existing treebank. Building a high-performing parser for a language is still crucially depends on the existence of large, in-domain texts as training data. While previous work has focused primarily on major languages, how to extend these results to other languages is the way to avoid working start from scratch. In general, a large monolingual corpus in a resource-rich source language labeled with lexico-syntactic information, and a very limited bilingual corpus are available. This paper addresses the problem of verb classification automatically in Tibetan using bilingual lexicon and translation information.
Dependency parsers, which are widely used in natural language processing tasks, employ a representation of syntax in which the structure of sentences is expressed in the form of directed links (dependencies) between their words. In this article, we introduce a new approach to transition‐based dependency parsing in which the parsing algorithm does not directly construct dependencies, but rather undirected links, which are then assigned a direction in a postprocessing step. We show that this alleviates error propagation, because undirected parsers do not need to observe the single‐head constraint, resulting in better accuracy. Undirected parsers can be obtained by transforming existing directed transition‐based parsers as long as they satisfy certain conditions. We apply this approach to obtain undirected variants of three different parsers (the Planar, 2‐Planar, and Covington algorithms) and perform experiments on several data sets from the CoNLL‐X shared tasks and on the Wall Street Journal portion of the Penn Treebank, showing that our approach is successful in reducing error propagation and produces improvements in parsing accuracy in most of the cases and achieving results competitive with state‐of‐the‐art transition‐based parsers.
In recent years, the emergence of English as an International Language (EIL) has paved the way for its global speakers to use it as a means of interacting globally, and representing themselves and their cultures internationally. Although English is globally considered as an international language and as a tool to be used in cross-cultural communication with people having various first languages from different parts of the world, native-speakers’ norms and cultures still dominate the language materials that are developed to be globally used. In fact, English language coursebooks insists on bombarding the ELT world with culturally-loaded native-speaker themes, such as actors in Hollywood (Coskun, 2009). Prodromou (1988) similarly underlines the issue that the majority of English language coursebooks are published by major Anglo-American publishers in Inner Circle countries and these coursebooks include cultural situations that most students will never come across, such as ‘finding a flat in London’ (p. 80). Considering the importance given to the growing role of EIL, the issue of linguistic norms and cultural content in language learning materials has remained one of the unresolved problems in the process of materials development. A group of scholars argues in favor of localizing the materials by using the learners’ experiences and making English language coursebooks culturally responsive to their needs. The opponents solely favor the integration of the linguistic and cultural norms of the native speakers of English in language learning materials. As far as EIL is concerned, there are several aspects that need to be taken into close account when language teaching materials are being prepared to be globally used. In a nutshell, in EIL era, while preparing English language coursebooks, rather than just integrating English of Specific Cultures, the linguistic and cultural norms of the native speakers of English, as the sole reference in the contents of the English language coursebook, at least a due attention should be paid to English for Specific Cultures, the linguistifc and cultural norms of non-native speakers of English. This study recommends a group of essential features for the future English language coursebooks in EIL era.
Abstract Discourse parsing has become an inevitable task to process information in the natural language processing arena. Parsing complex discourse structures beyond the sentence level is a significant challenge. This article proposes a discourse parser that constructs rhetorical structure (RS) trees to identify such complex discourse structures. Unlike previous parsers that construct RS trees using lexical features, syntactic features and cue phrases, the proposed discourse parser constructs RS trees using high‐level semantic features inherited from the Universal Networking Language (UNL). The UNL also adds a language‐independent quality to the parser, because the UNL represents texts in a language‐independent manner. The parser uses a naive Bayes probabilistic classifier to label discourse relations. It has been tested using 500 Tamil‐language documents and the Rhetorical Structure Theory Discourse Treebank, which comprises 21 English‐language documents. The performance of the naive Bayes classifier has been compared with that of the support vector machine (SVM) classifier, which has been used in the earlier approaches to build a discourse parser. It is seen that the naive Bayes probabilistic classifier is better suited for discourse relation labeling when compared with the SVM classifier, in terms of training time, testing time, and accuracy.
Abstract: Similarity is criteria of measuring nearness or proximity between two concepts. Several algorithmic approaches for computing similarity have been proposed. Among the existing Similarity measure, majority of them utilize WordNet as an underlying ontology for calculating semantic similarity. WordNet is a lexical database for English Language which was created and maintained by Congnitive Science Laboratory at Princeton University under the supervision of Professor George A. Miller. It is organized as a network which consists of concepts or terms called Synsets (list of synonyms terms) and the relationship between them. There are different type of relationship exists in WordNet such as is-a, part-of, synonym and antonym. It has thdatabases, one for noun, one for verb and one for adverb and adjective. This project work proposes a metric for semantic relatedness calculation between pair of concepts which uses Tversky’s feature based approach which takes into account the common and distinct feature of the two terms or concepts. If commonality is more as compared to differences the similarity between concepts is high otherwise similarity is low. Tversky’s theory is quantified by information content of two concepts and the Information content of most specific common ancestor of two concepts. As we move down in the WordNet hierarchy, more specific and more Informative concept are there, where as when we move up in the hierarchy more Generalized and less Informative concepts are there. So depth of a concept in the WordNet hierarchy is a critical factor in similarity calculation. We take into consideration the depth of the specific concept in the WordNet hierarchy which is the deciding factor for determining the relevance of distinct feature specific to a concept in similarity calculation. Introduction of depth reduces the impact of the less relevant dissimilarity indulge in similarity calculation thereby increase precision. We carried out our experiment of 28
The present contribution represents the first step in comparing the nature of syntactico-semantic relations present in the sentence structure to their equivalents in the discourse structure. The study is carried out on the basis of Czech manually annotated material collected in the Prague Dependency Treebank (PDT). According to the analysis of the underlying syntactic structure of a sentence (tectogrammatics) in the PDT, we distinguish various types of relations that can be expressed both within a single sentence (i.e. in a tree) and in a larger text, beyond the sentence boundary (between trees). We suggest that, on the one hand, semantic nature of each type of these relations corresponds both within a sentence and in a larger text (i.e. a causal relation remains a causal relation) but, on the other hand, according to the semantic properties of the relations, their distribution in a sentence or between sentences is very diverse. In this study, this observation is analyzed in detail for three cases (relations of condition, specification and opposition ) and further supported by similar behaviour of the English data from the Penn Discourse Treebank.
In current web applications, more businesses are gradually publishing their business as services over the web. This growing number of web services available within an organization and on the Web raises a new and challenging search problem: locating desired web services. Searching for web services with conventional web search engines is insufficient in this context. Automatically clustering Web Service Description Language (WSDL) files on the web into functionally similar homogeneous service groups can be seen as a bootstrapping step for creating a service search engine and, at the same time, reducing the search space for service discovery. In order to overcome some the limitations of pattern-matching approach, the proposed work uses two semantic approaches to cluster similar services. An experimental study based on an information retrieval technique known as latent semantic analysis is applied to the collection of WSDL files and the another semantic approach is based on WordNet which is a lexical database to cluster similar services, as a predecessor step to retrieve the relevant Web services for a user request by search engines. The baseline approach and the two approaches based on semantic is applied on a collection of WSDL documents consisting of to test the quality of clusters formed. As a result, WordNet based approach for clustering shows better cluster quality.
The present sociophonetic study examines the English variety in Michigan's Upper Peninsula (UP) based upon a 130-speaker sample from Marquette County. The linguistic variables of interest include seven monophthongs and four diphthongs: 1) front lax, 2) low back, and 3) high back monophthongs and 4) short and 5) long diphthongs. The sample is stratified by the predictor variables of heritage-location, bilingualism, age, sex and class. The aim of the thesis is two fold: 1) to determine the extent of potential substrate effects on a 71-speaker older-aged bilingual and monolingual subset of these UP English speakers focusing on the predictor variables of heritage-location and bilingualism, and 2) to determine the extent of potential exogenous influences on an 85-speaker subset of UP English monolingual speakers by focusing on the predictor variables of heritage-location, age, sex and class. All data were extracted from a reading passage task collected during a sociolinguistic interview and measured instrumentally. The findings of this apparent-time data reveal the presence of lingering effects from substrate sources and developing effects from exogenous sources based upon American and Canadian models of diffusion. The linguistic changes-in-progress from above, led by middle-class females, are taking shape in the speech of UP residents of whom are propagating linguistic phenomena typically associated with varieties of Canadian English (i.e., low-back merger, Canadian shift, and Canadian raising); however, the findings also report resistance of such norms by working-class females. Finally, the data also reveal substrate effects demonstrating cases of dialect leveling and maintenance. As a result, the speech spoken in Michigan's Upper Peninsula can presently be described as a unique variety of English comprised of lingering substrate effects as well as exogenous effects modeled from both American and Canadian English linguistic norms.
Taboada et al. (2008) propose a word-based method for extracting sentiment from text that relies on the most relevant parts of a text. The method predicts that opinion words found in the nuclei (more important parts) of a document are more significant for the overall sentiment, whereas opinion words found in the satellites (less important parts) only potentially interfere with the overall sentiment. However, as pointed out by Taboada et al. (2008) and Narayanan et al. (2009), for certain discourse relations (for instance, Condition relations), the calculation of sentiment should involve both parts of the relation. Based on our analysis of the affective content expressed by automatically extracted discourse relations from the Simon Fraser University Corpus (Taboada 2008) and the Penn Discourse Treebank (Prasad et al. 2008), we propose to classify all the discourse relations into four categories: (1) relations that reverse polarity, (2) intensify polarity, (3) downtone polarity, or (4) produce no change in polarity. We compare the performance of a sentiment analysis system (SO-CAL, Taboada et al. 2011) when opinion words are detected only in the nuclei with its performance when both parts of the relation are analyzed in combination with the opinion words. The results of the experiment show that extraction of both the nucleus and the satellite parts of texts does not improve the performance of a sentiment extraction system.
The syntactic ambiguity of a transitive verb (Vt) followed by a noun (N) has long been a problem in Chinese parsing. In this paper, we propose a classifier to resolve the ambiguity of Vt-N structures. The design of the classifier is based on three important guidelines, namely, adopting linguistically motivated features, using all available resources, and easy in-tegration into a parsing model. The lin-guistically motivated features include semantic relations, context, and morpho-logical structures; and the available re-sources are treebank, thesaurus, affix da-tabase, and large corpora. We also pro-pose two learning approaches that resolve the problem of data sparseness by auto-parsing and extracting relative knowledge from large-scale unlabeled data. Our experiment results show that the Vt-N classifier outperforms the cur-rent PCFG parser. Furthermore, it can be easily and effectively integrated into the PCFG parser and general statistical pars-ing models. Evaluation of the learning approaches indicates that world knowledge facilitates Vt-N disambigua-tion through data selection and error cor-rection. 1
We describe a contextual parser for the Robot Commands Treebank, a new crowdsourced resource. In contrast to previous semantic parsers that select the most-probable parse, we consider the different problem of parsing using additional situational context to disambiguate between different readings of a sentence. We show that multiple semantic analyses can be searched using dynamic programming via interaction with a spatial planner, to guide the parsing process. We are able to parse sentences in near linear-time by ruling out analyses early on that are incompatible with spatial context. We report a 34% upper bound on accuracy, as our planner correctly processes spatial context for 3,394 out of 10,000 sentences. However, our parser achieves a 96.53% exact-match score for parsing within the subset of sentences recognized by the planner, compared to 82.14% for a non-contextual parser.