Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article analyses the structure of Yoruba numerals and their derivation. Data are collected from the compilation of Yoruba numerals and observation of its use coupled with the researcher's intuitive knowledge of the language. The work dwells on the existing literature on numerals too. The author adopts a descriptive method in analysing the data. The work looks at the roles of affixes in realising odd numbers, multiples of 20, centenary, bicentenary, and so on in their order of increase. It is discovered that the direction of counting in Yoruba is largely progressive. Besides, the language adopts base 5, decimal (base 10) and vigesimal (base 20) systems of counting. It is equally discovered that the choice of either of the two variations is largely dependent on the articulatory parameter of the first vowel (V1) of the root word. It is noted that the Yoruba numeral system offers a suitable linguistic database for both the theoretical and empirical domains of linguistic study especially documentary linguistics. The current study has general pedagogic implications for the teaching and learning of Yoruba numerals.
This paper presents results of dependency parsing of Old French, a language which is poorly standardized at the lexical level, and which displays a relatively free word order. The work is carried out on five distinct sample texts extracted from the dependency treebank Syntactic Reference Corpus of Medieval French (SRCMF). Following Achim Stein's previous work, we have trained the Mate parser on each sub-corpus and cross-validated the results. We show that the parsing efficiency is diminished by the greater lexical variation of Old French compared to parse results on modern French. In order to improve the result of the POS tagging step in the parsing process, we applied a pre-treatment to the data, comparing two distinct strategies: one using a slightly post-treated version of the TreeTagger trained on Old French by Stein, and a CRF trained on the texts, enriched with external resources. The CRF version outperforms every other approach.
Abstract: Similarity is criteria of measuring nearness or proximity between two concepts. Several algorithmic approaches for computing similarity have been proposed. Among the existing Similarity measure, majority of them utilize WordNet as an underlying ontology for calculating semantic similarity. WordNet is a lexical database for English Language which was created and maintained by Congnitive Science Laboratory at Princeton University under the supervision of Professor George A. Miller. It is organized as a network which consists of concepts or terms called Synsets (list of synonyms terms) and the relationship between them. There are different type of relationship exists in WordNet such as is-a, part-of, synonym and antonym. It has thdatabases, one for noun, one for verb and one for adverb and adjective. This project work proposes a metric for semantic relatedness calculation between pair of concepts which uses Tversky’s feature based approach which takes into account the common and distinct feature of the two terms or concepts. If commonality is more as compared to differences the similarity between concepts is high otherwise similarity is low. Tversky’s theory is quantified by information content of two concepts and the Information content of most specific common ancestor of two concepts. As we move down in the WordNet hierarchy, more specific and more Informative concept are there, where as when we move up in the hierarchy more Generalized and less Informative concepts are there. So depth of a concept in the WordNet hierarchy is a critical factor in similarity calculation. We take into consideration the depth of the specific concept in the WordNet hierarchy which is the deciding factor for determining the relevance of distinct feature specific to a concept in similarity calculation. Introduction of depth reduces the impact of the less relevant dissimilarity indulge in similarity calculation thereby increase precision. We carried out our experiment of 28
Constituent Context Model (CCM) is an effective generative model for grammar induction, the aim of which is to induce hierarchical syntactic structure from natural text. The CCM simply defines the Multinomial distribution over constituents, which leads to a severe data sparse problem because long constituents are unlikely to appear in unseen data sets. This paper proposes a Bayesian method for constituent smoothing by defining two kinds of prior distributions over constituents: the Dirichlet prior and the Pitman-Yor Process prior. The Dirichlet prior functions as an additive smoothing method, and the PYP prior functions as a back-off smoothing method. Furthermore, a modified CCM is proposed to differentiate left constituents and right constituents in binary branching trees. Experiments show that both the proposed Bayesian smoothing method and the modified CCM are effective, and combining them attains or significantly improves the state-of-the-art performance of grammar induction evaluated on standard treebanks of various languages.
Abstract Discourse parsing has become an inevitable task to process information in the natural language processing arena. Parsing complex discourse structures beyond the sentence level is a significant challenge. This article proposes a discourse parser that constructs rhetorical structure (RS) trees to identify such complex discourse structures. Unlike previous parsers that construct RS trees using lexical features, syntactic features and cue phrases, the proposed discourse parser constructs RS trees using high‐level semantic features inherited from the Universal Networking Language (UNL). The UNL also adds a language‐independent quality to the parser, because the UNL represents texts in a language‐independent manner. The parser uses a naive Bayes probabilistic classifier to label discourse relations. It has been tested using 500 Tamil‐language documents and the Rhetorical Structure Theory Discourse Treebank, which comprises 21 English‐language documents. The performance of the naive Bayes classifier has been compared with that of the support vector machine (SVM) classifier, which has been used in the earlier approaches to build a discourse parser. It is seen that the naive Bayes probabilistic classifier is better suited for discourse relation labeling when compared with the SVM classifier, in terms of training time, testing time, and accuracy.
The Stuttgart-Tbingen Tagset (STTS) is a widely used POS annotation scheme for German which provides 54 different tags for the analysis on the part of speech level.The tagset, however, does not distinguish between adverbs and different types of particles used for expressing modality, intensity, graduation, or to mark the focus of the sentence.In the paper, we present an extension to the STTS which provides tags for a more fine-grained analysis of modification, based on a syntactic perspective on parts of speech.We argue that the new classification not only enables us to do corpus-based linguistic studies on modification, but also improves statistical parsing.We give proof of concept by training a data-driven dependency parser on data from the TiGer treebank, providing the parser a) with the original STTS tags and b) with the new tags.Results show an improved labelled accuracy for the new, syntactically motivated classification.
Recurrent neural network language models have solved the problems of data sparseness and dimensionality disaster which exist in traditional N-gram models. RNNLMs have recently demonstrated state-of-the-art performance in speech recognition, machine translation and other tasks. In this paper, we improve the model performance by providing contextual word vectors in association with RNNLMs. This method can reinforce the ability of learning long-distance information using vectors training from Skip-gram model. The experimental results show that the proposed method can improve the perplexity performance significantly on Penn Treebank data. And we further apply the models to speech recognition task on the Wall Street Journal corpora, where we achieve obvious improvements in word-error-rate.
This paper presents a set of Bilingual Dictionary Drafting (BDD) methods including manual extraction from existing lexical databases and corpus based NLP tools, as well as their evaluation on the example of German-Basque as language pair. Our aim is twofold: to give support to a German-Basque bilingual dictionary project by providing draft Bilingual Glossaries and to provide lexicographers with insight into how useful BDD methods are. Results show that the analysed methods can greatly assist on bilingual dictionary writing, in the context of medium-density language pairs.
The lexicon dynamics deals with the changes that occur in language from a historical stage to another. As part of a living organism, words emerge and fade away, bearing the mark of the linguistic norms. The linguists’ preoccupation with the accuracy of language goes back in time and is related to lexicon, semantics, grammar, spelling etc. The present study is designed as a concise presentation of certain mistakes frequently met in the current written press, with a view to correcting them. Some deviations from the norms are minor, others major, whereas their causes are numerous. Recent loans, neologisms, confusion of styles and excessive use of clichés are only a few aspects to be approached in the present work.
In recent years, the emergence of English as an International Language (EIL) has paved the way for its global speakers to use it as a means of interacting globally, and representing themselves and their cultures internationally. Although English is globally considered as an international language and as a tool to be used in cross-cultural communication with people having various first languages from different parts of the world, native-speakers’ norms and cultures still dominate the language materials that are developed to be globally used. In fact, English language coursebooks insists on bombarding the ELT world with culturally-loaded native-speaker themes, such as actors in Hollywood (Coskun, 2009). Prodromou (1988) similarly underlines the issue that the majority of English language coursebooks are published by major Anglo-American publishers in Inner Circle countries and these coursebooks include cultural situations that most students will never come across, such as ‘finding a flat in London’ (p. 80). Considering the importance given to the growing role of EIL, the issue of linguistic norms and cultural content in language learning materials has remained one of the unresolved problems in the process of materials development. A group of scholars argues in favor of localizing the materials by using the learners’ experiences and making English language coursebooks culturally responsive to their needs. The opponents solely favor the integration of the linguistic and cultural norms of the native speakers of English in language learning materials. As far as EIL is concerned, there are several aspects that need to be taken into close account when language teaching materials are being prepared to be globally used. In a nutshell, in EIL era, while preparing English language coursebooks, rather than just integrating English of Specific Cultures, the linguistic and cultural norms of the native speakers of English, as the sole reference in the contents of the English language coursebook, at least a due attention should be paid to English for Specific Cultures, the linguistifc and cultural norms of non-native speakers of English. This study recommends a group of essential features for the future English language coursebooks in EIL era.
Dependency parsers, which are widely used in natural language processing tasks, employ a representation of syntax in which the structure of sentences is expressed in the form of directed links (dependencies) between their words. In this article, we introduce a new approach to transition‐based dependency parsing in which the parsing algorithm does not directly construct dependencies, but rather undirected links, which are then assigned a direction in a postprocessing step. We show that this alleviates error propagation, because undirected parsers do not need to observe the single‐head constraint, resulting in better accuracy. Undirected parsers can be obtained by transforming existing directed transition‐based parsers as long as they satisfy certain conditions. We apply this approach to obtain undirected variants of three different parsers (the Planar, 2‐Planar, and Covington algorithms) and perform experiments on several data sets from the CoNLL‐X shared tasks and on the Wall Street Journal portion of the Penn Treebank, showing that our approach is successful in reducing error propagation and produces improvements in parsing accuracy in most of the cases and achieving results competitive with state‐of‐the‐art transition‐based parsers.
The article is devoted to the problem of translation from French into Polish from the perspective of the object oriented approach proposed by Wiesław Banyś. The author takes into consideration some of the problems, both of theoretical and practical nature, which appear while working on the formation of contrastive lexical database for automatic translation of texts. While analyzing specific chosen examples which may cause different kinds of problems in the description, the author offers their interpretations in the target language in accordance with the adopted approach.
Recent work has sparked new interest in type-supervised part-of-speech tagging, a data setting in which no labeled sen-tences are available, but the set of allowed tags is known for each word type. This paper describes observational initializa-tion, a novel technique for initializing EM when training a type-supervised HMM tagger. Our initializer allocates probabil-ity mass to unambiguous transitions in an unlabeled corpus, generating token-level observations from type-level supervision. Experimentally, observational initializa-tion gives state-of-the-art type-supervised tagging accuracy, providing an error re-duction of 56 % over uniform initialization on the Penn English Treebank. 1
We investigate whether parsers can be used for self-monitoring in surface realization in order to avoid egregious errors involving "vicious" ambiguities, namely those where the intended interpretation fails to be considerably more likely than alternative ones. Using parse accuracy in a simple reranking strategy for selfmonitoring, we find that with a stateof-the-art averaged perceptron realization ranking model, BLEU scores cannot be improved with any of the well-known Treebank parsers we tested, since these parsers too often make errors that human readers would be unlikely to make. However, by using an SVM ranker to combine the realizer's model score together with features from multiple parsers, including ones designed to make the ranker more robust to parsing mistakes, we show that significant increases in BLEU scores can be achieved. Moreover, via a targeted manual analysis, we demonstrate that the SVM reranker frequently manages to avoid vicious ambiguities, while its ranking errors tend to affect fluency much more often than adequacy.
Automatically acquiring semantic verb classes from corpora is a challenging task, especially with no existing treebank. Building a high-performing parser for a language is still crucially depends on the existence of large, in-domain texts as training data. While previous work has focused primarily on major languages, how to extend these results to other languages is the way to avoid working start from scratch. In general, a large monolingual corpus in a resource-rich source language labeled with lexico-syntactic information, and a very limited bilingual corpus are available. This paper addresses the problem of verb classification automatically in Tibetan using bilingual lexicon and translation information.
Search engines have become the main way for people to get expected information, most of them are based on keyword search. However, keyword search is based on computing the similarity of letters of the keywords, instead of semantic meaning, therefore the searching results often include irrelevant information to user intention. This paper aims to find a way on improving keyword search efficiency. Using Wikipedia, which is the largest online encyclopedia, this paper explores the relations of terms through computing the semantic relatedness between words, and presents an algorithm called WLA in the light of link structure and text message in Wikipedia. What is more, we design a terms query platform through which users will be able to get all the meanings about the concepts. By making a comparison with lexical database WordNet, it has demonstrated the feasibility on our methods.
The well-established memory bias for arousing-negative stimuli seems to be enhanced in high trait-anxious persons and persons suffering from anxiety disorders. We monitored the emergence and development of such a bias during and after learning, in high and low trait anxious participants. A word-learning paradigm was applied, consisting of spoken pseudowords paired either with arousing-negative or neutral pictures. Learning performance during training evidenced a short-lived advantage for arousing-negative associated words, which was not present at the end of training. Cued recall and valence ratings revealed a memory bias for pseudowords that had been paired with arousing-negative pictures, immediately after learning and two weeks later. This held even for items that were not explicitly remembered. High anxious individuals evidenced a stronger memory bias in the cued-recall test, and their ratings were also more negative overall compared to low anxious persons. Both effects were evident, even when explicit recall was controlled for. Regarding the memory bias in anxiety prone persons, explicit memory seems to play a more crucial role than implicit memory. The study stresses the need for several time points of bias measurement during the course of learning and retrieval, as well as the employment of different measures for learning success.
We propose a novel approach for learning image representation based on qualitative assessments of visual aesthetics. It relies on a multi-node multi-state model that represents image attributes and their relations. The model is learnt from pair wise image preferences provided by annotators. To demonstrate the effectiveness we apply our approach to fashion image rating, i.e., comparative assessment of aesthetic qualities. Bag-of-features object recognition is used for the classification of visual attributes such as clothing and body shape in an image. The attributes and their relations are then assigned learnt potentials which are used to rate the images. Evaluation of the representation model has demonstrated a high performance rate in ranking fashion images.
Resumen Este artículo analiza las actitudes lingüísticas de hablantes nativos de español de la Ciudad Autónoma de Buenos Aires, hacia al español de la Argentina y el español de los otros países hispanohablantes. El artículo es parte de los resultados del Proyecto LIAS (Linguistic Identity and Attitudes in Spanish-speaking Latin America), financiado por El Consejo Noruego de Investigación (RCN). La recolección de los datos se realizó en la capital del país, entrevistando a una muestra de 400 informantes previamente estratificada con las variables de edad, sexo y nivel socioeconómico. El procesamiento estadístico de los datos de campo recolectados arrojó resultados de interés en torno a la mayoría de los tópicos analizados y especialmente en lo referente a aspectos tales como la valoración positiva de la propia variedad lingüística; la resistencia a identificar a España como la única fuente de la norma lingüística de la lengua española; el rechazo a la unificación de la lengua y, por consiguiente, la defensa de la diversidad lingüística como portadora de riqueza cultural. Abstract This article analyzes the linguistic attitudes of native Spanish speakers from Buenos Aires City, towards Spanish spoken in Argentina and in the other Spanish-speaking countries. It is a result of the LIAS-Project (Linguistic Identity and Attitudes in Spanish-speaking Latin America), funded by The Research Council of Norway (RCN). The data were gathered in the capital of the country, interviewing a stratified sample of 400 respondents, based on the variables of age, sex and socioeconomic status. The analysis of the data rendered interesting results on most of the analyzed topics; especially important was the positive appraisal of Argentineans' own linguistic variety; the strong resistance against identifying Spain as the only source of the linguistic norm for the Spanish language; and the rejection of language unification, defending in this way linguistic diversity as an important conveyor of cultural richness.
The authors focus on how to segment semantic units in Chinese discourse and how to label relations among semantic units automatically. During the parsing process, several sequence labelling methods are compared for discourse segmentation, while a maximum entropy-based training and decoding algorithm is specially proposed. Experiments are done based on Tsinghua Chinese Treebank, which is annotated with logical and semantic relations at complex-sentence level. Experimental results show that F-score of discourse segmentation reaches 89.1%. When parsing discourses with no more than 6 relations included, the labeling F-score can achieve 63%.
For languages such as English, several constituent-to-dependency conversion schemes are pro-posed to construct corpora for dependency parsing. It is hard to determine which scheme is better because they reflect different views of dependency analysis. We usually obtain dependen-cy parsers of different schemes by training with the specific corpus separately. It neglects the correlations between these schemes, which can potentially benefit the parsers. In this paper, we study how these correlations influence final dependency parsing performances, by proposing a joint model which can make full use of the correlations between heterogeneous dependencies, and finally we can answer the following question: parsing heterogeneous dependencies jointly or separately, which is better? We conduct experiments with two different schemes on the Penn Treebank and the Chinese Penn Treebank respectively, arriving at the same conclusion that joint-ly parsing heterogeneous dependencies can give improved performances for both schemes over the individual models.
Part-of-speech (POS) taggers can be quite accurate, but for practical use, accuracy often has to be sacrificed for speed. For example, the maintainers of the Stanford tagger (Toutanova et al., 2003; Manning, 2011) recommend tagging with a model whose per tag error rate is 17% higher, relatively, than their most accurate model, to gain a factor of 10 or more in speed. In this paper, we treat POS tagging as a single-token independent multiclass classification task. We show that by using a rich feature set we can obtain high tagging accuracy within this framework, and by employing some novel feature-weight-combination and hypothesis-pruning techniques we can also get very fast tagging with this model. A prototype tagger implemented in Perl is tested and found to be at least 8 times faster than any publicly available tagger reported to have comparable accuracy on the standard Penn Treebank Wall Street Journal test set.
What factors contribute to subjective experiences of familiarity, and are these subject to unconscious selection? We investigated the circumstances under which judgments of familiarity are sensitive to task-irrelevant sources using the artificial grammar learning paradigm, a task known to be heavily reliant on familiarity-based responding. In 2 experiments, we manipulated ‘free-floating feelings of familiarity’ by subliminally priming participants with either a subjectively familiar stimulus (their surname) or unfamiliar stimulus (a random letter string). In Experiment 1, after training on an artificial grammar, participants were required to rate the familiarity of a new set of grammar strings where the subliminal priming manipulation preceded each rating. Under these instructions the manipulation significantly altered ratings of familiarity. In Experiment 2, the training, the request for familiarity ratings, and the subliminal manipulation were all unchanged. In addition, however, participants were informed about the presence of rules dictating the structure of the training strings and were required to judge both whether each test-string conformed to those rules and to report the basis for their judgment. This broader decision context eliminated the effect of subliminal primes on ratings of familiarity even when participants’ reported basis for their judgments revealed no conscious knowledge of the rule structure. These results demonstrate that unconscious sources of familiarity can be selected or excluded according to conscious task contexts. The findings are incompatible with theories that equate familiarity with automaticity and those that state people must always be aware of the structural antecedents of metacognition.
We propose a new simple but effective method for building Tibetan-Chinese machine Translation corpus and a novel Tibetan-Chinese Machine Translation model integrating Tibetan syntactic cues which is based on the Treebank, this model can be used on the system of Tibetan-Chinese Machine Translation successfully. Keywords: syntactic Treebank; Tibetan syntactic cues; Machine Translation;
The global spread of English and the advent of a need for English as an International Language has become one of the hotly-debated issues in recent years. This owes much to the fact that English speakers today are more likely to be non-native speakers of English than native speakers, and most likely to use English in communication with other non-native speakers of English than native speakers. A significant number of scholars (e.g., Honna, 2003; Widdowson, 2003) even believe that English is no longer the sole property of its native speakers. Nevertheless, majority of English language teaching coursebooks are still being published by major Anglo-American publishers and are based on the linguistic norms and cultures of native English speaking countries, mainly the USA and the UK. Inevitably, criticism regarding an accurate presentation of cultural information and images about a variety of norms and cultures beyond the Anglo-Saxon and European world has risen. In fact, the English presented in these coursebooks has been seen as mainly representing the linguistic norms and culture of its native speakers, thereby offering ‘English of Specific Cultures’. The current discussions on the English language teaching and culture axis, however, make possible an understanding of an English language that has become first international and then global, thereby creating possibilities of portrayal of linguistic norms and cultures of Outer and Expanding circle countries especially through ELT coursebooks. Commissioned as such, then, English can be regarded as a language through which access to Englishes and cultures of the world accompanies its pedagogy, hence ‘ English for Specific Cultures’ (Yano, 2009). Discussing at length the role of English as an International Language and its cultural implications, this article investigates the varieties of Englishes in a series of EIL-based coursebooks, inquiring whether they are based on English of Specific Cultures or English for Specific Cultures.
We present a new dependency parsing method for Korean applying cross-lingual transfer learning and domain adaptation techniques. Unlike existing transfer learning methods relying on aligned corpora or bilingual lexicons, we propose a feature transfer learning method with minimal supervision, which adapts an existing parser to the target language by transferring the features for the source language to the target language. Specifically, we utilize the Triplet/Quadruplet Model, a hybrid parsing algorithm for Japanese, and apply a delexicalized feature transfer for Korean. Experiments with Penn Korean Treebank show that even using only the transferred features from Japanese achieves a high accuracy (81.6%) for Korean dependency parsing. Further improvements were obtained when a small annotated Korean corpus was combined with the Japanese training corpus, confirming that efficient crosslingual transfer learning can be achieved without expensive linguistic resources.
Languages that have no explicit word de-limiters often have to be segmented for sta-tistical machine translation (SMT). This is commonly performed by automated seg-menters trained on manually annotated corpora. However, the word segmentation (WS) schemes of these annotated corpora are handcrafted for general usage, and may not be suitable for SMT. An analysis was performed to test this hypothesis us-ing a manually annotated word alignment (WA) corpus for Chinese-English SMT. An analysis revealed that 74.60 % of the sentences in the WA corpus if segmented using an automated segmenter trained on the Penn Chinese Treebank (CTB) will contain conflicts with the gold WA an-notations. We formulated an approach based on word splitting with reference to the annotated WA to alleviate these con-flicts. Experimental results show that the refined WS reduced word alignment error rate by 6.82 % and achieved the highest BLEU improvement (0.63 on average) on the Chinese-English open machine trans-lation (OpenMT) corpora compared to re-lated work. 1
The conceptions of a linguistic norm for non-linguists form the focus of this article. These conceptions manifest themselves as mental values in the speakers’ minds and can be viewed as implicit reference values of linguistic judgements. There will hereby be an attempt to reconstruct the conceptual contours of people without any academic background in linguistics and to illustrate which criteria play a role in the assessments given and which structural domains can even be assessed. With respect to the method employed here, the study builds on the qualitative content analysis used for extracting and interpreting data. The result of which forms a categorical system that has been constructed deductively and rechecked inductively by the material. The system of categories constructed, which is empirically based on 56 qualitative interviews, also forms the data’s interpretative framework. With the aid of the acquired results, it is shown that non-linguists certainly have a clear picture of what constitutes a good language or how a good language should be; it is closely oriented on written language, invariant, understandable, and has a high communicative scope. In this way, a complex and multilayered conception of linguistic norms can be established. The contours of which will be sketched in this article.
Linguistic norms emerge in human communities because people imitate each other. A shared linguistic system provides people with the benefits of shared knowledge and coordinated planning. Once norms are in place, why would they ever change? This question, echoing broad questions in the theory of social dynamics, has particular force in relation to language. By definition, an innovator is in the minority when the innovation first occurs. In some areas of social dynamics, important minorities can strongly influence the majority through their power, fame, or use of broadcast media. But most linguistic changes are grassroots developments that originate with ordinary people. Here, we develop a novel model of communicative behavior in communities, and identify a mechanism for arbitrary innovations by ordinary people to have a good chance of being widely adopted. To imitate each other, people must form a mental representation of what other people do. Each time they speak, they must also decide which form to produce themselves. We introduce a new decision function that enables us to smoothly explore the space between two types of behavior: probability matching (matching the probabilities of incoming experience) and regularization (producing some forms disproportionately often). Using Monte Carlo methods, we explore the interactions amongst the degree of regularization, the distribution of biases in a network, and the network position of the innovator. We identify two regimes for the widespread adoption of arbritrary innovations, viewed as informational cascades in the network. With moderate regularization of experienced input, average people (not well-connected people) are the most likely source of successful innovations. Our results shed light on a major outstanding puzzle in the theory of language change. The framework also holds promise for understanding the dynamics of other social norms.
OBJECTIVE: Premenstrual dysphoric disorder (PMDD) is associated with increased pain, but there has been a lack of well-controlled research assessing pain responsivity, sex hormones, and their relationships in this group. This study was designed to address this gap in the literature. MATERIALS AND METHODS: Healthy, regularly cycling participants (14 PMDD, 14 non-PMDD) attended pain testing sessions during the mid-follicular, ovulatory, and late-luteal phases of the menstrual cycle (order counterbalanced) and salivary estradiol, progesterone, and testosterone were assessed at each testing session. Pain sensitivity was measured from electrocutaneous threshold/tolerance, ischemic threshold/tolerance, sensory and affective ratings of electrocutaneous and ischemic stimuli, and the nociceptive flexion reflex threshold (NFR, a measure of spinal nociception). RESULTS: Women with PMDD had higher sensory pain ratings of electrocutaneous stimuli and trends for lower ischemic thresholds and higher affective pain ratings of electrocutaneous stimuli. However, there were no group differences observed in NFR threshold. Testosterone levels were also lower during the mid-follicular and ovulatory phases in PMDD. Correlations between pain outcomes and estradiol and testosterone indicated that these hormones are hypoalgesic, with estradiol having a greater hypoalgesic effect within the PMDD group. DISCUSSION: Overall, women with PMDD may have a phase-independent hyperalgesia, with pain amplification likely occurring at the supraspinal level rather than the spinal level, given the lack of group differences in NFR threshold. Because testosterone was hypoalgesic and lower in women with PMDD, and there were strong associations between pain and estradiol in PMDD, sex hormones may play a role in PMDD-related hyperalgesia.
This paper describes experiments for statistical dependency parsing using two different parsers trained on a recently extended dependency treebank for Greek, a language with a moderately rich morphology. We show how scores obtained by the two parsers are influenced by morphology and dependency types as well as sentence and arc length. The best LAS obtained in these experiments was 80.16 on a test set with manually validated POS tags and lemmas. 1
BACKGROUND: Careful observation of the longitudinal course of bipolar disorders is pivotal to finding optimal treatments and improving outcome. A useful tool is the daily prospective Life-Chart Method, developed by the National Institute of Mental Health. However, it remains unclear whether the patient version is as valid as the clinician version. METHODS: We compared the patient-rated version of the Lifechart (LC-self) with the Young-Mania-Rating Scale (YMRS), Inventory of Depressive Symptoms-Clinician version (IDS-C), and Clinical Global Impression-Bipolar version (CGI-BP) in 108 bipolar I and II patients who participated in the Naturalistic Follow-up Study (NFS) of the German centres of the Bipolar Collaborative Network (BCN; formerly Stanley Foundation Bipolar Network). For statistical evaluation, levels of severity of mood states on the Lifechart were transformed numerically and comparison with affective scales was performed using chi-square and t tests. For testing correlations Pearson´s coefficient was calculated. RESULTS: Ratings for depression of LC-self and total scores of IDS-C were found to be highly correlated (Pearson coefficient r = -.718; p <.001), whilst the correlation of ratings for mania with YMRS compared to LC-self were slightly less robust (Pearson coefficient r =.491; p =.001). These results were confirmed by good correlations between the CGI-BP IA (mania), IB (depression) and IC (overall mood state) and the LC-self ratings (Pearson coefficient r =.488, r =.721 and r =.65, respectively; all p <.001). CONCLUSIONS: The LC-self shows a significant correlation and good concordance with standard cross sectional affective rating scales, suggesting that the LC-self is a valid and time and money saving alternative to the clinician-rated version which should be incorporated in future clinical research in bipolar disorder. Generalizability of the results is limited by the selection of highly motivated patients in specialized bipolar centres and by the open design of the study.
The sentiment mining approaches can typically be divided into lexicon and machine learning approaches. Recently there are an increasing number of approaches which combine both to improve the performance when used separately. However, this still lacks contextual understanding which led to the introduction of deep learning approaches which allows for semantic compositionality over a sentiment treebank. This paper enhances the deep learning approach with semantic lexicon so that scores can be computed in-stead merely nominal classification. Besides, neutral classification is also improved. Results suggest that the approach outperforms its original.
Automatic text categorisation systems is a type of software that every day it is receiving more interest, due not only to its use in documentaries environments but also to its possible application to tag properly documents on the Web. Many options have been proposed to face this subject using statistical approaches, natural language processing tools, ontologies and lexical databases. Nevertheless, there have been no too many empirical evaluations comparing the influence of the different tools used to solve these problems, particularly in a multilingual environment. In this paper we propose a multi-language rule-based pipeline system for automatic document categorisation and we compare empirically the results of applying techniques that rely on statistics and supervised learning with the results of applying the same techniques but with the support of smarter tools based on language semantics and ontologies, using for this purpose several corpora of documents. GENIE is being applied to real environments, which shows the potential of the proposal.
The web today is huge and enormous collection of data today and it goes on increasing day by day. Thus, searching for some particular data in this collection has a significant impact. Researches taking place give prominence to the relevancy and relatedness of the data that is found. Inspite of their relevance pages for any search topic, the results are still huge to be explored. Another important issue to be kept in mind is the users’ standpoint differs from time to time from topic to topic. Effective relevance prediction can help avoid downloading and visiting many irrelevant pages. The performance of a crawler depends mostly on the opulence of links in the specific topic being searched. This paper reviews the researches on web crawling algorithms used for searching. Keywords— Web Crawling Algorithms, Crawling Algorithm Survey, Search Algorithms, Lexical Database, Metadata, Semantic. __________________________________________________*****_________________________________________________
Abstract We investigated discrimination in the context of evaluating advertisements, based on the suppression model (justification‐suppression model [ JSM ]) of prejudice expression. Previous research has demonstrated that when people are given an opportunity to give a high rating to an ad featuring a Black model, a sense of nonprejudice is created, which, in turn, provides an opportunity to discriminate subsequently without feeling prejudiced. We extended the JSM by investigating whether the acquisition of legitimacy credits (a moral authority earned by demonstrating nonprejudice) is a sufficient condition to release the expression of prejudice. We found that subjects who first evaluated a high‐quality ad featuring a Black model felt eligible to use legitimacy credits in subsequent evaluations. But in a subsequent study, participants who acquired these credits evaluated Black model ads more negatively than White model ads only when these ads were of low quality. The implications for evaluating the subtle way that prejudice affects rating of models of color in advertisements are discussed.
With growing interest in the creation and search of linguistic annotations that form general graphs (in contrast to formally simpler, rooted trees), there also is an increased need for infrastructures that support the exploration of such representations, for example logical-form meaning representations or semantic dependency graphs. In this work, we heavily lean on semantic technologies and in particular the data model of the Resource Description Framework (RDF) to represent, store, and efficiently query very large collections of text annotated with graph-structured representations of sentence meaning. Keywords:Semantic Dependency Graphs, Treebank Search, Resource Description Framework 1.
This paper introduces a new technique for phrase-structure parser analysis, catego-rizing possible treebank structures by inte-grating regular expressions into derivation trees. We analyze the performance of the Berkeley parser on OntoNotes WSJ and the English Web Treebank. This provides some insight into the evalb scores, and the problem of domain adaptation with the web data. We also analyze a “test-on-train ” dataset, showing a wide variance in how the parser is generalizing from differ-ent structures in the training material. 1
Natural language is a fundamental thing of human-society to communicate and interact with one another. In this globalization era, we interact with different regional people as per our interest in social, cultural, economical, educational and professional domain. There are thousands of natural languages exist in our earth. It is quite tough, rather impossible to know all the languages. So we need a computerized approach to convert one natural language to another as per our necessity. This computerized conversion among multiple languages is known as multilingual machine translation. But in this paper we work with a bilingual model, where we concern with two languages: English and Bengali. We use soft computational approach where fuzzy If-Then rule is applied to choose a lemma from prior knowledge; Penn TreeBank PoS tags and HMM tagger are used as lexical class marker to each word in corpora.
This paper investigates the recognition of unknown words in Chinese parsing. Two methods are proposed to handle this problem. One is the modification of a character-based model. We model the emission probability of an unknown word using the first and last characters in the word. It aims to reduce the POS tag ambiguities of unknown words to improve the parsing performance. In addition, a novel method, using graph-based semisupervised learning (SSL), is proposed to improve the syntax parsing of unknown words. Its goal is to discover additional lexical knowledge from a large amount of unlabeled data to help the syntax parsing. The method is mainly to propagate lexical emission probabilities to unknown words by building the similarity graphs over the words of labeled and unlabeled data. The derived distributions are incorporated into the parsing process. The proposed methods are effective in dealing with the unknown words to improve the parsing. Empirical results for Penn Chinese Treebank and TCT Treebank revealed its effectiveness.
Chunking or shallow syntactic parsing is proving to be a task of interest to many natural language processing applications. The problem gets worse for the Arabic language because of its specific features that make it quite different and even more ambiguous than other natural languages when processed. In this paper, we present a method for chunking Arabic texts based on supervised learning. We use the Conditional Random Fields algorithm and the Penn Arabic Treebank to train the model. For the experimentation, we use over than 10,100 sentences as training data and 2,524 sentences for the test. The evaluation of the method consists of the calculation of the generated model accuracy and the results are very encouraging.
One way of teaching grammar, namely morphology and syntax, is to visualize sentences as diagrams capturing relationships between words. Similarly, such relationships are captured in a more complex way in treebanks serving as key building stones in modern natural language processing. However, building them is very time consuming, thus we have been seeking for an alternative cheaper and faster way, like crowdsourcing. The purpose of our work is to explore possibility to get sentence diagrams produced by students and teachers. In our pilot study, the object language is Czech, where sentence diagrams are part of elementary school curriculum.
In this paper, the development and evaluation of the Urdu parser is presented along with the comparison of existing resources for the language variants Urdu/Hindi. This parser was given a linguistically rich grammar extracted from a treebank. This context free grammar with sufficient encoded information is comparable with the state of the art parsing requirements for morphologically rich and closely related language variants Urdu/Hindi. The extended parsing model and the linguistically rich grammar together provide us promising parsing results for both the language variants. The parser gives 87% of f-score, which outperforms the multi-path shift-reduce parser for Urdu and a simple Hindi dependency parser with 4.8% and 22% increase in recall, respectively.