Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Discourse relations bind smaller linguistic elements into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked sentences. A more subtle challenge is that it is not enough to represent the meaning of each sentence of a discourse relation, because the relation may depend on links between lower-level elements, such as entity mentions. Our solution computes distributional meaning representations by composition up the syntactic parse tree. A key difference from previous work on compositional distributional semantics is that we also compute representations for entity mentions, using a novel downward compositional pass. Discourse relations are predicted not only from the distributional representations of the sentences, but also of their coreferent entity mentions. The resulting system obtains substantial improvements over the previous state-of-the-art in predicting implicit discourse relations in the Penn Discourse Treebank.
In this paper we present a system for experimenting with combinations of dependency parsers. The system supports initial training of different parsing models, creation of parsebank(s) with these models, and different strategies for the construction of ensemble models aimed at improving the output of the individual models by voting. The system employs two algorithms for construction of dependency trees from several parses of the same sentence and several ways for ranking of the arcs in the resulting trees. We have performed experiments with state-of-the-art dependency parsers including MaltParser (Nivre et al., 2006), MSTParser (McDonald, 2006), TurboParser (Martins et al., 2010), and MATEParser (Bohnet, 2010), on the data from the Bulgarian treebank – BulTreeBank. Our best result from these experiments is slightly better then the best result reported in the literature for this language (Martins et al.,
Incorporating knowledge for training a parser has been shown to remedy the weaknesses of probabilistic context-free grammar. Previous parsing systems have exploited content words semantic resource and word-formation knowledge. However, they are limited in that they do not take into account conjunction category refinement, which stands out to be helpful in predicting the syntactic structure and syntactic label in Chinese. We define a conjunction taxonomy representing intrinsic syntactic constraints, and show that refined categories in the taxonomy for conjunctions contribute to improved parsing performance. The taxonomy is used to supervise the splitting of these refined tags, and the automatic hierarchical state-split approach is employ to compensate the limitation in the scope and refinement degree of the taxonomy. The experiments are carried out on Penn Chinese Treebank, which show that our method can improve parsing performance significantly.
Prior research suggested the possibility of establishing systematic linkages between some intrinsic features of a presupposition and textual and pragmatic functions that it can carry out with greater probability. This study aims, firstly, to provide an organic view of semantics of presupposition triggers, thanks to a lexical database comprising 19,500 entries. Secondly, the database was used to investigate a corpus of chat conversations including about 200,000 tokens. The results show that triggers occur mainly as non-informative, maintaining an information already known by all participants of the communication; but, depending on their different features, some of them are systematically associated to a function of anaphora and textual cohesion; others to strengthen social conventions and stereotypes. The informative function, although in a minority proportion, is quantitatively significant only in correspondence to a single class of presupposition triggers.
Over the past 15 years, one of the basic functions of the Linguistic Landscape (LL) has been to measure and assess the visibility of multiple languages in public spaces. In research carried out around the world, English has recurrently appeared as one of the most important features of the LL, though its classification, and the extent to which its presence renders a space multilingual, remains controversial. In the light of Seargeant’s (2011) and Yano’s (2011) recent claims that English may be considered a feature of the local landscape rather than a language in its own right, this paper explores code classification in the LL of Toulouse. It challenges the concept of ‘global’ languages (Crystal, 1997), and illustrates the ways in which various codes become affected, consumed, and enveloped by the dominant language, French. Not only is this the case for English, but also the regional tongue, Occitan, whose visibility, despite appearing autonomous, may be more accurately considered a part of the French landscape. This illustrates the shortcomings in classifying languages according to fixed, subjective (and often ad hoc) criteria, because even the presentation of so-called ‘foreign’ languages in Toulouse is influenced by, and subject to, local linguistic norms. The implications of this are profound, as we see that French is the driving force behind every language in the city. Moreover, this calls into question the very definition of multilingualism, as LL authors exhibit a preference for the national code even when writing in others. As such, this research will have a meaningful impact on language coding in the LL, and contribute to the advancement of the methodologies by which we measure and analyse multilingualism, language beliefs, and language practices in public spaces.
The French Lexical Network (fr-LN) is a global model of the French lexicon presently under construction. The fr-LN accounts for lexical knowledge as a lexical network structured by paradigmatic and syntagmatic relations holding between lexical units. This paper describes how morphological knowledge is presently being introduced into the fr-LN through the implementation and lexicographic exploitation of a dynamic morphological model. Section 1 presents theoretical and practical justifications for the approach which we believe allows for a cognitively sound description of morphological data within semantically-oriented lexical databases. Section 2 gives an overview of the structure of the dynamic morphological model, which is constructed through two complementary processes: a Morphological Process--section 3--and a Lexicographic Process--section 4.
In the context of processing Bengali words through a computer, there may arise several issues that are directly linked with surface structure of words. These issues may create problems in manual and computer-based counting of number of words in a corpus. They can also create problems in morphological processing of words. These issues come up because there is hardly any consistency in orthographic representation of words in written Bengali texts. The high irregularities in writing of inflected words, proper names, adjectival forms, adverbial forms, compound words, reduplicated words, onomatopoeic words, hyphenated words, etc. present a daunting task before an investigator in normalizing the surface forms of words for generating a lexical database as well as developing a word processing system for the works of language technology.
In this study, a dictionary-based method is used to extract expressive concepts from documents. So far, there have been many studies concerning concept mining in English, but this area of study for Turkish, an agglutinative language, is still immature. We used dictionary instead of WordNet, a lexical database grouping words into synsets that is widely used for concept extraction. The dictionaries are rarely used in the domain of concept mining, but taking into account that dictionary entries have synonyms, hypernyms, hyponyms and other relationships in their meaning texts, the success rate has been high for determining concepts. This concept extraction method is implemented on documents, that are collected from different corpora.
The purpose of our work is to explore the possibility of using sentence diagrams produced by schoolchildren as training data for automatic syntactic analysis. We have implemented a sentence diagram editor that schoolchildren can use to practice morphology and syntax. We collect their diagrams, combine them into a single diagram for each sentence and transform them into a form suitable for training a particular syntactic parser. In this study, the object language is Czech, where sentence diagrams are part of elementary school curriculum, and the target format is the annotation scheme of the Prague Dependency Treebank. We mainly focus on the evaluation of individual diagrams and on their combination into a merged better version.
We suggest a new annotation scheme for unlexicalized PCFGs that is inspired by formal language theory and only depends on the structure of the parse trees. We evaluate this scheme on the TüBa-D/Z treebank w.r.t. several metrics and show that it improves both parsing accuracy and parsing speed considerably. We also show that our strategy can be fruitfully com-bined with known ones like parent annota-tion to achieve accuracies of over 90 % la-beled F1 and leaf-ancestor score. Despite increasing the size of the grammar, our annotation allows for parsing more than twice as fast as the PCFG baseline. 1
This paper reports on a corpus-based quantitative study of the use of nominalizations across China English and British English in two comparable media corpora. In contrast to previous corpus-based studies of nominalizations, we start by using a syntactic approach and proceed with some methodological innovations incorporating large lexical databases and syntactically annotated corpora. The data show that there are significant differences in the use of nominalizations across these two English varieties. It is hoped that this research will offer useful insights on variations in nominalization across different English varieties and also on the understanding of the two English varieties in question. 1
In current web applications, more businesses are gradually publishing their business as services over the web. This growing number of web services available within an organization and on the Web raises a new and challenging search problem: locating desired web services. Searching for web services with conventional web search engines is insufficient in this context. Automatically clustering Web Service Description Language (WSDL) files on the web into functionally similar homogeneous service groups can be seen as a bootstrapping step for creating a service search engine and, at the same time, reducing the search space for service discovery. In order to overcome some the limitations of pattern-matching approach, the proposed work uses two semantic approaches to cluster similar services. An experimental study based on an information retrieval technique known as latent semantic analysis is applied to the collection of WSDL files and the another semantic approach is based on WordNet which is a lexical database to cluster similar services, as a predecessor step to retrieve the relevant Web services for a user request by search engines. The baseline approach and the two approaches based on semantic is applied on a collection of WSDL documents consisting of to test the quality of clusters formed. As a result, WordNet based approach for clustering shows better cluster quality.
For reinforcing city sports park informatization management level and the society service quality of sports park, this paper researches combined multi-agent technology, and structures multimedia active service system framework of city sports park. It has elaborated functions of every feature and workflow of the system, and put forward Agent design procedure based on JADE. Whats more, it also adopts FIPAACl linguistic norms between Agents communication and puts forward some realized advice which makes the information based on the fundamental of Agent.
The present contribution represents the first step in comparing the nature of syntactico-semantic relations present in the sentence structure to their equivalents in the discourse structure. The study is carried out on the basis of Czech manually annotated material collected in the Prague Dependency Treebank (PDT). According to the analysis of the underlying syntactic structure of a sentence (tectogrammatics) in the PDT, we distinguish various types of relations that can be expressed both within a single sentence (i.e. in a tree) and in a larger text, beyond the sentence boundary (between trees). We suggest that, on the one hand, semantic nature of each type of these relations corresponds both within a sentence and in a larger text (i.e. a causal relation remains a causal relation) but, on the other hand, according to the semantic properties of the relations, their distribution in a sentence or between sentences is very diverse. In this study, this observation is analyzed in detail for three cases (relations of condition, specification and opposition ) and further supported by similar behaviour of the English data from the Penn Discourse Treebank.
This article analyses the structure of Yoruba numerals and their derivation. Data are collected from the compilation of Yoruba numerals and observation of its use coupled with the researcher's intuitive knowledge of the language. The work dwells on the existing literature on numerals too. The author adopts a descriptive method in analysing the data. The work looks at the roles of affixes in realising odd numbers, multiples of 20, centenary, bicentenary, and so on in their order of increase. It is discovered that the direction of counting in Yoruba is largely progressive. Besides, the language adopts base 5, decimal (base 10) and vigesimal (base 20) systems of counting. It is equally discovered that the choice of either of the two variations is largely dependent on the articulatory parameter of the first vowel (V1) of the root word. It is noted that the Yoruba numeral system offers a suitable linguistic database for both the theoretical and empirical domains of linguistic study especially documentary linguistics. The current study has general pedagogic implications for the teaching and learning of Yoruba numerals.
This paper presents results of dependency parsing of Old French, a language which is poorly standardized at the lexical level, and which displays a relatively free word order. The work is carried out on five distinct sample texts extracted from the dependency treebank Syntactic Reference Corpus of Medieval French (SRCMF). Following Achim Stein's previous work, we have trained the Mate parser on each sub-corpus and cross-validated the results. We show that the parsing efficiency is diminished by the greater lexical variation of Old French compared to parse results on modern French. In order to improve the result of the POS tagging step in the parsing process, we applied a pre-treatment to the data, comparing two distinct strategies: one using a slightly post-treated version of the TreeTagger trained on Old French by Stein, and a CRF trained on the texts, enriched with external resources. The CRF version outperforms every other approach.
Abstract: Similarity is criteria of measuring nearness or proximity between two concepts. Several algorithmic approaches for computing similarity have been proposed. Among the existing Similarity measure, majority of them utilize WordNet as an underlying ontology for calculating semantic similarity. WordNet is a lexical database for English Language which was created and maintained by Congnitive Science Laboratory at Princeton University under the supervision of Professor George A. Miller. It is organized as a network which consists of concepts or terms called Synsets (list of synonyms terms) and the relationship between them. There are different type of relationship exists in WordNet such as is-a, part-of, synonym and antonym. It has thdatabases, one for noun, one for verb and one for adverb and adjective. This project work proposes a metric for semantic relatedness calculation between pair of concepts which uses Tversky’s feature based approach which takes into account the common and distinct feature of the two terms or concepts. If commonality is more as compared to differences the similarity between concepts is high otherwise similarity is low. Tversky’s theory is quantified by information content of two concepts and the Information content of most specific common ancestor of two concepts. As we move down in the WordNet hierarchy, more specific and more Informative concept are there, where as when we move up in the hierarchy more Generalized and less Informative concepts are there. So depth of a concept in the WordNet hierarchy is a critical factor in similarity calculation. We take into consideration the depth of the specific concept in the WordNet hierarchy which is the deciding factor for determining the relevance of distinct feature specific to a concept in similarity calculation. Introduction of depth reduces the impact of the less relevant dissimilarity indulge in similarity calculation thereby increase precision. We carried out our experiment of 28
Constituent Context Model (CCM) is an effective generative model for grammar induction, the aim of which is to induce hierarchical syntactic structure from natural text. The CCM simply defines the Multinomial distribution over constituents, which leads to a severe data sparse problem because long constituents are unlikely to appear in unseen data sets. This paper proposes a Bayesian method for constituent smoothing by defining two kinds of prior distributions over constituents: the Dirichlet prior and the Pitman-Yor Process prior. The Dirichlet prior functions as an additive smoothing method, and the PYP prior functions as a back-off smoothing method. Furthermore, a modified CCM is proposed to differentiate left constituents and right constituents in binary branching trees. Experiments show that both the proposed Bayesian smoothing method and the modified CCM are effective, and combining them attains or significantly improves the state-of-the-art performance of grammar induction evaluated on standard treebanks of various languages.
Abstract Discourse parsing has become an inevitable task to process information in the natural language processing arena. Parsing complex discourse structures beyond the sentence level is a significant challenge. This article proposes a discourse parser that constructs rhetorical structure (RS) trees to identify such complex discourse structures. Unlike previous parsers that construct RS trees using lexical features, syntactic features and cue phrases, the proposed discourse parser constructs RS trees using high‐level semantic features inherited from the Universal Networking Language (UNL). The UNL also adds a language‐independent quality to the parser, because the UNL represents texts in a language‐independent manner. The parser uses a naive Bayes probabilistic classifier to label discourse relations. It has been tested using 500 Tamil‐language documents and the Rhetorical Structure Theory Discourse Treebank, which comprises 21 English‐language documents. The performance of the naive Bayes classifier has been compared with that of the support vector machine (SVM) classifier, which has been used in the earlier approaches to build a discourse parser. It is seen that the naive Bayes probabilistic classifier is better suited for discourse relation labeling when compared with the SVM classifier, in terms of training time, testing time, and accuracy.
The Stuttgart-Tbingen Tagset (STTS) is a widely used POS annotation scheme for German which provides 54 different tags for the analysis on the part of speech level.The tagset, however, does not distinguish between adverbs and different types of particles used for expressing modality, intensity, graduation, or to mark the focus of the sentence.In the paper, we present an extension to the STTS which provides tags for a more fine-grained analysis of modification, based on a syntactic perspective on parts of speech.We argue that the new classification not only enables us to do corpus-based linguistic studies on modification, but also improves statistical parsing.We give proof of concept by training a data-driven dependency parser on data from the TiGer treebank, providing the parser a) with the original STTS tags and b) with the new tags.Results show an improved labelled accuracy for the new, syntactically motivated classification.
Recurrent neural network language models have solved the problems of data sparseness and dimensionality disaster which exist in traditional N-gram models. RNNLMs have recently demonstrated state-of-the-art performance in speech recognition, machine translation and other tasks. In this paper, we improve the model performance by providing contextual word vectors in association with RNNLMs. This method can reinforce the ability of learning long-distance information using vectors training from Skip-gram model. The experimental results show that the proposed method can improve the perplexity performance significantly on Penn Treebank data. And we further apply the models to speech recognition task on the Wall Street Journal corpora, where we achieve obvious improvements in word-error-rate.
This paper presents a set of Bilingual Dictionary Drafting (BDD) methods including manual extraction from existing lexical databases and corpus based NLP tools, as well as their evaluation on the example of German-Basque as language pair. Our aim is twofold: to give support to a German-Basque bilingual dictionary project by providing draft Bilingual Glossaries and to provide lexicographers with insight into how useful BDD methods are. Results show that the analysed methods can greatly assist on bilingual dictionary writing, in the context of medium-density language pairs.
The lexicon dynamics deals with the changes that occur in language from a historical stage to another. As part of a living organism, words emerge and fade away, bearing the mark of the linguistic norms. The linguists’ preoccupation with the accuracy of language goes back in time and is related to lexicon, semantics, grammar, spelling etc. The present study is designed as a concise presentation of certain mistakes frequently met in the current written press, with a view to correcting them. Some deviations from the norms are minor, others major, whereas their causes are numerous. Recent loans, neologisms, confusion of styles and excessive use of clichés are only a few aspects to be approached in the present work.
In recent years, the emergence of English as an International Language (EIL) has paved the way for its global speakers to use it as a means of interacting globally, and representing themselves and their cultures internationally. Although English is globally considered as an international language and as a tool to be used in cross-cultural communication with people having various first languages from different parts of the world, native-speakers’ norms and cultures still dominate the language materials that are developed to be globally used. In fact, English language coursebooks insists on bombarding the ELT world with culturally-loaded native-speaker themes, such as actors in Hollywood (Coskun, 2009). Prodromou (1988) similarly underlines the issue that the majority of English language coursebooks are published by major Anglo-American publishers in Inner Circle countries and these coursebooks include cultural situations that most students will never come across, such as ‘finding a flat in London’ (p. 80). Considering the importance given to the growing role of EIL, the issue of linguistic norms and cultural content in language learning materials has remained one of the unresolved problems in the process of materials development. A group of scholars argues in favor of localizing the materials by using the learners’ experiences and making English language coursebooks culturally responsive to their needs. The opponents solely favor the integration of the linguistic and cultural norms of the native speakers of English in language learning materials. As far as EIL is concerned, there are several aspects that need to be taken into close account when language teaching materials are being prepared to be globally used. In a nutshell, in EIL era, while preparing English language coursebooks, rather than just integrating English of Specific Cultures, the linguistic and cultural norms of the native speakers of English, as the sole reference in the contents of the English language coursebook, at least a due attention should be paid to English for Specific Cultures, the linguistifc and cultural norms of non-native speakers of English. This study recommends a group of essential features for the future English language coursebooks in EIL era.
Proof theory began in the 1920’s as a part of Hilbert’s program. That program aimed to secure the foundations of mathematics by modeling infinitary mathematics with formal axiomatic systems, and proving those systems consistent using restricted, “finitary” means. The program thus viewed mathematics as a system of reasoning with precise linguistic norms, governed by rules that can be described and studied in concrete terms. Such a viewpoint, today, has applications in mathematics, computer science, and the philosophy of mathematics.
Dependency parsers, which are widely used in natural language processing tasks, employ a representation of syntax in which the structure of sentences is expressed in the form of directed links (dependencies) between their words. In this article, we introduce a new approach to transition‐based dependency parsing in which the parsing algorithm does not directly construct dependencies, but rather undirected links, which are then assigned a direction in a postprocessing step. We show that this alleviates error propagation, because undirected parsers do not need to observe the single‐head constraint, resulting in better accuracy. Undirected parsers can be obtained by transforming existing directed transition‐based parsers as long as they satisfy certain conditions. We apply this approach to obtain undirected variants of three different parsers (the Planar, 2‐Planar, and Covington algorithms) and perform experiments on several data sets from the CoNLL‐X shared tasks and on the Wall Street Journal portion of the Penn Treebank, showing that our approach is successful in reducing error propagation and produces improvements in parsing accuracy in most of the cases and achieving results competitive with state‐of‐the‐art transition‐based parsers.
The article is devoted to the problem of translation from French into Polish from the perspective of the object oriented approach proposed by Wiesław Banyś. The author takes into consideration some of the problems, both of theoretical and practical nature, which appear while working on the formation of contrastive lexical database for automatic translation of texts. While analyzing specific chosen examples which may cause different kinds of problems in the description, the author offers their interpretations in the target language in accordance with the adopted approach.
Recent work has sparked new interest in type-supervised part-of-speech tagging, a data setting in which no labeled sen-tences are available, but the set of allowed tags is known for each word type. This paper describes observational initializa-tion, a novel technique for initializing EM when training a type-supervised HMM tagger. Our initializer allocates probabil-ity mass to unambiguous transitions in an unlabeled corpus, generating token-level observations from type-level supervision. Experimentally, observational initializa-tion gives state-of-the-art type-supervised tagging accuracy, providing an error re-duction of 56 % over uniform initialization on the Penn English Treebank. 1
We investigate whether parsers can be used for self-monitoring in surface realization in order to avoid egregious errors involving "vicious" ambiguities, namely those where the intended interpretation fails to be considerably more likely than alternative ones. Using parse accuracy in a simple reranking strategy for selfmonitoring, we find that with a stateof-the-art averaged perceptron realization ranking model, BLEU scores cannot be improved with any of the well-known Treebank parsers we tested, since these parsers too often make errors that human readers would be unlikely to make. However, by using an SVM ranker to combine the realizer's model score together with features from multiple parsers, including ones designed to make the ranker more robust to parsing mistakes, we show that significant increases in BLEU scores can be achieved. Moreover, via a targeted manual analysis, we demonstrate that the SVM reranker frequently manages to avoid vicious ambiguities, while its ranking errors tend to affect fluency much more often than adequacy.
Automatically acquiring semantic verb classes from corpora is a challenging task, especially with no existing treebank. Building a high-performing parser for a language is still crucially depends on the existence of large, in-domain texts as training data. While previous work has focused primarily on major languages, how to extend these results to other languages is the way to avoid working start from scratch. In general, a large monolingual corpus in a resource-rich source language labeled with lexico-syntactic information, and a very limited bilingual corpus are available. This paper addresses the problem of verb classification automatically in Tibetan using bilingual lexicon and translation information.
Search engines have become the main way for people to get expected information, most of them are based on keyword search. However, keyword search is based on computing the similarity of letters of the keywords, instead of semantic meaning, therefore the searching results often include irrelevant information to user intention. This paper aims to find a way on improving keyword search efficiency. Using Wikipedia, which is the largest online encyclopedia, this paper explores the relations of terms through computing the semantic relatedness between words, and presents an algorithm called WLA in the light of link structure and text message in Wikipedia. What is more, we design a terms query platform through which users will be able to get all the meanings about the concepts. By making a comparison with lexical database WordNet, it has demonstrated the feasibility on our methods.
The well-established memory bias for arousing-negative stimuli seems to be enhanced in high trait-anxious persons and persons suffering from anxiety disorders. We monitored the emergence and development of such a bias during and after learning, in high and low trait anxious participants. A word-learning paradigm was applied, consisting of spoken pseudowords paired either with arousing-negative or neutral pictures. Learning performance during training evidenced a short-lived advantage for arousing-negative associated words, which was not present at the end of training. Cued recall and valence ratings revealed a memory bias for pseudowords that had been paired with arousing-negative pictures, immediately after learning and two weeks later. This held even for items that were not explicitly remembered. High anxious individuals evidenced a stronger memory bias in the cued-recall test, and their ratings were also more negative overall compared to low anxious persons. Both effects were evident, even when explicit recall was controlled for. Regarding the memory bias in anxiety prone persons, explicit memory seems to play a more crucial role than implicit memory. The study stresses the need for several time points of bias measurement during the course of learning and retrieval, as well as the employment of different measures for learning success.
We propose a novel approach for learning image representation based on qualitative assessments of visual aesthetics. It relies on a multi-node multi-state model that represents image attributes and their relations. The model is learnt from pair wise image preferences provided by annotators. To demonstrate the effectiveness we apply our approach to fashion image rating, i.e., comparative assessment of aesthetic qualities. Bag-of-features object recognition is used for the classification of visual attributes such as clothing and body shape in an image. The attributes and their relations are then assigned learnt potentials which are used to rate the images. Evaluation of the representation model has demonstrated a high performance rate in ranking fashion images.
Resumen Este artículo analiza las actitudes lingüísticas de hablantes nativos de español de la Ciudad Autónoma de Buenos Aires, hacia al español de la Argentina y el español de los otros países hispanohablantes. El artículo es parte de los resultados del Proyecto LIAS (Linguistic Identity and Attitudes in Spanish-speaking Latin America), financiado por El Consejo Noruego de Investigación (RCN). La recolección de los datos se realizó en la capital del país, entrevistando a una muestra de 400 informantes previamente estratificada con las variables de edad, sexo y nivel socioeconómico. El procesamiento estadístico de los datos de campo recolectados arrojó resultados de interés en torno a la mayoría de los tópicos analizados y especialmente en lo referente a aspectos tales como la valoración positiva de la propia variedad lingüística; la resistencia a identificar a España como la única fuente de la norma lingüística de la lengua española; el rechazo a la unificación de la lengua y, por consiguiente, la defensa de la diversidad lingüística como portadora de riqueza cultural. Abstract This article analyzes the linguistic attitudes of native Spanish speakers from Buenos Aires City, towards Spanish spoken in Argentina and in the other Spanish-speaking countries. It is a result of the LIAS-Project (Linguistic Identity and Attitudes in Spanish-speaking Latin America), funded by The Research Council of Norway (RCN). The data were gathered in the capital of the country, interviewing a stratified sample of 400 respondents, based on the variables of age, sex and socioeconomic status. The analysis of the data rendered interesting results on most of the analyzed topics; especially important was the positive appraisal of Argentineans' own linguistic variety; the strong resistance against identifying Spain as the only source of the linguistic norm for the Spanish language; and the rejection of language unification, defending in this way linguistic diversity as an important conveyor of cultural richness.
This paper describes experiments for statistical dependency parsing using two different parsers trained on a recently extended dependency treebank for Greek, a language with a moderately rich morphology. We show how scores obtained by the two parsers are influenced by morphology and dependency types as well as sentence and arc length. The best LAS obtained in these experiments was 80.16 on a test set with manually validated POS tags and lemmas. 1
BACKGROUND: Careful observation of the longitudinal course of bipolar disorders is pivotal to finding optimal treatments and improving outcome. A useful tool is the daily prospective Life-Chart Method, developed by the National Institute of Mental Health. However, it remains unclear whether the patient version is as valid as the clinician version. METHODS: We compared the patient-rated version of the Lifechart (LC-self) with the Young-Mania-Rating Scale (YMRS), Inventory of Depressive Symptoms-Clinician version (IDS-C), and Clinical Global Impression-Bipolar version (CGI-BP) in 108 bipolar I and II patients who participated in the Naturalistic Follow-up Study (NFS) of the German centres of the Bipolar Collaborative Network (BCN; formerly Stanley Foundation Bipolar Network). For statistical evaluation, levels of severity of mood states on the Lifechart were transformed numerically and comparison with affective scales was performed using chi-square and t tests. For testing correlations Pearson´s coefficient was calculated. RESULTS: Ratings for depression of LC-self and total scores of IDS-C were found to be highly correlated (Pearson coefficient r = -.718; p <.001), whilst the correlation of ratings for mania with YMRS compared to LC-self were slightly less robust (Pearson coefficient r =.491; p =.001). These results were confirmed by good correlations between the CGI-BP IA (mania), IB (depression) and IC (overall mood state) and the LC-self ratings (Pearson coefficient r =.488, r =.721 and r =.65, respectively; all p <.001). CONCLUSIONS: The LC-self shows a significant correlation and good concordance with standard cross sectional affective rating scales, suggesting that the LC-self is a valid and time and money saving alternative to the clinician-rated version which should be incorporated in future clinical research in bipolar disorder. Generalizability of the results is limited by the selection of highly motivated patients in specialized bipolar centres and by the open design of the study.
The sentiment mining approaches can typically be divided into lexicon and machine learning approaches. Recently there are an increasing number of approaches which combine both to improve the performance when used separately. However, this still lacks contextual understanding which led to the introduction of deep learning approaches which allows for semantic compositionality over a sentiment treebank. This paper enhances the deep learning approach with semantic lexicon so that scores can be computed in-stead merely nominal classification. Besides, neutral classification is also improved. Results suggest that the approach outperforms its original.
Automatic text categorisation systems is a type of software that every day it is receiving more interest, due not only to its use in documentaries environments but also to its possible application to tag properly documents on the Web. Many options have been proposed to face this subject using statistical approaches, natural language processing tools, ontologies and lexical databases. Nevertheless, there have been no too many empirical evaluations comparing the influence of the different tools used to solve these problems, particularly in a multilingual environment. In this paper we propose a multi-language rule-based pipeline system for automatic document categorisation and we compare empirically the results of applying techniques that rely on statistics and supervised learning with the results of applying the same techniques but with the support of smarter tools based on language semantics and ontologies, using for this purpose several corpora of documents. GENIE is being applied to real environments, which shows the potential of the proposal.
The web today is huge and enormous collection of data today and it goes on increasing day by day. Thus, searching for some particular data in this collection has a significant impact. Researches taking place give prominence to the relevancy and relatedness of the data that is found. Inspite of their relevance pages for any search topic, the results are still huge to be explored. Another important issue to be kept in mind is the users’ standpoint differs from time to time from topic to topic. Effective relevance prediction can help avoid downloading and visiting many irrelevant pages. The performance of a crawler depends mostly on the opulence of links in the specific topic being searched. This paper reviews the researches on web crawling algorithms used for searching. Keywords— Web Crawling Algorithms, Crawling Algorithm Survey, Search Algorithms, Lexical Database, Metadata, Semantic. __________________________________________________*****_________________________________________________
Abstract We investigated discrimination in the context of evaluating advertisements, based on the suppression model (justification‐suppression model [ JSM ]) of prejudice expression. Previous research has demonstrated that when people are given an opportunity to give a high rating to an ad featuring a Black model, a sense of nonprejudice is created, which, in turn, provides an opportunity to discriminate subsequently without feeling prejudiced. We extended the JSM by investigating whether the acquisition of legitimacy credits (a moral authority earned by demonstrating nonprejudice) is a sufficient condition to release the expression of prejudice. We found that subjects who first evaluated a high‐quality ad featuring a Black model felt eligible to use legitimacy credits in subsequent evaluations. But in a subsequent study, participants who acquired these credits evaluated Black model ads more negatively than White model ads only when these ads were of low quality. The implications for evaluating the subtle way that prejudice affects rating of models of color in advertisements are discussed.
With growing interest in the creation and search of linguistic annotations that form general graphs (in contrast to formally simpler, rooted trees), there also is an increased need for infrastructures that support the exploration of such representations, for example logical-form meaning representations or semantic dependency graphs. In this work, we heavily lean on semantic technologies and in particular the data model of the Resource Description Framework (RDF) to represent, store, and efficiently query very large collections of text annotated with graph-structured representations of sentence meaning. Keywords:Semantic Dependency Graphs, Treebank Search, Resource Description Framework 1.
This paper introduces a new technique for phrase-structure parser analysis, catego-rizing possible treebank structures by inte-grating regular expressions into derivation trees. We analyze the performance of the Berkeley parser on OntoNotes WSJ and the English Web Treebank. This provides some insight into the evalb scores, and the problem of domain adaptation with the web data. We also analyze a “test-on-train ” dataset, showing a wide variance in how the parser is generalizing from differ-ent structures in the training material. 1
Natural language is a fundamental thing of human-society to communicate and interact with one another. In this globalization era, we interact with different regional people as per our interest in social, cultural, economical, educational and professional domain. There are thousands of natural languages exist in our earth. It is quite tough, rather impossible to know all the languages. So we need a computerized approach to convert one natural language to another as per our necessity. This computerized conversion among multiple languages is known as multilingual machine translation. But in this paper we work with a bilingual model, where we concern with two languages: English and Bengali. We use soft computational approach where fuzzy If-Then rule is applied to choose a lemma from prior knowledge; Penn TreeBank PoS tags and HMM tagger are used as lexical class marker to each word in corpora.
This paper investigates the recognition of unknown words in Chinese parsing. Two methods are proposed to handle this problem. One is the modification of a character-based model. We model the emission probability of an unknown word using the first and last characters in the word. It aims to reduce the POS tag ambiguities of unknown words to improve the parsing performance. In addition, a novel method, using graph-based semisupervised learning (SSL), is proposed to improve the syntax parsing of unknown words. Its goal is to discover additional lexical knowledge from a large amount of unlabeled data to help the syntax parsing. The method is mainly to propagate lexical emission probabilities to unknown words by building the similarity graphs over the words of labeled and unlabeled data. The derived distributions are incorporated into the parsing process. The proposed methods are effective in dealing with the unknown words to improve the parsing. Empirical results for Penn Chinese Treebank and TCT Treebank revealed its effectiveness.