Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This paper presents the work that we have carried out in inves tigating the purpose of discourse structure forwhy-question answering (why-QA). We developed a system for answer- ing why-questions that employs the discourse relations in a pre-an notated document collection (the RST Treebank). With this method, we obtain a recall of 53.3% with a mean reciprocal rank (MRR) of 0.662. We argue that the maximum recall that can be ob tained from the use of RST relations as proposed in the present paper is 58.0%. If we dis card the questions that require world knowledge, maximum recall is 73.9%. We conclude that d iscourse structure can play an important role in complex question answering, but that more forms of linguistic processing are needed for increasing recall.
Recent work identifies two properties that appear particularly relevant to the characterization of graph-based dependency models of syntactic structure: the absence of interleaving substructures (well-nestedness) and a bound on a type of discontinuity (gap-degree ≤ 1) successfully describe more than 99% of the structures in two dependency treebanks (Kuhlmann and Nivre 2006). Bodirsky et al. (2005) establish that every dependency structure with these two properties can be recast as a lexicalized Tree Adjoining Grammar (LTAG) derivation and vice versa. However, multi-component extensions of TAG (MC-TAG), argued to be necessary on linguistic grounds, induce dependency structures that do not conform to these two properties (Kuhlmann and Möhl 2006). In this paper, we observe that several types of MC-TAG as used for linguistic analysis are more restrictive than the formal system is in principle. In particular, tree-local MC-TAG, tree-local MC-TAG with flexible composition (Kallmeyer and Joshi 2003), and special cases of set-local TAG as used to describe certain linguistic phenomena satisfy the well-nested and gap degree ≤ 1 criteria. We also observe that gap degree can distinguish between prohibited and allowed wh-extractions in English, and report some preliminary work comparing the predictions of the graph approach and the MC-TAG approach to scrambling.
Emotions are currently conceptualized as ongoing temporal processes. Consistent with this view, an important target of attempts at emotion regulation are the temporal characteristics of an emotional response. The process model of emotion regulation (Gross, 1998a) distinguishes between antecedent- and response-focused emotion regulation strategies, depending on when during the unfolding emotional response they act. Two strategies that exemplify this distinction are cognitive reappraisal and expressive suppression. The present study explored the effects of the interaction between habitual engagement in reappraisal and suppression and their voluntary manipulation. Using a between-subjects design, 122 participants selected based on their self-reported habitual emotion regulation strategy (reappraisal, suppression, or both strategies without clear preference for one over the other) received instructions to engage in reappraisal, suppression, or merely watch emotion-eliciting images. Chronometric analyses of emotion-related psychophysiological measures (startle reflex modulation, corrugator electromyography, and skin conductance) were conducted in order to further characterize the differences in the time course of these two strategies during the down-regulation of negative emotion. As expected, instructions to reappraise resulted in lower unpleasantness and arousal ratings, as well as less overall corrugator electromyographic activity, compared to instructions to suppress. No differences between instruction conditions were observed on startle reflex or skin conductance. Moreover, no differences were observed in the chronometry of any of the physiological measures. Habitual emotion regulation style had no direct effect on any of the dependent variables, and it did not interact with instruction condition. The implications for the study of the chronometry of emotion regulation are discussed.
The need for the representation of both semantics and common sense and its organization in a lexical database or knowledge base has motivated the development of large projects, such as Wordnets, CYC and Mikrokosmos. Besides the generic bases, another approach is the construction of ontologies for specific domains. Among the advantages of such approach there is the possibility of a greater and more detailed coverage of a specific domain and its terminology. Domain ontologies are important resources in several tasks related to the language processing, especially in those related to information retrieval and extraction in textual bases. Information retrieval or even question and answer systems can benefit from the domain knowledge represented in an ontology. Besides embracing the terminology of the field, the ontology makes the relationships among the terms explicit.
This paper describes a method for automatic acquisition of wide-coverage treebank-based deep linguistic resources for Japanese, as part of a project on treebankbased induction of multilingual resources in the framework of Lexical-Functional Grammar (LFG). We automatically annotate LFG f-structure functional equations (i.e. labelled dependencies) to the Kyoto Text Corpus version 4.0 (KTC4) (Kurohashi and Nagao 1997) and the output of of Kurohashi-Nagao Parser (KNP) (Kurohashi and Nagao 1998), a dependency parser for Japanese. The original KTC4 and KNP provide unlabelled dependencies. Our method also includes zero pronoun identification. The performance of the f-structure annotation algorithm with zero-pronoun identification for KTC4 is evaluated against a manually-corrected Gold Standard of 500 sentences randomly chosen from KTC4 and results in a pred-only dependency f-score of 94.72%. The parsing experiments on KNP output yield a pred-only dependency f-score of 82.08%.
In this paper we will describe VIT (Venice Italian Treebank), created at the University of Venice. We will focus on the syntactic-semantic features and on the quantitative analysis of the data of our treebank comparing them to other treebanks. In general, we will try to substantiate the claim that treebanking grammars or parsers is dramatically dependent on the chosen treebank; and eventually this process seems to be dependent either from substantial factors such as the adopted linguistic framework for structural description or, ultimately, the described language.
This paper proposes a novel Chinese syntactic parsing model based on semantic class, which is a variant of normal lexicalized statistical model. It attempts to make use of the syntactic and semantic similarity between Chinese words and then produces a more knowledgeable estimate of the probability of grammar rules. A simple but effective unsupervised method is designed to determine the proper semantic class of given words. Semantic class is used to improve the performance of parsing model. We evaluate our methods on the widely used Penn Chinese Treebank. Experimental results show that it outperforms a famous lexicalized model significantly on appropriate semantic class levels.
Query expansion(QE) has been proved to be one of effective methods for improving the performance of the information retrieval(IR) system.Therefore,a new fuzzy QE method based on synonymy thesaurus is proposed,and the synonymy thesaurus is built based on the famous lexical database WordNet.In the synonymy thesaurus,the similarity between the synonyms is,which is obtained by Tanimoto coefficient.By using this synonymy thesaurus,query expansion can be done well.Then the fuzzy QE method is introduced into the document information retrieval system together with the modified vector space model.The experimental results show that the developed information retrieval system has got more effective performance than before by using the fuzzy query expansion method.One feature of the proposed information retrieval model is that it can be treated as one of simple semantic models.Another feature is that the expansion degree is controllable based on different thresholds.
Reviewed by: Antología conmemorativa: Nueva Revista de Filología Hispánica. Cincuenta tomos ed. by Alejandro Rivas and Yliana Rodríguez Natalya I. Stolova Antología conmemorativa: Nueva Revista de Filología Hispánica. Cincuenta tomos. Vol. 2. Ed. by Alejandro Rivas and Yliana Rodríguez. (Publicaciones de la Nueva Revista de Filología Hispánica 9.) Mexico City: El Colegio de México, 2003. Pp. viii, 651. ISBN 9681211154. The present work is the second of the two anniversary collections published by Nueva Revista de Filología Hispánica (NRFH) to celebrate the appearance of its fiftieth volume. Volume 2 contains thirty-two selected articles that have been published in the journal over the years. These include eighteen papers on literary topics and fourteen papers concerned with linguistics. I limit my description to the linguistic articles, giving in parentheses the original publication date. A series of papers in the collection adopt a diachronic or philological perspective. Eugenio Coseriu (1961) challenges the Arabic origin of several Spanish and Rumanian expressions, arguing that these originated within the Romance language family. Rafael Lapesa (1961) traces the development of the Latin demonstratives into the Spanish and French articles. Margherita Morreale (1963–64) offers a philological commentary on the Evangelio de San Mateo según el manuscrito escurialense I-j-6: Texto, gramática y vocabulario published by Thomas Montgomery in 1962. Germán de Granda (1978) examines the history behind the verbal diphthongized voseo forms. Yakov Malkiel (1988) describes the demise of Old Spanish nozir, nuzir ‘harm’ during the Late Middle Ages. The different varieties of Spanish constitute the focus of a cluster of five papers. María Josefa Canellada de Zamora and Alonso Zamora Vicente (1960), as well as Juan M. Lope Blanch (1963–64), treat the reduction and the loss of unstressed vowels in Mexican Spanish. Tracy D. Terrell (1978) focuses on the aspiration and the elision of the implosive and final /s/ in the Spanish of Puerto Rico. Manuel Alvar (1988) challenges the notion of el dialecto andaluz ‘the Andalusian dialect’. Guillermo L. Guitarte (1992) and María Beatriz Fontanella de Weinberg (1995) take up the phenomenon of rehilamiento in the nineteenth-century Spanish of Buenos Aires. Two articles employ Spanish data to consider issues related to linguistic terminology and linguistic theory. Bernard Pottier (1961) explores the notion of auxiliary verb. José Pedro Rona (1973) addresses the question of linguistic norm in the context of the different local, regional, national, and pan-Latin American features. Finally, Antonio Quilis (1982) provides a description of the grammar of Tagalog Arte y reglas de la lengua tagala (1610) written by missionary Francisco de San José Blancas and places this work within the context of the missionary linguistics of the Philippines. This volume is a valuable source for Hispanists, and for linguists interested in key works on the Spanish language published from the 1960s to the 1990s. Natalya I. Stolova Colgate University Copyright © 2007 Linguistic Society of America
Since noun phrases are the most popular phrases in texts, noun phrase identification is one of vital subtasks of natural language processing. Generally Chinese noun phrases have hierarchical inner structures. This paper proposes an approach of defining various levels of granularity for noun phrases, catering for different application demands. Three levels of granularity noun phrases are proposed, that is, concept noun phrase, base noun phrase and entire noun phrase. The task of noun phrase identification is to label word sequences with phrase tags. All granularity noun phrase identifications are cast as classification problem under certain encoding schemes. The experimental dataset is acquired empirically from Chinese Penn Treebank 5.1. F, measure of concept noun phrase, base noun phrase and entire noun phrase identification reaches 92.12%, 84.13% and 85.32% respectively.
In this paper, we address the issue of improving a Chinese chunking system with rich lexicalized information. A method that incorporates statistical information based on distributional similarity between words obtained from large unlabeled corpus and morphological knowledge into a state-of-the-art CRF-based chunking model is proposed to tackle the data sparseness problem given limited amount of labeled training data. Evaluations are performed on the latest release of Chinese Treebank, and experimental results show that our method outperforms the chunking models based on features over word and automatically assigned POS tags when using the same amount of training data.
The paper aims at the complexity of syntactic network and the feasibility that the complex network work as a means of linguistic studies.The paper proposes the method how to build a syntactic based on dependency treebank and investigates the complexity of Chinese syntactic dependency network based on two Chinese treebanks with different genres.The results show that syntactic networks have similar average path length and diameter with the random networks,but cluster coefficients of syntactic networks are much greater than that of random networks,and degree distributions of syntactic networks also obey the power law.The paper reveals that two syntactic networks have the same diameter,but with different average degree,path length,cluster coefficients and power exponent.
Abstract. Despite their widespread use in Natural Language Processing applications, lexical databases and wordnets in particular do not yet contribute satisfactorily to the difficult problem of automatic word sense discrimination. Having built a number of lexical databases ourselves, we are keenly aware of still unresolved fundamental theoretical issues. In this paper we examine some of these questions and suggests preliminary answers concerning the nature of lexical elements and the conceptualsemantic and lexical relations that interconnect them. Our perspective is multilingual, and our goal is to formulate a proposal for a “Global Wordnet Grid ” that will meet the challenge of mapping the lexicons of many languages in interesting and useful ways. 1
The purpose of the Chinese PropBank (CPB) project is to add a layer of annotation to the hand-parsed sentences in the Chinese Treebank (CTB) (Xue et al., 2005). This layer of annotation assigns predicatespecific argument labels to the constituents in a parse tree. The arguments of each predicate in the sentence, which are limited to verbs and their nominalizations in the work we report here, receive an argument label in the form of ArgN, where N is an integer between 0 and 5. These numbered arguments represent core arguments that are defined in relation to the predicate, which is labeled as Rel. Each core argument plays a unique role with regard to the predicate and generally the total number of core arguments for each predicate does not exceed 6. The core arguments annotated for the verb N (”investigate”) in Example (1) are the NPs (”the police”) and (”accident”) I(”cause”), which are labeled as Arg0 and Arg1 respectively. The semantic role labels added to the parse tree are in bold.
Semantic analysis has become a bottleneck of many natural language applications. Machine translation, automatic question answering, dialog management, and others rely on high quality semantic analysis. Verbs are central elements of clauses with strong influence on the realization of whole sentences. Therefore the semantic analysis of verbs plays a key role in the analysis of natural language. We believe that solid disambiguation of verb senses can boost the performance of many real-life applications. In this thesis, we investigate the potential of statistical disambiguation of verb senses. Each verb occurrence can be described by diverse types of information. We investigate which information is worth considering when determining the sense of verbs. Different types of classification methods are tested with regard to the topic. In particular, we compared the Naive Bayes classifier, decision trees, rule-based method, maximum entropy, and support vector machines. The proposed methods are thoroughly evaluated on two different Czech corpora, VALEVAL and the Prague Dependency Treebank. Significant improvement over the baseline is observed.
The is the reported phenomenon of increased spatial abilities after listening to that composer's music. However, subsequent research suggests that the Mozart effect may be an artifactual consequence of heightened arousal and mood rather than the music of Mozart per se (e.g., Thompson, Schellenberg, & Husain, 2001). The present study considers if performance improvements in a scored computer game are consistent with the mood and arousal hypothesis. Indeed, the use of a computer game as the experimental vehicle makes this work notably the most ecologically valid study of the Mozart effect to date. Specifically, in this work, ratings of musical preference as well as the game performance of individuals listening to different types of music are compared. If arousal and mood are the real we hypothesized that the performance level of participants would increase when listening to the selections they most enjoy. Results supported this hypothesis. ********** The Mozart Effect (e.g., Rauscher, Shaw, & Ky, 1993) is the reported phenomenon that listening to Mozart would temporarily increase spatial reasoning ability by the equivalent of 8-9 points on the Stanford-Binet. Rauscher, Shaw, and Ky explain the Mozart Effect by suggesting that exposure to musical compositions that are structurally complex excites certain cortical firing patterns comparable to those activated when completing spatial-temporal tasks. The Mozart Effect has also been seized upon by the media, and even distorted into the claim that simply listening to Mozart would make people smarter. Indeed, the idea that passively listening to Mozart might increase IQ scores has sparked the development of many educational books and music products (see McKelvie & Low, 2002). As Nantals and Schellenberg (1999) note, one Governor even budgeted for a compact disc or cassette for each infant born in his state. Despite issues with face validity, the Mozart Effect has been seriously discussed in such prestigious publications as Science and Nature, and still frequents the pages of respected psychology journals. At times, there have been problems replicating the basic but it has been suggested by Rauscher, Shaw, and Ky (1998) that inconsistent results by other researchers can be attributed to methodological differences. However, Nantais and Schellenberg (1999) had no difficulty replicating the basic finding: That is, they found a significant increase in performance on spatial-temporal tasks for subjects that heard a musical piece; but, there was no marked difference between those that heard Mozart or those who heard Schubert. Likewise, other researchers (e.g., Ashby, Isben, & Turken, 1999; Steele, Bass, & Crook, 1999) also observed that changes in mood can have a significant effect on cognitive performance, and that the original experimental conditions (e.g., listening to Mozart, relaxation music, or silence) likely each have an affect on mood and arousal. As such, the argument emerged that observed performance differences may occur due to improvements in mood and arousal rather than from neurophysiological priming. Consistent, Thompson, Schellenberg, and Husain (2001) reported that individuals that listened to Mozart performed better on spatial tasks, but also scored higher on positive mood and arousal ratings. Subjects that scored low on mood and arousal showed no effect of the music. By examining participant's spatial abilities after listening to a Mozart sonata (expected to produce positive mood), and an adagio by Albinoni (a sad piece), they were able to provide additional support for the arousal and mood hypothesis. In summary, the most current explanation for the Mozart Effect would suggest that an individual's mood/preference for a particular piece of music should correlate with any cognitive gains (Steele, 2000). In fact, if arousal and mood produce the effect, then equally pleasant stimuli other than music should have the same result. …
Vernacularisation et traduction des textes pragmatiques en Afrique — La traduction des textes comportant des lacunes d'ordre grammatical, lexical, stylistique ou idiomatique présente habituellement des difficultés particulières, lesquelles sont amplifiées lorsqu'elles sont attribuables à la vernacularisation d'une langue étrangère. Dans les sociétés postcoloniales, l'absence ou la non-disponibilité des études linguistiques sur la plupart des langues locales rend ardue l'analyse des interférences entre ces dernières et les langues officielles étrangères. Cette situation, ajoutée à la grande diversité ethnolinguistique ambiante, ne facilite pas l'interprétation des textes produits par les personnes semi-lettrées. Le traducteur de ces textes se présente davantage comme un rédacteur qui, à partir de l'idée globale qui se dégage de l'original, conçoit et produit un texte répondant aux normes de la langue cible. L'évaluation d'un tel travail ne peut se faire qu'en comparant la finalité des deux textes.
In the paper, we describe methods for exploitation of a new lexical database of valency frames (VerbaLex) in relation to Transparent Intensional Logic (TIL). We present a detailed description of the Complex Valency Frames (CVF) as they appear in VerbaLex including basic ontology of the VerbaLex semantic roles.
The correct attachment of prepositional phrases (PPs) is a central disambiguation problem when parsing natural languages. This paper compares the baseline situation for French as exemplified in the Le Monde treebank with earlier findings for English, German and Swedish. We perform uniform treebank queries and show that the noun attachment rate for French prepositions is strongly influenced by the preposition de which is by far the most frequent preposition and has a strong tendency for noun attachment. We therefore also compute the noun attachment rate for the other prepositions separately as well as for the many complex prepositions that are explicitly marked in this treebank.
We describe how the British National Corpus (BNC), a one hundred million word balanced corpus of British English, was parsed into Lexical Functional Grammar (LFG) c-structures and f-structures, using a treebank-based \nparsing architecture. The parsing architecture uses a state-of-the-art statistical parser and reranker trained on the Penn Treebank to produce context-free phrase structure trees, and an annotation algorithm to automatically annotate \nthese trees into LFG f-structures. We describe the pre-processing steps which were taken to accommodate the differences between the Penn Treebank and the BNC. Some of the issues encountered in applying the parsing \narchitecture on such a large scale are discussed. The process of annotating a gold standard set of 1,000 parse trees is described. We present evaluation results obtained by evaluating the c-structures produced by the statistical parser against the c-structure gold standard. We also present the results obtained by evaluating the f-structures produced by the annotation algorithm against an \nautomatically constructed f-structure gold standard. The c-structures achieve an f-score of 83.7% and the f-structures an f-score of 91.2%.
This licentiate thesis deals with automatic syntactic analysis, or parsing, of natural languages. A parser constructs the syntactic analysis, which it learns by looking at correctly analyzed sentences, known as training data. The general topic concerns manipulations of the training data in order to improve the parsing accuracy. Several studies using constituency-based theories for natural languages in such automatic and data-driven syntactic parsing have shown that training data, annotated according to a linguistic theory, often needs to be adapted in various ways in order to achieve an adequate, automatic analysis. A linguistically sound constituent structure is not necessarily well-suited for learning and parsing using existing data-driven methods. Modifications to the constituency-based trees in the training data, and corresponding modifications to the parser output, have successfully been applied to increase the parser accuracy. The topic of this thesis is to investigate whether similar modifications in the form of tree transformations to training data, annotated with dependency-based structures, can improve accuracy for data-driven dependency parsers. In order to do this, two types of tree transformations are in focus in this thesis. The first one concerns non-projectivity. The full potential of dependency parsing can only be realized if non-projective constructions are allowed, which pose a problem for projective dependency parsers. On the other hand, non-projective parsers tend, among other things, to be slower. In order to maintain the benefits of projective parsing, a tree transformation technique to recover non-projectivity while using a projective parser is presented here. The second type of transformation concerns linguistic phenomena that are possible but hard for a parser to learn, given a certain choice of dependency analysis. This study has concentrated on two such phenomena, coordination and verb groups, for which tree transformations are applied in order to improve parsing accuracy, in case the original structure does not coincide with a structure that is easy to learn. Empirical evaluations are performed using treebank data from various languages, and using more than one dependency parser. The results show that the benefit of these tree transformations used in preprocessing and postprocessing to a large extent is language, treebank and parser independent.
Youngster violence/violence on youngsters. Crossed views on the perception of verbal violence among immigrant school populations Among the different forms of violence, from and addressed to the youth, those exerted within the school framework are many and various. The form that has more particularly caught our attention, is the so-called normative violence connected to the language practices among immigrant school populations.Treating the issue of violence in relation to the norm, whether it be of a linguistic kind or another, is legitimate as far as the former constitutes and destroys the latter. Indeed, the school presents itself as « the place favouring a struggle to impose linguistic norms ». The way of speaking of school actors – a.o. learners and teachers – is an indicator of the relationship that the latter have with social rules in general and school norms in particular. Beyond the description of obscenities, language coarseness and vulgarity of the former, and of the French language « in the manner of Charles-Henri » of the others, these forms of verbal violence in the school framework are worth thinking about.Moreover, the social imagery related to immigration is so virulent that the equation « practice of languages different from "correct" French (the so-called "bon usage"), immigration and delinquency » may seem astonishing to some while it is admitted by others. The practice of the language would explain acts of delinquency: to act on the language would be preventive. « Young immigrants had better behave themselves », or even « they’d better talk correctly »! The eradication of violence seems to be at the cost of this simple solution.It is precisely beyond this naive optimism that we would like to go in this article, centred on the notion of youngster verbal violence, in the way it is perceived and lived in the school framework. To do so, the contribution of sociolinguistics can help in order to delimit and redefine language practices and to supply didactics of French with useful marks for efficient school norm teaching.
There are many methods to improve performance of statistical parsers. Resolving structural ambiguities is a major task of these methods. In the proposed approach, the parser produces a set of n-best trees based on a feature-extended PCFG grammar and then selects the best tree structure based on association strengths of dependency word-pairs. However, there is no sufficiently large Treebank producing reliable statistical distributions of all word-pairs. This paper aims to provide a self-learning method to resolve the problems. The word association strengths were automatically extracted and learned by parsing a giga-word corpus. Although the automatically learned word associations were not perfect, the constructed structure evaluation model improved the bracketed f-score from 83.09% to 86.59%. We believe that the above iterative learning processes can improve parsing performances automatically by learning word-dependence information continuously from web.
Our paper reports an attempt to apply an unsupervised clustering algorithm to a Hungarian treebank in order to obtain semantic verb classes. Starting from the hypothesis that semantic metapredicates underlie verbs' syntactic realization, we investigate how one can obtain semantically motivated verb classes by automatic means. The 150 most frequent Hungarian verbs were clustered on the basis of their complementation patterns, yielding a set of basic classes and hints about the features that determine verbal subcategorization. The resulting classes serve as a basis for the subsequent analysis of their alternation behavior.
As an extension of decades of syntactic theorizing, treebanks have inherited a small set of phrasal categories, which abstract over the environments that the categories can occur in. Extending ideas from Johnson (1998), we explore encoding information from the local tree context in each category. We then discuss two clustering techniques which preserve the distributionally relevant category distinction, forming linguistically relevant generalizations and improving PCFG parsing performance. 1
Multilingual lexicons are needed in various applications, such as cross-lingual information retrieval, machine translation, and some others. Often, these applications suffer from the ambiguity of dictionary items, especially when an intermediate natural language is involved in the process of the dictionary construction, since this language adds its ambiguity to the ambiguity of working languages. This paper aims to propose a new method for producing multilingual dictionaries without the risk of introducing additional ambiguity. As a disambiguated intermediate language we use the so-called Universal Words. A set of more than 200,000 unambiguous Universal Words have been constructed automatically on the basis of the well-known English lexical database WordNet. This approach is being used for the construction of a five language-dictionary in the field of cultural heritage within the framework of the PATRILEX project sponsored by the Spanish Research Council.
In 1993 the Ministers of Education in the Netherlands and Flanders decided to install a binational committee of experts in order to co-ordinate, streamline, improve and stimulate the production of bilingual dictionaries and lexical databases with Dutch as a source or target language. This committee, called Commissie voor Lexicografische Vertaalvoorzieningen (Committee for Interlingual Lexicographical Resources) or CLVV, has, under the presidency of W. Martin, set up Action Plans involving some twenty dictionary projects which have been finished or are nearly finished by now. In this article the general policy lines of the CLVV are presented next to the criteria for the selection of language pairs, the infrastructure used, the results obtained and the lessons to be drawn from this ‘Dutch’ approach. The article also serves as a framework in which to situate the articles that follow.
In pluralistic nation such as ours, the function of government should be to foster and support the similarities that unite us, rather than institutionalize the differences that divide us. ProEnglish (http://www.proenglish.org/main/gen-info.htm) They tell me, 'Go back to Mexico, don't speak Spanish' Juan, Latino student at Junction High School Introduction The purpose of this article is to describe the prevalent linguistic ideology of certain members of dominant Euro-American group. (1) This linguistic ideology was encountered during an approximate eight month critical ethnographic action research project. In response to reported experiences of prejudice and racial discrimination by transnational newcomer students, seven teacher inquirers (2) engaged in an intercultural peace curricula development project that was facilitated by the author during the 2004-2005 school years at U.S Midwestern High School. Though the original dissertation research study design was not focused on mapping the prevalent linguistic ideology at Junction High School, attitudes about non-English language use quickly became central to our peacebuilding efforts. (3) Data presented here relays these attitudes as well as cultural assimilationist orientations exhibited by some students, teachers and administrators who were members of the dominant Euro-American population. The attitudes and linguistic normative monitoring of members of this dominant social group at Junction High School (4) created non-peaceful (5) school and classroom environment for newcomer students whose first languages included: Spanish; Japanese; Mandarin; and Arabic. Related research further examines everyday understandings of peace and non-peace at Junction High School (Brantmeier, 2007b) and also gives more in-depth description of the process of building intercultural empathy (Brantmeier, 2007a). In this article theoretical discussion of the terms linguistic ideology and cultural assimilation foregrounds description of the action research methodology employed in the dissertation study. Findings related to non-peaceful attitudes and behaviors, more specifically data related to attitudes about language and cultural assimilationist orientations, are then presented. A discussion follows that connects themes in the data analysis to wider cultural debates concerning language use and identity in the United States. Finally, call is made for further research that maps how dominant linguistic ideologies are enacted and countered. Theoretical Discussion: Linguistic Ideology and Cultural Assimilation Working conceptions are needed for the terms ideology and linguistic ideology. Apple (2004) describes functional understanding of ideology as a form of false consciousness which distorts one's perceptions of social reality and serves the interests of the dominant class in society (Apple 2004: 18-19). Understood in this light, an ideology is social construction that serves the interests of situated group of people within society; unequal power relationships are maintained through the propagation of an ideology. Apple focuses on class relations in the previous definition. The term linguistic ideology here is linked to broader focus on power and place, to race, to class, to regional dialects, to the language spoken, and to related status and power differentials in linguistically diverse environments. Rumsey (1990) describes linguistic ideology in terms of everyday understandings of language practices, or notions about the nature of language in the world (Rumsey 1990: 346). This commonsense understanding of right or correct language use can have consequences for those who lie outside the dominant linguistic norms. Thus, linguistic ideology can be understood here as dominant, everyday attitudes and practices concerning language use that serve to reinforce power and status differentials among members of population within situated social contexts. …
An approach for identifying the human source of a text by leveraging the significance of synonyms in language is presented. While others have attempted to identify authors in the past, they have focused on purely statistical approaches such as word length distribution, number of distinct words, and language models. We claim that an author's choice of synonyms is idiosyncratic and can be used in determining the identity of an author, which we demonstrate via our algorithm for recognizing authors. This algorithm uses synonym sets from the WordNet lexical database to give more weight to words that have many common synonyms. The results of this method applied to the task of identifying the authors of classic literature show that there is a correlation between an author's synonym choice and the author's identity. With this new author recognition technology, we may now explore new avenues of intelligent and meaningful interaction with users.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 61-72. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
This paper presents the first steps towards a statistical syntactic analyzer for Basque. The system is based on a syntactically dependency annotated treebank and an adaptation of the deterministic syntactic analyzer of Nivre et al. (2007), which relies on a shift/reduce deterministic analyzer together with a machine learning module that determines which one of 4 analysis options to take, giving a unique syntactic dependency analysis of an input sentence. The results are near to those obtained by similar systems.
We study the correlations in the connectivity patterns of large scale syntactic dependency networks. These networks are induced from treebanks: their vertices denote word forms which occur as nuclei of dependency trees. Their edges connect pairs of vertices if at least two instance nuclei of these vertices are linked in the dependency structure of a sentence. We examine the syntactic dependency networks of seven languages. In all these cases, we consistently obtain three findings. Firstly, clustering, i.e., the probability that two vertices which are linked to a common vertex are linked on their part, is much higher than expected by chance. Secondly, the mean clustering of vertices decreases with their degree — this finding suggests the presence of a hierarchical network organization. Thirdly, the mean degree of the nearest neighbors of a vertex x tends to decrease as the degree of x grows—this finding indicates disassortative mixing in the sense that links tend to connect vertices of dissimilar degrees. Our results indicate the existence of common patterns in the large scale organization of syntactic dependency networks.
<h3>Introduction</h3> Natural language applications like machine translation, question answering, and summarization currently are forced to depend on impoverished text models like bags of words or n-grams, while the decisions that they are making ought to be based on the meanings of those words in context. That lack of semantics causes problems throughout the applications. Misinterpreting the meaning of an ambiguous word results in failing to extract data, incorrect alignments for translation, and ambiguous language models. Incorrect coreference resolution results in missed information (because a connection is not made) or incorrectly conflated information (due to false connections). Some richer semantic representation is badly needed. The OntoNotes project is a collaborative effort between BBN Technologies, the University of Colorado, the University of Pennsylvania, and the University of Southern California's Information Sciences Institute to produce such a resource. It aims to annotate a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, use net, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information (syntax and predicate argument structure) and shallow semantics (word sense linked to an ontology and coreference). OntoNotes builds on two time-tested resources, following the Penn Treebank for syntax and the Penn PropBank for predicate-argument structure. Its semantic representation will include word sense disambiguation for nouns and verbs, with each word sense connected to an ontology, and coreference. The current goals call for annotation of over a million words each of English and Chinese, and half a million words of Arabic over five years. The authors wish to make this resource available to the natural language research community so that decoders for these phenomena can be trained to generate the same structure in new documents. Lessons learned over the years have shown that the quality of annotation is crucial if it is going to be used for training machine learning algorithms. Taking this cue, we ensure that each layer of annotation in OntoNotes will have at least 90% inter- annotator agreement. Our pilot studies have shown that predicate structure, word sense, ontology linking, and coreference can all be annotated rapidly and with better than 90% consistency. <h3>Samples</h3> The following screen captures provide examples of the data contained in this corpus. <ul> <li> <a href="./desc/addenda/LDC2007T21_eng_tbk.jpg" rel="nofollow">English tree</a>. </li> <li> <a href="./desc/addenda/LDC2007T21_sense_pred.jpg" rel="nofollow">English sense predicate structure</a>. </li> <li> <a href="./desc/addenda/LDC2007T21_chi_comp.jpg" rel="nofollow">Chinese tree and sense predicate structure</a>. </li> </ul><h3>Sponsorship</h3> This work was suppported in part by the Defense Research Advanced Projects Agency, GALE Program Grant No. HR0011-06-C-0022. The content of this publication does not necessarily reflect the position or policy of the Government, and no official endorsement should be inferred. </br> Portions © 1989 Dow Jones & Company, Inc., © 1996-2001 Sinorama Magazine, © 1994-1998 Xinhua News Agency, © 1995, 2005, 2006, 2007 Trustees of the University of Pennsylvania
In this paper we present a quantitative analysis of a bilingual lexical database which has been produced with OMBI, a tool for creating and editing bilingual dictionaries. OMBI has proven to be a valuable tool in the creation of rich bilingual multi-purpose lexical databases. One of the most distinctive features of the tool is reversal of source language and target language in order to create bilingual dictionaries in an economic and accurate way. We will focus on OMBI's reversal function, its initial concept and its results in practice. © 2007 Oxford University Press. All rights reserved.
This paper describes an effective approach to adapting an HPSG parser trained on the Penn Treebank to a biomedical domain. In this approach, we train probabilities of lexical entry assignments to words in a target domain and then incorporate them into the original parser. Experimental results show that this method can obtain higher parsing accuracy than previous work on domain adaptation for parsing the same data. Moreover, the results show that the combination of the proposed method and the existing method achieves parsing accuracy that is as high as that of an HPSG parser retrained from scratch, but with much lower training cost. We also evaluated our method in the Brown corpus to show the portability of our approach in another domain.
In this paper we propose an algorithm for converting dependency structures to phrase structures. This algorithm mainly concerns the characteristics of non-configurational languages. We review current works in the field and on the basis of these works we try to adopt a more flexible approach to the problem. 1
PURPOSE: This study examined the effect of socioeconomic status (SES) on the early lexical performance of African American children. METHOD: Thirty African American toddlers (30 to 40 months old) from low-SES (n = 15) and middle-SES (n = 15) backgrounds participated in the study. Their lexical-semantic performance was examined on 2 norm-referenced standardized tests of vocabulary, a measure of lexical diversity (number of different words) derived from language samples, and a fast mapping task that examined novel word learning. RESULTS: Toddlers from low-SES homes performed significantly poorer than those from middle-SES homes on standardized receptive and expressive vocabulary tests and on the number of different words used in spontaneous speech. No significant SES group differences were observed in their ability to learn novel word meanings on a fast mapping task. CONCLUSION: The influence of socioeconomic background on African American children's lexical semantic tasks varies with the type of measure used.
This paper presents a uniform approach to data extraction from syntactically annotated corpora encoded in XML. XQuery, which incorporates XPath, has been designed as a query language for XML. The combination of XPath and XQuery offers flexibility and expressive power, while corpus specific functions can be added to reduce the complexity of individual extraction tasks. We illustrate our approach using examples from dependency treebanks for Dutch.
Although using ontologies to assist information retrieval and text document processing has recently attracted more and more attention, existing ontologybased approaches have not shown advantages over the traditional keywords-based Latent Semantic Indexing (LSI) method. This paper proposes an algorithm to extract a concept forest (CF) from a document with the assistance of a natural language ontology, the WordNet lexical database. Using concept forests to represent the semantics of text documents, the semantic similarities of these documents are then measured as the commonalities of their concept forests. Performance studies of text document clustering based on different document similarity measurement methods show that the CF-based similarity measurement is an effective alternative to the existing keywords-based methods. In particular, this CFbased approach has obvious advantages over the existing keywords-based methods, including LSI, in processing short text documents or in P2P or live news environments where it is impractical to collect the entire document corpus for analysis.
How far can we get with unsupervised parsing if we make our training corpus several orders of magnitude larger than has hitherto be attempted? We present a new algorithm for unsupervised parsing using an all-subtrees model, termed U-DOP*, which parses directly with packed forests of all binary trees. We train both on Penn’s WSJ data and on the (much larger) NANC corpus, showing that U-DOP * outperforms a treebank-PCFG on the standard WSJ test set. While U-DOP * performs worse than state-of-the-art supervised parsers on handannotated sentences, we show that the model outperforms supervised parsers when evaluated as a language model in syntax-based machine translation on Europarl. We argue that supervised parsers miss the fluidity between constituents and non-constituents and that in the field of syntax-based language modeling the end of supervised parsing has come in sight. 1
Multiobjective evolutionary algorithms (MOEA) are an effective tool for solving search and optimization problems containing several incommensurable and possibly conflicting objectives. Unfortunately, many MOEAs face difficulties in solving problems when the number of objectives increases. In this paper, we investigate the efficacy of spatially structured MOEAs for scalable multiobjective problems. The algorithm is an extension of the standard cellular evolutionary algorithm, where the population is mapped to nodes of alternative complex networks. A selection regime based on a non-dominance rating and a crowding mechanism guides the evolutionary trajectory and an ε-dominance external archive is used to maintain a spread of solutions across the Pareto-optimal front. An important outcome of this work is the classification of the network models based on their impact on convergence speed and solution quality as the number of objectives increases for a given problem.
Based on the language of 17th century Bosnian Franciscan literature and enriched with features of the Neo-Štokavian folklore koine, the language of the 18th century writers represents a consistent system. Although it was not subject to willful codification, the language of the 18th century writers has codification elements. This is primarily implied by functional distribution, pronounced independence from common speech, especially on the syntactic level, compulsory use for all users and specific prescriptiveness in grammar handbooks of the time.
Iraj Mirza’s poetry occupies a special place in Persian literature as compared to the works of other poets of his time, due to his almost unrivalled use of language and rhetorics. Deviating from syntactic, semantic and pragmatic norms, he creates a new atmosphere with the simple language he uses, which draws his poetry close to the language of nature. The present paper examines Iraj Mirza’s poetry in terms of language function and his expert play with language. He deviates from the accepted linguistic norms of syntax, semantics and pragmatics, with an artistic courage, creating a new atmosphere in literary language: It is worth mentioning that he does so with such a simple language that one can claim, without unnecessary exaggeration that his poems are closer to the language of the nature than those of his contemporary poets. This article studies some of the language functions of Iraj's poems, revealing a small part of his skill in playing with the language.