Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We demonstrate that log-linear grammars with latent variables can be practically trained using discriminative methods. Central to efficient discriminative training is a hierarchical pruning procedure which allows feature expectations to be effi-ciently approximated in a gradient-based procedure. We compare L1 and L2 reg-ularization and show that L1 regularization is superior, requiring fewer iterations to converge, and yielding sparser solutions. On full-scale treebank parsing exper-iments, the discriminative latent models outperform both the comparable genera-tive latent models as well as the discriminative non-latent baselines. 1
Visual stimuli are judged for their emotional significance based on two fundamental dimensions, valence and arousal, and may lead to changes in neural and body functions like attention, affect, memory and heart rate. Alterations in behaviour and mood have been encountered in patients with Parkinson's disease (PD) undergoing functional neurosurgery, suggesting that electrical high-frequency stimulation of the subthalamic nucleus (STN) may interfere with emotional information processing. Here, we use the opportunity to directly record neuronal activity from the STN macroelectrodes in patients with PD during presentation of emotionally laden and neutral pictures taken from the International Affective Picture System (IAPS) to further elucidate the role of the STN in emotional processing. We found a significant event-related desynchronization of STN alpha activity with pleasant stimuli that correlated with the individual valence rating of the pictures. Our findings suggest involvement of the human STN in valence-related emotional information processing that can potentially be altered during high-frequency stimulation of the STN in PD leading to behavioural complications.
This paper examines whether a learningbased coreference resolver can be improved using semantic class knowledge that is automatically acquired from a version of the Penn Treebank in which the noun phrases are labeled with their semantic classes. Experiments on the ACE test data show that a resolver that employs such induced semantic class knowledge yields a statistically significant improvement of 2 % in F-measure over one that exploits heuristically computed semantic class knowledge. In addition, the induced knowledge improves the accuracy of common noun resolution by 2-6%. 1
Activity within the visual cortex can be influenced by the emotional salience of a stimulus, but it is not clear whether such cortical activity is modulated by the affective status of the individual. This study used functional magnetic resonance imaging (fMRI) to examine the relationship between affect ratings on the Positive and Negative Affect Schedule and activity within the occipital cortex of 13 normal-weight women while viewing images of high calorie and low calorie foods. Regression analyses revealed that when participants viewed high calorie foods, Positive Affect correlated significantly with activity within the lingual gyrus and calcarine cortex, whereas Negative Affect was unrelated to visual cortex activity. In contrast, during presentations of low calorie foods, affect ratings, regardless of valence, were unrelated to occipital cortex activity. These findings suggest a mechanism whereby positive affective state may affect the early stages of sensory processing, possibly influencing subsequent perceptual experience of a stimulus.
Previous studies in data-driven dependency parsing have shown that tree transformations can improve parsing accuracy for specific parsers and data sets. We investigate to what extent this can be generalized across languages/treebanks and parsers, focusing on pseudo-projective parsing, as a way of capturing non-projective dependencies, and transformations used to facilitate parsing of coordinate structures and verb groups. The results indicate that the beneficial effect of pseudo-projective parsing is independent of parsing strategy but sensitive to language or treebank specific properties. By contrast, the construction specific transformations appear to be more sensitive to parsing strategy but have a constant positive effect over several languages.
Abstract Affective judgments can often be influenced by emotional information people unconsciously perceive, but the neural mechanisms responsible for these effects and how they are modulated by individual differences in sensitivity to threat are unclear. Here we studied subliminal affective priming by recording brain potentials to surprise faces preceded by 30-msec happy or fearful prime faces. Participants showed valence-consistent changes in affective ratings of surprise faces, although they reported no knowledge of prime-face expressions, nor could they discriminate between prime-face expressions in a forced-choice test. In conjunction with the priming effect on affective evaluation, larger occipital P1 potentials at 145–175 msec were found with fearful than with happy primes, and source analyses implicated the bilateral extrastriate cortex in this effect. Later brain potentials at 300–400 msec were enhanced with happy versus fearful primes, which may reflect differential attentional orienting. Personality testing for sensitivity to threat, especially social threat, was also used to evaluate individual differences potentially relevant to subliminal affective priming. Indeed, participants with high trait anxiety demonstrated stronger affective priming and greater P1 differences than did those with low trait anxiety, and these effects were driven by fearful primes. Results thus suggest that unconsciously perceived affective information influences social judgments by altering very early perceptual analyses, and that this influence is accentuated to the extent that people are oversensitive to threat. In this way, perception may be subject to a variety of influences that govern social preferences in the absence of concomitant awareness of such influences.
In the previous chapter I defended the hypothesis of public linguistic norms by appeal to the nature of successful communication. The phenomenon to be accounted for was that whereby two participants in a speech exchange efficiently transmit very specific knowledge to one another through their speech, even under conditions in which they know nothing of each others' speech and interpretative dispositions save what is manifest in the brief speech exchange itself. The burden of chapter 2 was to argue that the reliable comprehension that is attained in such cases seems miraculous unless we suppose that there are such norms – norms which (given the concrete speech context) determine what the speaker literally said with her words, and which are at least implicitly exploited by the hearer in the process by which she arrives at a representation of the content of the speech she observed. Whereas that argument for public linguistic norms is thus an argument from Successful Communication, the present argument, in contrast, will be from a certain kind of unsuccessful communication – that arising from cases in which, though the speaker aims to be communicating knowledge, this aim is thwarted owing to misunderstanding. My central thesis will be that in these sorts of misunderstanding cases at least one of the parties is appropriately blamed for the breakdown in communication, and that (at least in RC-cases) warranted ascriptions of blame presuppose public linguistic norms.
We present a novel method for evaluating the output of Machine Translation (MT), based on comparing the dependency structures of the translation and reference rather than their surface string forms. Our method uses a treebank-based, widecoverage, probabilistic Lexical-Functional Grammar (LFG) parser to produce a set of structural dependencies for each translation-reference sentence pair, and then calculates the precision and recall for these dependencies. Our dependency-based evaluation, in contrast to most popular string-based evaluation metrics, will not unfairly penalize perfectly valid syntactic variations in the translation. In addition to allowing for legitimate syntactic differences, we use paraphrases in the evaluation process to account for lexical variation. In comparison with other metrics on 16,800 sentences of Chinese-English newswire text, our method reaches high correlation with human scores. An experiment with two translations of 4,000 sentences from Spanish-English Europarl shows that, in contrast to most other metrics, our method does not display a high bias towards statistical models of translation.
Historically, unsupervised learning techniques have lacked a principled technique for selecting the number of unseen components. Research into non-parametric priors, such as the Dirichlet process, has enabled instead the use of infinite models, in which the number of hidden categories is not fixed, but can grow with the amount of training data. Here we develop the infinite tree, a new infinite model capable of representing recursive branching structure over an arbitrarily large set of hidden categories. Specifically, we develop three infinite tree models, each of which enforces different independence assumptions, and for each model we define a simple direct assignmentsampling inference procedure. We demonstrate the utility of our models by doing unsupervised learning of part-of-speech tags from treebank dependency skeleton structure, achieving an accuracy of 75.34%, and by doing unsupervised splitting of part-of-speech tags, which increases the accuracy of a generative dependency parser from 85.11% to 87.35%.
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank. MorphoTrees is the idea of building effective and intuitive hierarchies over the information provided by computational morphological systems. MorphoTrees are implemented for Arabic as an extension to the TrEd annotation environment based on Perl. Encode Arabic libraries for Haskell and Perl serve for processing the non-trivial and multi-purpose ArabTEX notation that encodes Arabic orthographies and phonetic transcriptions in parallel.
The percentage of pupils fluent in both Swedish and Finnish in Swedish-speaking schools in Finland has grown and now includes a third of all pupils. This article focuses on the linguistic classroom discourse in a Swedish-speaking school in a strongly Finnish-dominated area. The purpose is to increase the understanding of how linguistic norms are maintained, and what these norms imply for the participation of bilingual pupils in classroom interaction. Videotaped lessons were analyzed by means of conversation analysis focusing on the interaction between the teacher and the bilingual pupils as well as on language-related sequences. The results show that the bilingual pupils cannot be regarded as victims of a language policy governed from above, but that they actively contribute to the construction and maintenance of a monolingual norm in the classroom. When using Finnish, they at the same time point at the “other-languageness” of the code-switched words. Monolingualism as a norm means a limitation restraining the pupils with gaps in their Swedish from participating with full competence in the classroom conversation. Between pupils, occasional Finnish words are used in an unproblematic manner. Code-switching in these cases works as a means to keep one's position on the conversational floor. By violating the monolingual norm, pupils can lodge a protest against the agenda of the teacher.
In the present chapter, my aim is to use the results of the previous two chapters to argue for three anti-individualistic doctrines in the philosophy of mind and language. These doctrines express anti-individualistic theses regarding speech content, linguistic meaning, and mental content/attitude individuation. The arguments themselves all share their basic structure: appealing to a thought experiment in which we vary the public linguistic norms in play while leaving intact all individualistic facts regarding the participants in a speech exchange, it is argued that these three properties – the content of the speech, the linguistic meanings of the expressions used, and the contents of the beliefs acquired in the exchange – will vary in a way reflecting the change in public linguistic norms. The result, of course, will be that the instantiation of the determinate properties in question is not fixed by the individualistic facts regarding the participants in a speech exchange. The facts regarding speech content etc. do not supervene on the set of individualistic facts regarding the speaker and hearer, respectively. (Putting the point in terms of supervenience connects our result with the standard way of formulating anti-individualistic doctrines.)
In order to determine novel information from raw text documents, a novelty detection recommender system was developed to explore the method of comparing various types of entities within sentences. We first detected novel sentences using named entity recognition to extract the entity types of person, place, time, and organization. In addition, part-of-speech tagging was performed to tag each word in the documents, allowing syntactic structures of noun, verb, and adjective to be used for comparisons. WordNet, an English lexical database of concepts and relations, was also incorporated to generate synonyms for the entities and parts of speech, as well as to determine the similarity of sentences. The novelty score of each sentence was determined by using two different metrics, UniqueComparison and ImportanceValue. UniqueComparison calculated the number of matched entities, whereas ImportanceValue took into account the total weight of matched words that coexisted in both the test and history sentences. The results look promising when compared to the benchmark scores for the Text Retrieval Conference’s (TREC) Novelty Track 2004. This demonstrated that the combination of named entity recognition and part-of-speech tagging is capable of detecting novelty with good results.
This paper investigates how the use of machine learning techniques can significantly predict the three major dimensions of learner-s emotions (pleasure, arousal and dominance) from brainwaves. This study has adopted an experimentation in which participants were exposed to a set of pictures from the International Affective Picture System (IAPS) while their electrical brain activity was recorded with an electroencephalogram (EEG). The pictures were already rated in a previous study via the affective rating system Self-Assessment Manikin (SAM) to assess the three dimensions of pleasure, arousal, and dominance. For each picture, we took the mean of these values for all subjects used in this previous study and associated them to the recorded brainwaves of the participants in our study. Correlation and regression analyses confirmed the hypothesis that brainwave measures could significantly predict emotional dimensions. This can be very useful in the case of impassive, taciturn or disabled learners. Standard classification techniques were used to assess the reliability of the automatic detection of learners- three major dimensions from the brainwaves. We discuss the results and the pertinence of such a method to assess learner-s emotions and integrate it into a brainwavesensing Intelligent Tutoring System.
This paper tests three factors that have been held to be responsible for the variable stress behavior of noun-noun constructs in English: argument structure, semantics, and analogy. In a large-scale investigation of some 4500 compounds extracted from the CELEX lexical database (Baayen et al. 1995), we show that traditional claims about noun-noun stress cannot be upheld. Argument structure plays a role only with synthetic compounds ending in the agentive suffix - er. The semantic categories and relations assumed in the literature to trigger rightward stress do not show the expected effects. As an alternative to the rule-based approaches, the data were modeled computationally and probabilistically using a memory-based analogical algorithm (TiMBL 5.1) and logistic regression, respectively. It turns out that probabilistic models and the analogical algorithm are more successful in predicting stress assignment correctly than any of the rules proposed in the literature. Furthermore, the results of the analogical modeling suggest that the left and right constituent are the most important factor in compound stress assignment. This is in line with recent findings on the semi-regular behavior of compounds in other languages.
Multiobjective evolutionary algorithms (MOEA) are an effective tool for solving search and optimization problems containing several incommensurable and possibly conflicting objectives. Unfortunately, many MOEAs face difficulties in solving problems when the number of objectives increases. In this paper, we investigate the efficacy of spatially structured MOEAs for scalable multiobjective problems. The algorithm is an extension of the standard cellular evolutionary algorithm, where the population is mapped to nodes of alternative complex networks. A selection regime based on a non-dominance rating and a crowding mechanism guides the evolutionary trajectory and an ε-dominance external archive is used to maintain a spread of solutions across the Pareto-optimal front. An important outcome of this work is the classification of the network models based on their impact on convergence speed and solution quality as the number of objectives increases for a given problem.
This paper reports on a hybrid architecture for computational anaphora resolution (CAR) of German that combines a rule-based pre-filtering component with a memory-based resolution module (using the Tilburg Memory Based Learner – TiMBL). The data source is provided by the TüBa-D/Z treebank of German newspaper text (Telljohann et al. 04) that is annotated with anaphoric relations. The CAR experiments performed on these treebank data corroborate the importance of modelling aspects of discourse structure for robust, data-driven anaphora resolution. The best result with an F-measure of 0.734 achieved by these experiments outperforms the results reported by (Schiehlen 04), the only other study of German CAR that is based on newspaper treebank data. 1
In Example Based Machine Translation research, many researchers apply Structure Based EBMT approach to annotate sentence structure in tree format. During training process, corpus become large and large and human tagging in Treebank creation becomes unrealistic. Therefore, it needs a tool to simplify and unify tagging process in order to enhance tagging performance and ensure sentence tree correctness. In this paper, we propose an automatic tool to create Treebank in TCT annotation schema.
We present results that show that incorporating lexical and structural semantic information is effective for word sense disambiguation. We evaluated the method by using precise information from a large treebank and an ontology automatically created from dictionary sentences. Exploiting rich semantic and structural information improves precision 2–3%. The most gains are seen with verbs, with an improvement of 5.7% over a model using only bag of words and n-gram features.
In this paper we present a quantitative analysis of a bilingual lexical database which has been produced with OMBI, a tool for creating and editing bilingual dictionaries. OMBI has proven to be a valuable tool in the creation of rich bilingual multi-purpose lexical databases. One of the most distinctive features of the tool is reversal of source language and target language in order to create bilingual dictionaries in an economic and accurate way. We will focus on OMBI's reversal function, its initial concept and its results in practice. © 2007 Oxford University Press. All rights reserved.
<h3>Introduction</h3> Natural language applications like machine translation, question answering, and summarization currently are forced to depend on impoverished text models like bags of words or n-grams, while the decisions that they are making ought to be based on the meanings of those words in context. That lack of semantics causes problems throughout the applications. Misinterpreting the meaning of an ambiguous word results in failing to extract data, incorrect alignments for translation, and ambiguous language models. Incorrect coreference resolution results in missed information (because a connection is not made) or incorrectly conflated information (due to false connections). Some richer semantic representation is badly needed. The OntoNotes project is a collaborative effort between BBN Technologies, the University of Colorado, the University of Pennsylvania, and the University of Southern California's Information Sciences Institute to produce such a resource. It aims to annotate a large corpus comprising various genres of text (news, conversational telephone speech, weblogs, use net, broadcast, talk shows) in three languages (English, Chinese, and Arabic) with structural information (syntax and predicate argument structure) and shallow semantics (word sense linked to an ontology and coreference). OntoNotes builds on two time-tested resources, following the Penn Treebank for syntax and the Penn PropBank for predicate-argument structure. Its semantic representation will include word sense disambiguation for nouns and verbs, with each word sense connected to an ontology, and coreference. The current goals call for annotation of over a million words each of English and Chinese, and half a million words of Arabic over five years. The authors wish to make this resource available to the natural language research community so that decoders for these phenomena can be trained to generate the same structure in new documents. Lessons learned over the years have shown that the quality of annotation is crucial if it is going to be used for training machine learning algorithms. Taking this cue, we ensure that each layer of annotation in OntoNotes will have at least 90% inter- annotator agreement. Our pilot studies have shown that predicate structure, word sense, ontology linking, and coreference can all be annotated rapidly and with better than 90% consistency. <h3>Samples</h3> The following screen captures provide examples of the data contained in this corpus. <ul> <li> <a href="./desc/addenda/LDC2007T21_eng_tbk.jpg" rel="nofollow">English tree</a>. </li> <li> <a href="./desc/addenda/LDC2007T21_sense_pred.jpg" rel="nofollow">English sense predicate structure</a>. </li> <li> <a href="./desc/addenda/LDC2007T21_chi_comp.jpg" rel="nofollow">Chinese tree and sense predicate structure</a>. </li> </ul><h3>Sponsorship</h3> This work was suppported in part by the Defense Research Advanced Projects Agency, GALE Program Grant No. HR0011-06-C-0022. The content of this publication does not necessarily reflect the position or policy of the Government, and no official endorsement should be inferred. </br> Portions © 1989 Dow Jones & Company, Inc., © 1996-2001 Sinorama Magazine, © 1994-1998 Xinhua News Agency, © 1995, 2005, 2006, 2007 Trustees of the University of Pennsylvania
We aim to improve the performance of a syntactic parser that uses a part-of-speech (POS) tagger as a preprocessor. Pipelined parsers consisting of POS taggers and syntactic parsers have several advantages, such as the capability of domain adaptation. However the performance of such systems on raw texts tends to be disappointing as they are affected by the errors of automatic POS tagging. We attempt to compensate for the decrease in accuracy caused by automatic taggers by allowing the taggers to output multiple answers when the tags cannot be determined reliably enough. We empirically verify the effectiveness of the method using an HPSG parser trained on the Penn Treebank. Our results show that ambiguous POS tagging improves parsing if outputs of taggers are weighted by probability values, and the results support previous studies with similar intentions. We also examine the effectiveness of our method for adapting the parser to the GENIA corpus and show that the use of ambiguous POS taggers can help development of portable parsers while keeping accuracy high. 1
Due to the data sparseness problem, the lexical information from a treebank for a lexicalized parser could be insufficient. This paper proposes an approach to learn head-modifier pairs from a raw corpus, and to integrate them into a lexicalized dependency parser to parse a Chinese Treebank. Experimental re-sults show that this approach not only enlarged the coverage of bi-lexical de-pendency, but also improved the accuracy of dependency parsing significantly.
This paper describes practical issues in the framework-independent evaluation of deep and shallow parsers. We focus on the use of two dependencybased syntactic representation formats in parser evaluation, namely, Carroll et al. (1998)’s Grammatical Relations and de Marneffe et al. (2006)’s Stanford Dependency scheme. Our approach is to convert the output of parsers into these two formats, and measure the accuracy of the resulting converted output. Through the evaluation of an HPSG parser and Penn Treebank phrase structure parsers, we found that mapping between different representation schemes is a non-trivial task that results in lossy conversions that may obscure important differences between different parsing approaches. We discuss sources of disagreements in the representation of syntactic structures in the two dependency-based formats, indicating possible directions for improved framework-independent parser evaluation.
We study the correlations in the connectivity patterns of large scale syntactic dependency networks. These networks are induced from treebanks: their vertices denote word forms which occur as nuclei of dependency trees. Their edges connect pairs of vertices if at least two instance nuclei of these vertices are linked in the dependency structure of a sentence. We examine the syntactic dependency networks of seven languages. In all these cases, we consistently obtain three findings. Firstly, clustering, i.e., the probability that two vertices which are linked to a common vertex are linked on their part, is much higher than expected by chance. Secondly, the mean clustering of vertices decreases with their degree — this finding suggests the presence of a hierarchical network organization. Thirdly, the mean degree of the nearest neighbors of a vertex x tends to decrease as the degree of x grows—this finding indicates disassortative mixing in the sense that links tend to connect vertices of dissimilar degrees. Our results indicate the existence of common patterns in the large scale organization of syntactic dependency networks.
We compare the accuracy of a statistical parse ranking model trained from a fully-annotated portion of the Susanne treebank with one trained from unlabeled partially-bracketed sentences derived from this treebank and from the Penn Treebank. We demonstrate that confidence-based semi-supervised techniques similar to self-training outperform expectation maximization when both are constrained by partial bracketing. Both methods based on partially-bracketed training data outperform the fully supervised technique, and both can, in principle, be applied to any statistical parser whose output is consistent with such partial-bracketing. We also explore tuning the model to a different domain and the effect of in-domain data in the semi-supervised training processes.
This paper presents the first steps towards a statistical syntactic analyzer for Basque. The system is based on a syntactically dependency annotated treebank and an adaptation of the deterministic syntactic analyzer of Nivre et al. (2007), which relies on a shift/reduce deterministic analyzer together with a machine learning module that determines which one of 4 analysis options to take, giving a unique syntactic dependency analysis of an input sentence. The results are near to those obtained by similar systems.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 61-72. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
An approach for identifying the human source of a text by leveraging the significance of synonyms in language is presented. While others have attempted to identify authors in the past, they have focused on purely statistical approaches such as word length distribution, number of distinct words, and language models. We claim that an author's choice of synonyms is idiosyncratic and can be used in determining the identity of an author, which we demonstrate via our algorithm for recognizing authors. This algorithm uses synonym sets from the WordNet lexical database to give more weight to words that have many common synonyms. The results of this method applied to the task of identifying the authors of classic literature show that there is a correlation between an author's synonym choice and the author's identity. With this new author recognition technology, we may now explore new avenues of intelligent and meaningful interaction with users.
In pluralistic nation such as ours, the function of government should be to foster and support the similarities that unite us, rather than institutionalize the differences that divide us. ProEnglish (http://www.proenglish.org/main/gen-info.htm) They tell me, 'Go back to Mexico, don't speak Spanish' Juan, Latino student at Junction High School Introduction The purpose of this article is to describe the prevalent linguistic ideology of certain members of dominant Euro-American group. (1) This linguistic ideology was encountered during an approximate eight month critical ethnographic action research project. In response to reported experiences of prejudice and racial discrimination by transnational newcomer students, seven teacher inquirers (2) engaged in an intercultural peace curricula development project that was facilitated by the author during the 2004-2005 school years at U.S Midwestern High School. Though the original dissertation research study design was not focused on mapping the prevalent linguistic ideology at Junction High School, attitudes about non-English language use quickly became central to our peacebuilding efforts. (3) Data presented here relays these attitudes as well as cultural assimilationist orientations exhibited by some students, teachers and administrators who were members of the dominant Euro-American population. The attitudes and linguistic normative monitoring of members of this dominant social group at Junction High School (4) created non-peaceful (5) school and classroom environment for newcomer students whose first languages included: Spanish; Japanese; Mandarin; and Arabic. Related research further examines everyday understandings of peace and non-peace at Junction High School (Brantmeier, 2007b) and also gives more in-depth description of the process of building intercultural empathy (Brantmeier, 2007a). In this article theoretical discussion of the terms linguistic ideology and cultural assimilation foregrounds description of the action research methodology employed in the dissertation study. Findings related to non-peaceful attitudes and behaviors, more specifically data related to attitudes about language and cultural assimilationist orientations, are then presented. A discussion follows that connects themes in the data analysis to wider cultural debates concerning language use and identity in the United States. Finally, call is made for further research that maps how dominant linguistic ideologies are enacted and countered. Theoretical Discussion: Linguistic Ideology and Cultural Assimilation Working conceptions are needed for the terms ideology and linguistic ideology. Apple (2004) describes functional understanding of ideology as a form of false consciousness which distorts one's perceptions of social reality and serves the interests of the dominant class in society (Apple 2004: 18-19). Understood in this light, an ideology is social construction that serves the interests of situated group of people within society; unequal power relationships are maintained through the propagation of an ideology. Apple focuses on class relations in the previous definition. The term linguistic ideology here is linked to broader focus on power and place, to race, to class, to regional dialects, to the language spoken, and to related status and power differentials in linguistically diverse environments. Rumsey (1990) describes linguistic ideology in terms of everyday understandings of language practices, or notions about the nature of language in the world (Rumsey 1990: 346). This commonsense understanding of right or correct language use can have consequences for those who lie outside the dominant linguistic norms. Thus, linguistic ideology can be understood here as dominant, everyday attitudes and practices concerning language use that serve to reinforce power and status differentials among members of population within situated social contexts. …
In 1993 the Ministers of Education in the Netherlands and Flanders decided to install a binational committee of experts in order to co-ordinate, streamline, improve and stimulate the production of bilingual dictionaries and lexical databases with Dutch as a source or target language. This committee, called Commissie voor Lexicografische Vertaalvoorzieningen (Committee for Interlingual Lexicographical Resources) or CLVV, has, under the presidency of W. Martin, set up Action Plans involving some twenty dictionary projects which have been finished or are nearly finished by now. In this article the general policy lines of the CLVV are presented next to the criteria for the selection of language pairs, the infrastructure used, the results obtained and the lessons to be drawn from this ‘Dutch’ approach. The article also serves as a framework in which to situate the articles that follow.
This paper describes an automatic prediction model of Chinese prepositional phrase boundary location based on HMM.It consists of two stages: automatically identify the phrase boundary using statistics from treebank,then,post-tune the results with dependency grammar knowledge generated by dependency treebank.Experimental results demonstrate a high rate of success for predicting boundary location(86.5% correct rate for close testing and 77.7% for open testing).
By raising the question of folk linguistics in the discussion of linguistic phenomena, it is possible to abandon the typically binary approach to norms (descriptive versus prescriptive), and to propose a third term: perception. In France, scientific truth is still defined along very Cartesian lines, and linguistic norms remain monopolized by purism and academic orthodoxy. As a result, perceptive norms has not been allowed to emerge as a field of research. And yet, data drawn from non-academic linguistic discourse could make a considerable contribution to work on linguistic norms and varieties of language. It proves worthy of study not as a naive, pre-scientific effort, but as one possible theory of language—thus causing us to rethink the way the rules of the system actually function, beyond their representations. Intuition, urbane linguistics, folk linguistics, proper noun, perception, folk sociolinguistics.
Multilingual lexicons are needed in various applications, such as cross-lingual information retrieval, machine translation, and some others. Often, these applications suffer from the ambiguity of dictionary items, especially when an intermediate natural language is involved in the process of the dictionary construction, since this language adds its ambiguity to the ambiguity of working languages. This paper aims to propose a new method for producing multilingual dictionaries without the risk of introducing additional ambiguity. As a disambiguated intermediate language we use the so-called Universal Words. A set of more than 200,000 unambiguous Universal Words have been constructed automatically on the basis of the well-known English lexical database WordNet. This approach is being used for the construction of a five language-dictionary in the field of cultural heritage within the framework of the PATRILEX project sponsored by the Spanish Research Council.
Our paper reports an attempt to apply an unsupervised clustering algorithm to a Hungarian treebank in order to obtain semantic verb classes. Starting from the hypothesis that semantic metapredicates underlie verbs' syntactic realization, we investigate how one can obtain semantically motivated verb classes by automatic means. The 150 most frequent Hungarian verbs were clustered on the basis of their complementation patterns, yielding a set of basic classes and hints about the features that determine verbal subcategorization. The resulting classes serve as a basis for the subsequent analysis of their alternation behavior.
The ability to detect similarity in conjunct heads is potentially a useful tool in helping to disambiguate coordination structures - a difficult task for parsers. We propose a distributional measure of similarity designed for such a task. We then compare several different measures of word similarity by testing whether they can empirically detect similarity in the head nouns of noun phrase conjuncts in the Wall Street Journal (WSJ) treebank. We demonstrate that several measures of word similarity can successfully detect conjunct head similarity and suggest that the measure proposed in this paper is the most appropriate for this task.
This licentiate thesis deals with automatic syntactic analysis, or parsing, of natural languages. A parser constructs the syntactic analysis, which it learns by looking at correctly analyzed sentences, known as training data. The general topic concerns manipulations of the training data in order to improve the parsing accuracy. Several studies using constituency-based theories for natural languages in such automatic and data-driven syntactic parsing have shown that training data, annotated according to a linguistic theory, often needs to be adapted in various ways in order to achieve an adequate, automatic analysis. A linguistically sound constituent structure is not necessarily well-suited for learning and parsing using existing data-driven methods. Modifications to the constituency-based trees in the training data, and corresponding modifications to the parser output, have successfully been applied to increase the parser accuracy. The topic of this thesis is to investigate whether similar modifications in the form of tree transformations to training data, annotated with dependency-based structures, can improve accuracy for data-driven dependency parsers. In order to do this, two types of tree transformations are in focus in this thesis. The first one concerns non-projectivity. The full potential of dependency parsing can only be realized if non-projective constructions are allowed, which pose a problem for projective dependency parsers. On the other hand, non-projective parsers tend, among other things, to be slower. In order to maintain the benefits of projective parsing, a tree transformation technique to recover non-projectivity while using a projective parser is presented here. The second type of transformation concerns linguistic phenomena that are possible but hard for a parser to learn, given a certain choice of dependency analysis. This study has concentrated on two such phenomena, coordination and verb groups, for which tree transformations are applied in order to improve parsing accuracy, in case the original structure does not coincide with a structure that is easy to learn. Empirical evaluations are performed using treebank data from various languages, and using more than one dependency parser. The results show that the benefit of these tree transformations used in preprocessing and postprocessing to a large extent is language, treebank and parser independent.
We describe how the British National Corpus (BNC), a one hundred million word balanced corpus of British English, was parsed into Lexical Functional Grammar (LFG) c-structures and f-structures, using a treebank-based \nparsing architecture. The parsing architecture uses a state-of-the-art statistical parser and reranker trained on the Penn Treebank to produce context-free phrase structure trees, and an annotation algorithm to automatically annotate \nthese trees into LFG f-structures. We describe the pre-processing steps which were taken to accommodate the differences between the Penn Treebank and the BNC. Some of the issues encountered in applying the parsing \narchitecture on such a large scale are discussed. The process of annotating a gold standard set of 1,000 parse trees is described. We present evaluation results obtained by evaluating the c-structures produced by the statistical parser against the c-structure gold standard. We also present the results obtained by evaluating the f-structures produced by the annotation algorithm against an \nautomatically constructed f-structure gold standard. The c-structures achieve an f-score of 83.7% and the f-structures an f-score of 91.2%.