Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Within the STEVIN1 project Large Scale Syntactic Annotation of written Dutch (LASSY), a manually corrected treebank of 1 million words is constructed. Lassy is part of a series of annotation projects for modern written and spoken Dutch. More specifically, it is an extension of the D-Coi and CGN projects,2 and constitutes the core of SoNaR, a 500 million words reference corpus of modern written Dutch.3 One of the goals of the latter project is to enrich the corrected treebank produced in Lassy4 with several semantic layers. For a general overview of the relations between D-Coi, Lassy and SoNaR, cf [19]. In this paper we will concentrate on the semantic layers of SoNaR core: (1) named entity labeling, (2) annotation of co-reference relations, (3) semantic role labeling and (4) annotation of spatial and temporal relations. Of these (2) originates from the STEVIN-project COREA,5 (3) and (4) from D-Coi, whereas (1) is a new area within STEVIN.
The Postmodern culture today breaks down historically a solid barrier, produced in the modern age, between the language and its users, signs and realities and subject and his object. This phenomenon brings about some translation problems at the same time; interpretative diversity, linguistic derivation in the mass media, permanent reproduction of translated text and meaning`s continuity, definition of translator etc. I examine these problems through a theoretical approach to the mediative nature of the act of translation. I stress on two inevitable aspects of the postmodern linguistic tendencies: connotation generalized in our translation activities and its re-mediative culture. Focusing on a cycling perspective of our translating activities, I could reach the conclusion that the translated text have no relation with the denotative meanings which have been comprised closed and original with the realities from the modern age. A corrected model could be proposed. I call this cycle model of translating process which comprises ① Mediation or Creation ② Re-mediation or Translation ③ Re-re-mediation or Comprehension ④ Verification, Deduction or Correction. Each activity has not only its own translating process but its circulated role for the time in which everyone can participate equally as a reader, a sender, a translator and an individual who makes its contextual needs and desires. This model could show that the translating process is not for fixing a linguistic sign to some closed meanings but for expanding its pertinent meanings to diverse situations. Translating activity is not for making a linguistic norm by itself, but for making appropriate communication with as much of the population as possible.
Chemical sensitivity (CS) is common in the adult population and implies negative effects (e.g. physical symptoms, negative affects and behavioral disruptions) of odorous chemical substances. In Study 1, relations between self-reported CS, negative affect and neuroticism were investigated among Swedish university students (n = 103). CS and neuroticism were positively correlated, suggesting that highly neurotic persons, compared to low- neurotic persons, more easily respond negatively to environmental odors. In Study 2 (n = 40), relations between CS, self-reported noise sensitivity, physical symptoms and odor perception (pleasantness, intensity and familiarity) were examined in a sub-sample of high (HCS) and low (LCS) chemical sensitivity. The HCS group reported higher odor intensity and noise sensitivity. No differences were found in odor pleasantness and familiarity ratings, and in physical symptoms. Overall, the results suggest that normal variation in CS has a general, sensory-unspecific basis.
Linguistic items are meaningful and most of our uses of them are meaningful too. In order for this to be the case, language is bound by a set of correctness-conditions according to which uses of language can be categorized as either correct or incorrect – that is, there are linguistic rules – or so it is at least plausible to suppose. Our question now is not whether or not language is meaningful and normative in this sense but whether those correctness-conditions normatively bind speakers' use of language: does this normativity of language entail the normativity of linguistic usage? Ought speakers of a language to abide by its correctness-conditions? Do linguistic rules entail linguistic norms, or, more pithily, do linguistic rules rule? I shall argue that provided we focus on the appropriate correctness-conditions – correctness-conditions framed at the level of sense rather than that of reference – these do indeed deliver oughts governing use.
The Turin University Treebank (TUT) is a treebank with dependency-based annotations of 2,400 Italian sentences. By converting TUT to binary constituency trees, it is possible to produce a treebank of derivations of Combinatory Categorial Grammar (CCG), with an algorithm that traverses a tree in a top-down manner, employing a stack to record argument structure, using Part of Speech tags to determine the lexical categories. This method reaches a coverage of 77%, resulting in a CCGbank for Italian comprising 1,837 sentences, with an average length of 22,9 tokens. The CCGbank for English has proven to be a useful tool for developing efficient wide-coverage parsers for semantic interpretation, and the Italian CCGbank is expected to be an equally useful linguistic resource for training statistical parsers.
We propose in this paper a formalization of the BDef lexicographic definitions pro- posed by (Altman et Polguere, 2003), in order to carry calculus on them. The information richness of such definitions suggest many calculus that are useful for lexicographic practice as well as for a more general reflection on modelling lexical semantics. Such calculus cannot be realized without a thorough formalization of such definitions that we propose to represent as typed feature structures (Carpenter, 1992). The formalization proposed allow to automati- cally check for the lexicon consistency, which suggest a new methodology for lexical description based on successive enrichment of the lexical database and the meta data that describe it. MOTS-CLES: definitions lexicographiques formalisees, regles lexicales, structures de traits typees, calculs lies aux descriptions lexicales.
The part-of-speech determination is necessary for resolving the part-of-speech ambiguity in English-Korean machine translation. The part-of-speech ambiguity causes high parsing complexity and makes the accurate translation difficult. In order to solve the problem, the resolution of the part-of-speech ambiguity must be performed after the lexical analysis and before the parsing. This paper proposes the CatAmRes model, which resolves the part-of-speech ambiguity, and compares the performance with that of other part-of-speech tagging methods. CatAmRes model determines the part-of-speech using the probability distribution from Bayesian network training and the statistical information, which are based on the Penn Treebank corpus. The proposed CatAmRes model consists of Calculator and POSDeterminer. Calculator calculates the degree of appropriateness of the partof-speech, and POSDeterminer determines the part-of-speech of the word based on the calculated values. In the experiment, we measure the performance using sentences from WSJ, Brown, IBM corpus.
영한 기계번역에서 영어 단어의 품사결정은 번역할 문장에 사용된 어휘의 품사 모호성을 해소하기 위해 필요하다. 어휘의 품사 모호성은 구문 분석을 복잡하게 하고 정확한 번역을 생성하는 것을 어렵게 한다. 본 논문에서는 이러한 문제점을 해결하기 위해 어휘 분석 이후 구문 분석 이전에 품사 모호성을 해소하려 하였으며 품사 모호성을 해소하기 위한 CatAmRes 모델을 제안하고 다른 품사태깅 방법과 성능 비교를 하였다. CatAmRes는 Penn Treebank 말뭉치를 이용하여 Bayesian Network를 학습하여 얻은 확률 분포와 말뭉치에서 나타나는 통계 정보를 이용하여 영어 단어의 품사를 결정을 한다. 본 논문에서 제안한 영어 품사결정 모델 CatAmRes는 결정할 품사의 적정도 값을 계산하는 Calculator와 계산된 적정도 값에 근거하여 품사를 결정하는 POSDeterminer로 구성된다. 실험에서는 CatAmRes의 동작과 성능을 테스트 하기 위해 WSJ, Brown, IBM 영역의 말뭉치에서 추출한 테스트 데이터를 이용하여 품사결정의 정확도를 평가하였다.
How an Embodied Mind Perspective can Influence the Study of Emotion Joshua Ian Davis (jdavis at barnard.edu), Moderator Department of Psychology, Barnard College of Columbia University 3009 Broadway, New York, NY 10027, USA Christian Keysers (c.m.keysers at rug.nl) Department of Neuroscience, University Medical Center Groningen & BCN NeuroImaging Center, University of Groningen A. Deusinglaan 2, 9713AW Groningen, The Netherlands Fritz Strack (strack at psychologie.uni-wuerzburg.de) University of Wurzburg LS Psychologie II, Rontgenring 10, 97070 Wurzburg, Germany Jamil Zaki (jamil at psych.columbia.edu) Department of Psychology, Columbia University 1190 Amsterdam Avenue, New York, NY, 10027, USA Keywords: Embodied Mind, Embodied Cognition, Embodiment, Emotion, Affect, Empathy, Botox, Facial Feedback, Neuroimaging, Psychophysiology Embodied approaches to studying the mind have received increasing attention over the last decade (e.g. Barsalou, 2008; Damasio, 1999; Lakoff & Johnson, 1999; Niedenthal, 2007). Common to these approaches is the goal of understanding how our processes for perception and action form a basis for our higher-level thoughts and emotions. While there has been a great deal of interest in embodied cognition as a theoretical stance, there have only been hints at the value of this approach to researchers across the cognitive sciences. This symposium is intended to help illustrate how an embodied mind perspective can change the research questions we ask and the ways we interpret results when studying emotion. We present four lines of research, bringing to bear psychophysiological, muscle paralysis, neuroimaging, and behavioral methods. We aim to broaden the discussion on the role of an embodied approach in emotion research and the cognitive sciences more generally. Jamil Zaki will begin the symposium by presenting research on the physiological mechanisms underlying empathic accuracy. Joshua Davis will then discuss research on the paralyzing effects of Botox on emotional experience. In the third talk, Christian Keysers will explore the embodied nature of the brain mechanisms by which we understand the emotions of others. Finally, Fritz Strack will discuss the various psychological mechanisms by which bodily actions can influence affect. Zaki: Shared physiological states and interpersonal understanding Empathy research has demonstrated that when perceivers and targets share autonomic arousal, perceivers can more accurately recognize those targets’ emotions (empathic accuracy). In particular, empathic accuracy has been shown when changes in skin conductance response for both perceivers and targets are correlated over time (known as physiological linkage). This finding suggests that embodying the states of others is an important route to interpersonal understanding. We replicated and extended this finding, illustrating how these effects depend on qualities of the target. Targets differ in emotional coherence (correlation between targets’ arousal and targets’ affect ratings), and empathic accuracy is highest when perceivers view targets with greater emotional coherence. Targets whose emotional states are accompanied by congruent bodily states become affectively “readable,” partially through facilitating shared arousal. We will discuss these findings and what they reveal about the mechanisms underlying the role of physiological linkage in empathic accuracy. Jamil Zaki will receive his Ph.D. in Psychology from Columbia University in 2010. His work on empathic accuracy has appeared in Psychological Science (Zaki, Bolger, & Ochsner, 2008). He is the recipient of a 2008 Autism Speaks Pre-Doctoral Award. Davis: The effects of Botox on emotional experience Prior research suggests that facial expressions are more than outward signals of what a person feels, but can also influence a person’s emotional experience. Earlier experiments guided participants to voluntarily pose or inhibit facial expressions. The voluntary nature of these tasks can lead to issues (e.g. participant awareness of the hypothesis, distraction, and various cognitive mediating variables) that can make interpretation of the findings more difficult. We compared participants receiving Botox – to treat facial wrinkles – to those receiving a control treatment. Crucially, Botox injections leave muscles in a state of flaccid paralysis by disrupting neural transmission at the
In this thesis, I present arguments for a model of language acquisition with three characteristics. These are (1) Continuity in the abstract principles of Universal Grammar; (2) Lexical Learning, or the setting of syntactic parameters based upon the acquisition of morphology; and (3) Morpholexical Learning, which is the abstraction of morphological patterns and generalizations from a lexical database. Continuity accounts for what is invariant in language development. Lexical Learning accounts for what is languageparticular, and which therefore must be learned. Morpholexical Learning accounts for the sequence of developmental stages observed in child language data. The main goal of this thesis is to demonstrate that Morpholexical Learning, in conjunction with paradigmatic structure in the lexicon, provides a model for the acquisition of inflectional morphology. I demonstrate this proposal with data on the acquisition of subject-verb agreement morphology in German. In Chapter One, I present an introduction to the concerns and main proposals of this thesis. In Chapter Two, I motivate the existence of paradigmatic structure with three diachronic case studies. In Chapter Three, I return to the acquisitional debates introduced in Chapter One. I argue that the Continuity Hypothesis represents a preferable alternative to the Maturational Hypothesis. Next, I show that Lexical Learning of clausal representations is superior to the Lexical Projection Hypothesis and the Full Competence Hypothesis. I argue that Morpholexical Learning provides an answer to the Developmental Problem which Continuity and Lexical Learning create. In Chapter Four, I provide an extended case study of the acquisition of subjectverb agreement in German. The construction of word-specific paradigms during the stages under examination accounts for the pattern of agreement errors which German children produce. In Chapter Five, I continue the analysis of paradigm mixture begun in Chapter Two. The patterns of paradigm mixture attested Latin, German, and Icelandic are in essence identical to one another, which suggests universal principles of inflectional organization. In Chapter Six, I conclude the thesis with a sketch of how children develop from the "word-specific paradigm" stage to the "general paradigm stage".
Collocations constitute a subclass of multi-word expressions that are particularly problematic for machine translation, due 1) to their omnipresence in texts, and 2) to their morpho-syntactic properties, allowing virtually unlimited variation and leading to long-distance dependencies.Since existing MT systems incorporate mostly local information, these are arguably ill-suited for handling those collocations whose items are not found in close proximity.In this article, we describe an integrated environment in which collocations (and possibly their translation equivalents) are first identified from text corpora and stored in the lexical database of a translation system, then they are employed by this system, which is capable of dealing with syntactic transformations as it is based on a deep linguistic approach.We compare the performance of our system (in terms of collocation translation adequacy) with that of two major MT systems, one statistical, and the other rule-based.Our results confirm that syntactic variation affects translation quality and show that a deep syntactic approach is more robust in this sense, especially for languages with freer word order (e.g., German) and richer morphology (e.g., Italian) than English.
The Columbia Arabic Treebank (CATiB) is a database of syntactic analyses of Arabic sentences. CATiB contrasts with previous ap-proaches to Arabic treebanking in its emphasis on faster production with some constraints on linguistic richness. Two basic ideas inspire the CATiB approach. First, CATiB avoids the annotation of redundant linguistic information that is determinable automatically from syntax and morphological analysis, e.g., nominal case. And secondly, CATiB uses linguistic representation and terminology inspired by the long tradition of Arabic syntactic studies. This makes it easier to train annotators and not be restricted to hire annotators who have degrees in linguistics. This paper describes CATiB’s representation and compares it to other Arabic treebanking efforts. 1.
Computer-aided Acquisition of Semantic Knowledge (CASK) is aimed at describing a number of semantic fields of a few European languages using data mining techniques elaborated within the framework of the new paradigm of computation known as Knowledge Discovery in Databases (KDD). CASK's motivation is to dig deeper in order to find building blocks which could be used in various sophisticated ways. The project is interdisciplinary involving scientific cooperation of experts in linguistics with information engineers. The task of linguists consists in an interactive (computer-aided) discovery of ontology-based definitions of feature structures using the SEMANA (Semantic Analyser) software which was designed especially in order to build linguistic databases with semantic knowledge.
Un assunto sempre più condiviso nell’ambito degli studi sull’acquisizione sia di L1 che di L2 è che l’evidenza empirica privilegiata debba essere rappresentata da corpora di produzioni scritte o orali degli apprendenti, estensivamente annotate a molteplici livelli di rappresentazione linguistica. Più in generale, corpora lemmatizzati e annotati a livello morfosintattico fanno ormai parte dello strumentario comune del linguista. Accanto ad essi, si fa però strada l’esigenza di disporre di risorse testuali più sofisticate dal punto di vista delle modalità di esplorazione linguistica, come ad esempio corpora annotati a livello sintattico (le cosiddette treebank). Questi consentono infatti di osservare i processi di convergenza degli apprendenti verso la lingua “obiettivo” anche a livello di specifici tratti grammaticali astratti o di macro-strutture linguistiche.
In this paper we describe and evaluate a top-down transfer component of a hybrid example-based machine translation system with an architecture similar to that of transfer MT systems, but with automatically derived transfer-rules and dictionary entries based on a parallel treebank. The tests were applied on the translation pair Dutch to English. Evaluation and error analysis have shown that the top-down transfer process has a number of shortcomings on which we wish to report and which we will try to solve in future work by applying bottom-up transfer.
Abstract Recently, there has been a growing interest in regional variation within African American English. This study reviews a work done on local speech in Pittsburgh, Pennsylvania, discussing trends for both African American and White ethnic groups. Just as scholars have found in other geographic regions, in Pittsburgh, African Americans and Whites share a number of feature characteristics of the local dialect, but remain distinct in a number of other ways. Research in Pittsburgh, as elsewhere, highlights the complexity, rather than the homogeneity, of African American speech across the country, as speakers exhibit alignment to both regional and supraregional ethnic linguistic norms.
It is unlikely that Standard Afrikaans has been based on one relatively uniform vernacular. Ana Deumert has convincingly argued that what we recognise as Standard Afrikaans today is a construction to be attributed to language entrepreneurs who strove for a unique South African identity towards the end of the nineteenth and early 20th century. This led to deliberately discarding some of the then metropolitan Dutch linguistic norms. The Afrikaans negative and diminutive systems will be shown to be the linguistic outcome of these conceptions of identity and purity.
We present an approach for smoothing treebank-PCFG lexicons by interpolating treebank lexical parameter estimates with estimates obtained from unannotated data via the Inside-outside algorithm. The PCFG has complex lexical categories, making relative-frequency estimates from a treebank very sparse. This kind of smoothing for complex lexical categories results in improved parsing performance, with a particular advantage in identifying obligatory arguments subcategorized by verbs unseen in the treebank.
Az eladasban a Szeged Treebank fuggsegi fa formatumra torten atalakitasanak folyamatat mutatjuk be. Az eredetileg frazisstrukturalt treebankbl automatikus konverzio eredmenyekeppen letrejott fuggsegi fakat kezi uton ellenriztuk es javitottuk, letrehozva ezzel az els magyar nyelv kezzel annotalt dependenciakorpuszt. Jelenleg az uzleti hireket, ujsaghireket es jogi szovegeket tartalmazo alkorpuszok annotacioja fejezdott be, de terveink kozott szerepel a teljes korpusz atalakitasa fuggsegi fa formatumra. Az elkeszult adatbazis hasznosithato tobbek kozott az informaciokinyeresben es a gepi forditasban is.
This article aims to show the effectiveness of evolutionary algorithms in automatically parsing sentences of real texts. Parsing methods based on complete search techniques are limited by the exponential increase of the size of the search space with the size of the grammar and the length of the sentences to be parsed. Approximated methods, such as evolutionary algorithms, can provide approximate results, adequate to deal with the indeterminism that ambiguity introduces in natural language processing. This work investigates different alternatives to implement an evolutionary bottom-up parser. Different genetic operators have been considered and evaluated. We focus on statistical parsing models to establish preferences among different parses. It is not our aim to propose a new statistical model for parsing but a new algorithm to perform the parsing once the model has been defined. The training data are extracted from syntactically annotated corpora (treebanks) which provide sets of lexical and syntactic tags as well as the grammar in which the parsing is based. We have tested the system with two corpora: Susanne and Penn Treebank, obtaining very encouraging results.
Abstract The aim of the article is to test empirically predictions formulated in the Transitivity Hypothesis framework. Methodological problems of the original approach are discussed and some solutions are offered. For the testing of the hypotheses two corpora of Czech were used (Prague Spoken Corpus and Prague Dependency Treebank). The results question both the predicted impact of the language form on transitivity and, more importantly, the concept of the Transitivity Hypothesis in general.
This paper presents a basic analysis of syntactic annotation errors and inconsistencies in the Prague Dependency Treebank, the biggest corpus of Czech with manual syntactic annotation. The corpus is used for developing and testing of many syntactic analysers of Czech and the problems in the annotation have an essential impact on the evaluation of the quality of these parsers and the results of precision measurements. We identify some of the basic annotation problems and in some cases, we outline possible solutions.
Abstract This chapter discusses the theme of this volume which is about violence in the language of Victorian novels. It analyzes the works of several notable Victorian writers including Charles Dickens, Anne Brontë, George Eliot and Thomas Hardy using narratography. It explains that narratography is the apprehension of mediated narrative increments as traced out in prose or image by the analytic act of reading. This chapter argues that novel violence violates not the literary community but the linguistic norm through their calculated deviance.
The basic concept of semantic Web,ontology and semantic annotation are described.Then the semantic annotation technology and tool today are introduced and analyzed,and a way of automatic semantic annotation based on HTML documents that contain rich semantic data on the Web is presented.This method couples structural analysis of documents with semantic analysis incorporating domain ontologies and lexical database Hownet,discovers the semantic partition tree corresponding to documents,and annotates HTML documents with semantic lables.The experiment is based on the HTML documents of electronic products,the result shows the method is feasible.
This paper presents a simple and effective approach to improve dependency parsing by using subtrees from auto-parsed data. First, we use a baseline parser to parse large-scale unannotated data. Then we extract subtrees from dependency parse trees in the auto-parsed data. Finally, we construct new subtree-based features for parsing algorithms. To demonstrate the effectiveness of our proposed approach, we present the experimental results on the English Penn Treebank and the Chinese Penn Treebank. These results show that our approach significantly outperforms baseline systems. And, it achieves the best accuracy for the Chinese data and an accuracy which is competitive with the best known systems for the English data.
BACKGROUND: Memory impairment and verbal learning are the most common cognitive deficits associated with schizophrenia. Hopkins Verbal Learning Test (HVLT) is considered to be the most reliable test to asses memory and verbal learning in this mental illness. AIMS: to create one form of the HVLT which would suit our linguistic and cultural context and to study the characteristics of this test in a group of healthy subjects. METHODS: The HVLT consists of a list of 12 words belonging to 3 semantic categories and which are read orally to the subject with an immediate and differed recall. The first part of this work was to select words from a lexical database in order to create the list of the HVLT. The test was then administered to 103 subjects aged from 17- to 45-years-old (mean=27,4; SD =7,3) and having between 1 and 20 years of education ( mean=12,2; SD=5,3). RESULTS: No statistical difference was found within performances of the HVLT across gender and sex. Whereas, years of education was found to have an impact on performances. Although statistically difference was found across level of education. CONCLUSION: Our study permitted us to create one form of the HVLT which well suits our Tunisian context and which we could use to evaluate memory functions among people suffering from schizophrenia.
Impersonating Identity in Spaces of Difference is an ALTERnative discourse of dis/ruption, decolonization, deconstruction....Writing on the b/orders of theories, disciplines, genres, cultures....this re/search weaves together personal, familial, and societal stories of silence and silencings. By subverting conventional academic texts and hegemonic frameworks of Canadian "MultiCULTural institutions and Canadian "MultiCULTural society," this subaltern re/search claims a space for marginalized voices. Conventional ways of knowing, being, becoming are generatively disRupted in order to create awareness of the continuing legacies of colonialism, modernity, patriarchy...and to highlight the urgent need to provide genuine spaces of "belonging" that inclusively honour and respect the gifts of all individuals. Entangled in the in/visible hierarchical realities of "Canadianness," this in/quiry articulates decolonizing resistance. The layers of the impure academic textual body perform the multiple fragmented intertextual layers of the improper Indocanadian/ Can-indian body creating em(bodi)ed re-imag(e)inings of epistemology and pedagogy. This tense, restless landscape of multi-languages, multi-genres, multimeanings, multi-truths is an inter-Ruption of predetermined b/orders and predetermined bodies of predetermined purity, in a world of ever-changing multiplicity. Drawing on "difference" in multiCULTuralism, language, voice, and identity, this performative work travels in and out of questions of absence, "hybridity," foreignness, loss, displacement, marginalization, patriarchy, colonialism, modernity....In this powerful and liberating form of non-traditional in/quiry, meaning making takes precedence over conventional stylistic or pre-established structures and acknowledges personal "ethnic" experience as a valuable form of reliable "academic" knowledge. This trans-disciplinary, transformative, transcultural...in/quiry disRupts traditional hegemonic narratives and challenges the conventional notions of re/search and writing through its form and content. Through the braided weaving of English, French and Punjabi, personal stories, familial narratives, prose, letters, e-mails, cross cultural conversations, visual imagery, historical documents, "subtexts," "surtexts," intertexts, collaborative texts...undermine and upset hegemonic linguistic norms. Fiction, fantasy, History, herstory, theirstories, memory...are juxtaposed through mixed non-linear genres and codes in protest of violent acts of com(form)ity, exclusion and censorship. Stories of India and Canada find themselves interwoven unexpectedly, betraying the lies and the truths of patriarchy, colonialism, modernity, multiCULTuralism, transCULTuralisms...disrupting clean, linear readings of writing and of research. This experiment with nonstructure and typography attempts to actively decolonize, deconstruct/re-construct imposed academic and social identities in a more meaningful way and provide readers with a sense of living in the "transculturality" of the "diasporic in-between." Only by provoking a critical, cross-cultural "INNERstanding" of History, POWer, systemic marginalization, colonialism and modernity can we begin to relationally take on the collective response-abilty for social justice in policy and practice ~ in the word and in the wor(l)d.
Syntactic parsing is a central problem and a challenge in the field of natural language processing. It attracts many studies and consequently there exists the effective parsers for several popular languages such as English and Chinese. For Vietnamese parsing, there have been a few studies focusing on this problem, these studies lack of applying modern techniques, and no popular parser has been released. This paper presents the first study on developing a Vietnamese wide coverage parser based on lexicalized probabilistic context free grammar (LPCFG) and using a standard parsed corpus (similar to Penn Treebank). In this paper the Bikel's parser is modified to analyze Vietnamese. We also provide a comparison based on investigating different parsing models and different linguistic features. The best configuration achieves around 78% of F-score. © 2009 IEEE. Index Keywords: Central problems; Experimental studies; F-score; Linguistic features; NAtural language processing; Parsed corpora; Probabilistic context free grammars; Syntactic parsing; Treebanks; Computational linguistics; Context free grammars; Knowledge engineering; Natural language processing systems; Query languages; Standardization; Systems engineering; Formal languages
The process of marking up the syntactic structure of the sentences in a corpus is facilitated by having a graphical tree editor. A type specification constraints the nature of acceptable trees, and visual indication of non-conformities makes corrections easy. Projectivity checking can also be a useful way to notice errors, depending on the linguistic design of the tree type specification. A small test corpus has been marked up using this method. Only the English language has been investigated so far, and the tool takes no other input than the type specification and the corpus. Adding a lexical database is an obvious way to make it more generally useful.
In the Computational Linguistics community, much work is put into the creation of large, high-quality linguistic resources, often with complex annotation. In order to make these resources accessible to non-technical audiences, formalisms for searching and filtering are needed, like the TIGER corpus query language. Recently, augmented treebanks have been published, including the SALSA corpus which features frame semantic annotation on top of syntactic structure. We design an extension for the TIGER language which allows searching for frame structures along with syntactic annotation. To achieve this, the TIGER ob ject model is expanded to include frame semantics, while remaining fully backwards-compatible and add these extensions to our own implementation of TIGER.
Business English is characterized by a specialized vocabulary,polysemy,stylistic norms of formal,nicety,preciseness,concision and emerging new words.It can improve the study effectiveness to paying attention to the accumulation of professional knowledge,to understand vocabulary through context clues,and to learn new words by chunk approach and concern about the latest business information.
espanolComo es bien sabido, aunque para los hablantes de una lengua las variedades dialectales resulten mas evidentes en los planos lexico, fonetico o fonologico, ellas se advierten en todos los niveles del lenguaje, orbita de la que, por supuesto, no escapa la sintaxis. Asi, en el caso particular del espanol de Buenos Aires, el uso del Preterito Perfecto Compuesto del Modo Indicativo difiere sensiblemente de la norma castellana, a la vez que la conciencia de los hablantes de la lengua respecto de el es practicamente nula: o lo niegan por completo, alegando que prefieren siempre el Preterito Perfecto Simple, o bien aducen que lo emplean segun la norma de Madrid; lo cual, como se vera a lo largo de nuestro trabajo, no resulta de ese modo en ninguno de los dos casos. Asi pues, intentaremos problematizar las cuestiones de norma y uso, en relacion con la conciencia de los hablantes portenos respecto de su empleo de los tiempos pasados. Para ello, partiremos de un trabajo de campo que hemos realizado y que nos ha permitido esbozar algunos matices caracteristicos del uso del tiempo verbal que nos ocupa, es decir, el Preterito Perfecto Compuesto del Modo Indicativo del dialecto rioplatense. EnglishIt is well known that dialectal language variations appear at every level of language including syntax. However, speakers are usually aware of lexical, phonetics, and phonological variations only. In this particular case, as expected, the use of perfect tenses in Buenos Aires (Argentina) is very different from that of Madrid (Spain). The problem is that most Argentinean speakers know how to use the Present Perfect according to Spanish rules they have learned in school, but their speech do not matches their learning. Most Argentinean speakers would say (and they believe) that they do not use the Present Perfect in everyday life, when they actually do, albeit in a different way. That is why I conducted a survey among speakers of all kind of age, in order to distinguish some specific characteristics of the Present Perfect use in rioplatense dialect. Finally, I intend to discuss the concept of language norm and use related to speakers' awareness in Buenos Aires.
The thesis studies the translation process for the laws of Finland as they are translated from Finnish into Swedish. The focus is on revision practices, norms and workplace procedures. The translation process studied covers three institutions and four revisions. In three separate studies the translation process is analyzed from the perspective of the translations, the institutions and the actors. The general theoretical framework is Descriptive Translation Studies. For the analysis of revisions made in versions of the Swedish translation of Finnish laws, a model is developed covering five grammatical categories (textual revisions, syntactic revisions, lexical revisions, morphological revisions and content revisions) and four norms (legal adequacy, correct translation, correct language and readability). A separate questionnaire-based study was carried out with translators and revisers at the three institutions. \n\nThe results show that the number of revisions does not decrease during the translation process, and no division of labour can be seen at the different stages. This is somewhat surprising if the revision process is regarded as one of quality control. Instead, all revisers make revisions on every level of the text. Further, the revisions do not necessarily imply errors in the translations but are often the result of revisers following different norms for legal translation. \n\nThe informal structure of the institutions and its impact on communication, visibility and workplace practices was studied from the perspective of organization theory. The results show weaknesses in the communicative situation, which affect the co-operation both between institutions and individuals. Individual attitudes towards norms and their relative authority also vary, in the sense that revisers largely prioritize legal adequacy whereas translators give linguistic norms a higher value. Further, multi-professional teamwork in the institutions studied shows a kind of teamwork based on individuals and institutions doing specific tasks with only little contact with others. This shows that the established definitions of teamwork, with people co-working in close contact with each other, cannot directly be applied to the workplace procedures in the translation process studied. Three new concepts are introduced: flerstegsrevidering (multi-stage revision), revideringskedja (revision chain) and normsyn (norm attitude). \n\nThe study seeks to make a contribution to our knowledge of legal translation, translation processes, institutional translation, revision practices and translation norms for legal translation. \n\n Keywords: legal translation, translation of laws, institutional translation, revision, revision practices, norms, teamwork, organizational informal structure, translation process, translation sociology, multilingual.
A text is the carrier of language,and the foundation of cognizing and distinguishing its styles and functions.Text typology offers TQA objective and theoretical underpinnings.Classification and cognition of text types is a fundamental and cognitive approach to text types and has the original guiding effects on TQA.Different text types are meant not only to make up different writing forms including lexical features,styles and norms of writing,rhetorical devices,etc,but also to form different language functions,different text focuses,different TT purposes and translation methods.Obviously,these differences require us to establish different assessment criteria and principles,which provide us with TQA references,and lend themselves to further explore and study TQA model so as to make it maneuverable and practical.
Portuguese, the national language of Portugal and Brazil, belongs to the Romance language group. It is descended from the Vulgar Latin of the western Iberian Peninsula (the regions of Gallaecia and Lusitania of the Roman Empire), as is Galician, often wrongly considered a dialect of Spanish. Portugal originated as a county of the Kingdom of Galicia, the westernmost area of theChristian north of the peninsula, the south having been under Arabic rule since the eighth century. Its name derived from the towns of Porto (Oporto) and Gaia (< CALE) at the mouth of the Douro river. As Galicia was definitively incorporated into the Kingdom of Castile and León, Portugal achieved independence under the Burgundian nobility to whom the county was granted in the eleventh century. Alfonso Henriques, victor of the battle of Sa-o Mamede (1128), was the first to take the title of King of Portugal. Apart from a short period of Castilian rule (1580-1640), Portugal was to remain an independent state. The speed of the Portuguese reconquest of the Arabic areas played an important partin the development of the language. The centre of the kingdom was already in Christian hands, after the fall of Coimbra (1064), and many previously depopulated areas had been repopulated by settlers from the north. The capture of Lisbon in 1147 and Faro in 1249 completed the Portuguese Reconquest, nearly 250 years before its Spanish counterpart, bringing northern and central settlers into the Mozarabic (arabised Romance) areas. The political centre of the kingdom also moved south, Guimarães being supplanted first by Coimbra, and subsequently by Lisbon as capital and seat of the court. The establishment of the university in Lisbon and Coimbra in 1288, to move between the two cities until its eventual establishment in Coimbra in 1537, made the centre and south the intellectual centre (although Braga in the north remained the religious capital). The form of Portuguese which eventually emerged as standard was the result of the interaction of northern and southern varieties, which gives Portuguese dialects their relative homogeneity. For several centuries after the independence of Portugal, the divergence of Portugueseand Galician was slight enough for them to be considered variants of the sameCastilian as a lyric poetry until the middle of the fourteenth century. Portuguese first appears as the language of legal documents at the beginning of the thirteenth century, coexisting with Latin throughout that century and finally replacing it during the reign of D. Dinis (1279-1325). In the fifteenth and sixteenth centuries the spread of the Portuguese Empire establishedPortuguese as the language of colonies in Africa, India and South America. A Portuguesebased pidgin was widely used as a reconnaissance language for explorers and later as a lingua franca for slaves shipped from Africa to America and the Caribbean. Some Portuguese lexical items, e.g. pikinini ‘child’ (pequeninho, diminutive of pequeno ‘small’), save ‘know’ (saber), are common to almost all creoles. Caribbean creoles have a larger Portuguese element, whose origin is controversial – the Spanish-based Papiamentu of Curaçao is the only clear case of large-scale relexification of an originally Portuguesebased creole. Brazilian Portuguese (BP), phonologically conservative, and lexically affected by the indigenous Tupi languages and the African languages of the slave population, was clearly distinct from European Portuguese (EP) by the eighteenth century. Continued emigration from Portugal perpetuated the European norm beside Brazilian Portuguese, especially in Rio de Janeiro, where D. João and his court took refuge in 1808. After Brazil gained its independence in 1822, there was great pressure from literary and political circles to establish independent Brazilian norms, in the face of a conservative prescriptive grammatical tradition based on European Portuguese. With approaching 200 million speakers in the eight member states of the Comunidadedos Países de Língua Portuguesa (CPLP), Portuguese is reckoned to be the sixth most widely spoken language in the world. It is spoken by 10 million people in Portugal and over 180 million in Brazil (following estimates based on the 2000 census figure of 170 million), and is the official language of Angola, Mozambique, Guiné-Bissau, São ToméPríncipe, Cape Verde and East Timor. It is spoken in isolated pockets in Goa, Malacca and Macau, and in expatriate communities in Europe and North America. Portuguese-based creoles are widely found in W. Africa and the Caribbean; Cape Verdean creole notably has official status beside Portuguese. The standard form of European Portuguese is traditionally defined as the speech ofLisbon and Coimbra. The distinctive traits of Lisbon phonology (centralisation of /e/ to /ɐ/ in palatal contexts; uvular /R/ in place of alveolar /r/) have more recently become dominant as a result of diffusion by the mass media. Unless otherwise stated, all phonetic citation forms are of European Portuguese. Of the two main urban accents of Brazilian Portuguese, Carioca (Rio de Janeiro) showsa greater approximation towards European norms than Paulista (São Paulo). While the extreme north and south show considerable conservatism, regional differences in Brazilian Portuguese are still less marked than class-based differences; non-standard varieties and informal speech show considerable simplification of inflectional morphology and concord, which has invited comparison with creoles.
All too often work in computational linguistics on the acquisition of conceptual descriptions takes place in isolation from work on concepts in psychology and neural science. We feel this is a mistake as evidence from these related disciplines can provide us with better ways of evaluating our results. In the talk I will present work in CIMEC on using cognitive evidence to evaluate the results of lexical acquisition work - specifically, using feature norms to evaluate the acquisition of features, and using EEG data to evaluate the results of categorization experiments.
Background. The aim of the present study was to investigate whether word lexical and semantic properties may differently affect the timing and topographical distribution of ERP components. In particular, we focused on the neural processing of abstract and concrete words. Previous studies have provided controversial evidence about the way words having a different concreteness degree are represented in the brain. The reason of that may be traced, at least in part, in methodological heterogeneity and poor control of stimuli. Our efforts were directed to overcome such limitations. Methods. A group of 15 volunteers was engaged in a lexical decision task (word/non-word discrimination). 600 linguistic stimuli were created. They consisted of 300 words (150 abstract, 150 concrete) and 300 legal pseudo-words. All stimuli were balanced in terms of length. Abstract and concrete words were also balanced in terms of frequency of occurrence (Bertinetto, 2006) and familiarity, while they significantly differed in terms of concreteness and imageability (ratings were obtained from three independent groups of 30 subjects with 5-point scales). EEG was recorded from 128 scalp sites at a sampling rate of 512 Hz. ERPs were time-locked to stimulus onset. Low Resolution Electromagnetic Tomography (LORETA) was performed on ERP difference waves. Results. RTs to words were faster than RTs to pseudo-words (the so-called “word superiority effect”). Words were discriminated from pseudo-words since 300 ms post-stimulus with larger N2 responses to words than to pseudo-words over the left occipito-temporal areas, namely the BA37, possibly indexing the activity of the so-called VWFA. Similarly, at later processing stages (between 500-600 ms), corresponding to deeper linguistic processing and over the left temporal parietal electrode sites, ERP responses to words were larger than to pseudo-words. RTs to concrete words were slightly faster than to abstract words. The two categories were discriminated as early as 350 ms post-stimulus, with larger responses to concrete words than to abstract words over the medial occipital regions. Concreteness-related ERP differences were also observed in the amplitudes of the later anterior LP component (between 370-570 ms), with larger responses to abstract words than to concrete words. Conclusions. Overall, these data show how ERPs can dissociate between lexical and higher level processes. Our results indicate that semantic processing may take place near-simultaneously and in different brain regions with the processing of information about the form of a word and its lexical properties. The concreteness effect seems not to be strictly bound neither to some confounding linguistic properties of the stimuli, as the word frequency of occurrence or familiarity, nor to task demands. These data provide support for the hypothesis that concrete, imaginable concepts activate perceptually based representations not available to abstract concepts.
Word Sense Disambiguation is the most critical issue in natural language processing. Although it has been addressed by many researchers, no satisfactory results are reported. Rule based systems alone can not handle this issue due to ambiguous nature of the natural language. Knowledge-based systems are therefore essential to find the intended sense of a word form. Machine readable dictionaries have been widely used in word sense disambiguation. The problem with this approach is that the dictionary entries for the target words are very short. WordNet is the most developed and widely used lexical database for English. The entries are always updated and many tools are available to access the database on all sorts of platforms. The WordNet database can be converted in MySQL format and we have modified it as per our requirement. Sense's definitions of the specific word, "Synset" definitions, the "Hypernymy" relation, and definitions of the context features (words in the same sentence) are retrieved from the WordNet database and used as an input of our Disambiguation algorithm.
Calabrese, Relevant southern features here are NC > NN (monno vs mondo < MUNDUM ‘world’, piommo vs piombo < PLUMBUM ‘lead’); characteristic patterns of both tonic and atonic vowel development; use of postposed possessives (figliomo vs mio figlio ‘my son’); extensive use of the preterit; etc. A number of features mark off Tuscan from its neighbours: absence of metaphony (umlaut); -VriV-> -ViV-(IANUARIUM > gennaio, cf. Gennaro, patron saint of Naples); fricativisation of intervocalic voiceless stops – the socalled gorgia toscana ‘Tuscan throat’ – which yields pronunciations such as [la harta] la carta ‘the paper’, [kauo] capo ‘head’, [lo hiro] lo tiro ‘I pull it’; etc. Such divisions reflect both geographical and administrative boundaries. The La Spezia-Rimini line corresponds very closely both to the Apennine mountains and to the southern limit of the Archbishopric of Milan. The line between central and southern dialects approximates to the boundary between the Lombard Kingdom of Italy and the Norman Kingdom of Sicily, and to a point where the Apennines broaden out to form a kind of mountain barrier between the two parts of the peninsula. The earliest texts are similarly regional in nature. The first in which undisputed vernacular material occurs is the Placito Capuano of 960, a Latin document reporting the legal proceedings relating to the ownership of a piece of land, in the middle of which an oath sworn by the witnesses is recorded verbatim: sao ko kelle terre, per kelle fini que ki contene, trenta anni le possette parte Sancti Benedicti ‘I know that those lands, within those boundaries which are here stated, thirty years the party of Saint Benedict owned them.’ The textual evidence gradually increases, and by the thirteenth century it is clear that there are wellrooted literary traditions in a number of centres up and down the land. These are touched on briefly by the Florentine Dante (1265-1321) in a celebrated section of this treatise De Vulgari Eloquentia, but it is the poetic supremacy of his Divine Comedy, rapidly followed in the same city by the achievements of Petrarch (1304-74) and Boccaccio (1313-75), which ensured that literary, and thus linguistic, pre-eminence should go to Tuscan. There ensued a centuries-long debate about the language of literature – la questionedella lingua ‘the language question’, with Tuscan being kept in the forefront as a result of the theoretical writings of the influential Venetian (!) Pietro Bembo (1470-1547), especially his Prose della volgar lingua (1525). His ideas were adopted by the members of the Accademia della Crusca, founded in Florence in 1582-3, which produced its first dictionary in 1612 and which still survives as a centre for research into the Italian language. Meanwhile, although the affairs of day-to-day existence were largely conducted in dialect, the sociopolitical dimension of the question increased in importance in the eighteenth and nineteenth centuries, assuming a particular urgency after unification in 1861. The new government appointed the author Alessandro Manzoni (1785-1873) – himself born in Milan but yet another enthusiastic non-native advocate of Florentine usage – to head a commission, which in due course recommended Florentine as the linguistic standard to be adopted in the new national school system. This suggestion was not without its critics, notably the great Italian comparative philologist, Graziadio Ascoli (1829-1907), and a number of the specific recommendations were hopelessly impractical, but in any case the core of literary usage was so thoroughly Tuscan that the language taught in schools was bound to be similar. Education was, of course, crucial since the history of standardisation is essentially the history of increased literacy. On the most conservative estimate only 2.5 per cent of the population would have been literate in any meaningful sense of the word in 1861, although a more recent and moreper had 91.5 per cent by 1961, the centenary of unification and the thousandth anniversary of the first text. Even so, there is no guarantee that those who can use Italian do so as their normal daily means of communication, and it was only in 1982 that opinion polls recorded a figure of more than 50 per cent of those interviewed claiming that their first language was the standard rather than a dialect. Yet the opposition language/dialect greatly oversimplifies matters. For most speakers it is a question of ranging themselves at some point of a continuum from standard Italian through regional Italian and regional dialect to the local dialect, as circumstances and other participants seem to warrant. Note too that the term dialect means something rather different when used of the more or less homogeneous means of spoken communication in an isolated rural community and when used to refer to something such as Milanese or Venetian, both of which have fully fledged literary and administrative traditions of their own, and hence a good deal of internal social stratification. Another significant factor in promoting a national language was conscription, firstbecause it brought together people from different regions, and second because the army is statutorily required to provide education equivalent to three years of primary school to anyone who enters the service illiterate. Indeed, it is out of the analysis of letters written by soldiers in the First World War that some scholars have been led to recognise italiano popolare ‘popular Italian’ as a kind of national substandard, a language which is neither the literary norm nor yet a dialect tied to a particular town or region. Among the features which characterise it are: the extension of gli ‘to him’ to replace le ‘to her’ and loro ‘to them’, and, relatedly, of suo ‘his/her’ to include ‘their’; a reduction in the use of the subjunctive in complement clauses, where it is replaced by the indicative, and in conditional apodoses, where the imperfect subjunctive is replaced by the conditional, and the pluperfect subjunctive is replaced by either the conditional perfect or the imperfect indicative (thus standard se fosse venuto, mi avrebbe aiutato (‘if he had come he would have helped me’) becomes either se sarebbe venuto, mi avrebbe aiutato or se veniva, mi aiutava, the latter having an imperfect indicative in the protasis too; the use of che ‘that’ as a general marker of subordination; plural instead of singular verbs after nouns like la gente ‘people’. Some of these uses – e.g. gli for loro, the reduction in the use of the subjunctive and the use of the imperfect in irrealis conditionals – have also begun to penetrate upwards into educated colloquial usage, and it is likely that the media, another powerful force for linguistic unification, will spread other emergent patterns in due course. Industrialisation, too, has had its effect in redrawing the linguistic boundaries, both social and geographical. In addition to the standard language, the dialects and the claimed existence of ita-liano popolare, there are no less than eleven other languages spoken within the peninsula and having, according to one recent but probably rather high estimate, a total of nearly 2.75 million speakers. Of these, more than two million represent speakers of other Romance languages: Catalan, French, Friulian, Ladin, Occitan and Sardinian. The remaining languages are: Albanian, German, Greek, Serbo-Croat and Slovene. Amidst this heterogeneity, the Italian national and regional constitutions recognise the rights of four linguistic minorities: French speakers in the autonomous region of the Valle d’Aosta (approx. 75,000), German speakers in the province of Bolzano (approx. 225,000), Slovene speakers in the provinces of Trieste and Gorizia (approx. 100,000), Ladin speakers in the province of Bolzano (approx. 30,000). Yet French (and Occitan – approx. 200,000) and German speakers outside the stated areas are not protected in the same way. Norof closely related the two in turn being sub-branches of the Rhaeto-Romance group. The recognised linguistic minorities are, not surprisingly, in areas where the borders of the Italian state(s) have oscillated historically. In contrast, the southern part of the peninsula is peppered with individual villages which preserve linguistically the traces of that region’s turbulent past. It is here that we find Italy’s 100,000 Albanian, 20,000 Greek and 3,500 SerboCroat speakers, as well as a number of communities whose northern dialects reflect the presence of mediaeval settlers and mercenaries. Sardinia too contains a few Ligurian-speaking villages and 20,000 Catalan speakersin the port of Alghero as evidence of former colonisation. More importantly, the island has almost 1,000,000 speakers of Sardinian, a separate Romance language which has suffered undue neglect ever since Dante said of the inhabitants that they imitated Latin tanquam simie homines ‘as monkeys do men’. What he was referring to was the way in which Sardinian, both in structure and vocabulary, reveals itself to be the most conservative of the Romance vernaculars. Thus, we find a vowel system with no mergers apart from the loss of Latin phonemic vowel length; an absence of palatalisation of k and g; preservation of final s (with important morphological consequences); a definite article su, sa, etc. which derives from Latin IPSE rather than ILLE. Old Sardinian also maintained direct reflexes of the Latin pluperfect indicative and imperfect subjunctive. and the language is one of the few not to retain a future periphrasis from Latin infinitive + HABEO, using instead a reflex of Latin DEBERE ‘to have to’, e.g. des essere ‘you will be’. On the lexical side we have petere ‘to ask’, imbennere ‘to find’ (cf. Lat. INVENIRE), domo/domu ‘house’, albu ‘white’, etc. (contrast It. chiedere, trovare, casa, bianco). The presence of Italian outside the boundaries of the modern Italian state is due totwo rather different types of circumstance. First, it may be spoken in areas either geographically continuous with or at some time part of Italy, as in the independent Republic of San Marino (population 30,000), enclosed within the region of Emilia-Romagna, and in Canton Ticino (population approx. 325,000), the entirely italophone part of Switzerland. Both have local dialects, Romagnolo in San Marino and Lombard in Ticino, as well as the standard language of education and administration. Elsewhere, the historical continuity is reflected at the level of dialect, but with the superimposition of a different standard language. Thus, in Corsica (population approx. 280,000) the dialects are either Tuscan (following partial colonisation from Pisa in the eleventh century) or Sardinian in type, but the official language has since 1769 been French. The same situation obtains for those Italian dialects spoken in the areas of Istria and Dalmatia now part of Slovenia and Croatia. The second circumstance arises when Italian, or more often Italian dialects, has beencarried overseas, mainly to the New World. In the USA about one million Italian speakers constitute the second largest linguistic minority (after Hispano-Americans). They are concentrated for the most part either in New York, where they are mainly of southern origin and where a kind of southern Italian dialectal koine has emerged, and in the San Francisco Bay area, where northern and central Italians predominate, and where the peninsular standard has had more influence. Italian language media include a number of newspapers, radio stations and television programmes. The current signs of a reawakening of interest in their linguistic heritage amongst Italo-Americans are paralleled in Canada and Australia, each with about half a million Italian speakers according to official figures. There were also in excess of three million émigrés to South America, mostly to Argentina, and this has led, on the River Plate, to the developmentItalian in the Australia had its origins in the language of an underprivileged and often uneducated immigrant class, in Africa – specifically Ethiopia and Somalia and until recently Libya – Italian survives as a typical relic of a colonial situation. Ethiopia also has the only documented instance of an Italian-based pidgin, used not only between Europeans and local inhabitants but also between speakers of mutually unintelligible indigenous languages. The position of Italian in Malta is similarly due to penetration at a higher rather than a lower social level. Research is only now beginning into the linguistic consequences of the postwar migration of, again mainly southern, Italian labour as ‘Gastarbeiter’ in Switzerland and Germany. Finally, two curiosities are the discovery by a group of Italian ethnomusicologists in 1973 in the village of S˘tivor in northern Bosnia of a community of 470 speakers of a dialect from the northern Italian province of Trento, and the case of a group of émigrés from two coastal villages near Bari in Puglia, who settled in Kerch in the Crimea in the 1860s and whose dialectophone descendants died out only in the late twentieth century.
SEER, Vol.87,M. 3, July2009 Reviews Marder,Stephen. A Supplementary Russian-English Dictionary (ASRED2). Second edition. SlavicaPublishers, Bloomington, IN, 2007.xxv+ 736pp. $44.95. The first editionofthisdictionary came out in 1992.A corrected reprint of 1994was followed by a Moscow mirror editionin 1995.For some readers, oftennativeRussianspecialists, thedictionary was something ofa shock,a shockforrecovery from whichtheMoscowmirror edition ismorethanample evidence.Peoplelikethepresent reviewer, somewhat takenaback(something regrettably reflected in his reviewof 1994: SEER, 72, 1, pp. 161-62),but nonetheless massively impressed, wenton to use the dictionary morethan regularly overthenext, well,fifteen years.His first-edition copymaywellstill be inpretty goodcondition, butthere canbe no doubtthatthissecondedition will,whilethefirst onewillforall sorts ofreasonscontinue tobe used,replace it. Time has passedand,within thattime,numerous dictionaries ofRussian have appeared,ofall sorts- in thefirst place,ASRED 1 anticipated them, openedmanyeyesinthesociety whoselanguageitwasrevealing, andinspired them.AndASRED 2, quitedifferent from them, providing invaluable, lexical and grammatical information and employing toolsto facilitate easyuse and accessibility, reappears, considerably expanded,toteachthemand takethem further. The first edition as producedbySlavicawasextremely durable;thissecond looksindestructible. It is a hardback, beautifully producedby thepublisher and,particularly, bytheauthor(though theauthorofa dictionary has to be a 'compiler', itseemsthatinthiscase we really aremuchclosertohavingan 'author').There are otherthings to do in life,alas, and theymustrender it put-downable, but it is mostcertainly all but unput-downable. And the wonderful listof such adjectivesin Russian on p. 79, under vnusabel'nyj, reinforces sucha conviction. Thereis a setofpreliminary pages:after theContents (no page number) comesthe Introduction to the Second Edition(ASRED 2), pp. i-iii,where theauthor buildsonASRED 1and further justifies it.Therefollows theIntroductionto theFirstEdition,pp. v-x, giving invaluableinformation on how besttousethedictionary and reminding us thatthedictionary's original guidingprinciple was 'to fillan alarming - and exasperating gap between whathasbeenrecorded andwhatitispossibletorecord', something inwhich it succeededmostnotably.On pp. xi-xxwe have a SelectedBibliography (Updated),withRussiansources(pp. xi-xviii)and Englishsources(pp. xviiixx ).We have thenAcknowledgments (SecondEdition), p. xxi,Acknowledgments (First Edition), p. xxii,ListofAbbreviations and Conventional Symbols (Updated),pp. xxiii-xxiv, and a RussianAlphabet, p. xxv.The bodyofthe dictionary is in thefollowing 736pages. Importantly, thereareextremely clearand fullentries: theheadwordsand other wordsandphrases within entries areall inboldcharacters and stressed; REVIEWS 527 morphological juncturebetweenstemand endingis marked;thereis exhaustivecross -referring to relatedentries; and an enormousamountofcultural information ofall sorts is given.One can open anypage to giveexamplesof thelast:autizm andAFE on p. 23,pénsija pò [vózrast]u on p. 83,kosój on p. 263, mitëk, mitrofánuska, andMít'kaonp. 323,ocepjátka onp. 406,Pjatèrocka onp. 506, slivát' on p. 579,urjük on p. 665, utjug on p. 669, and cetvërka on p. 706,each page opened withoutsearchingand additionalexamplesbeing citableon almost eachpage.Manyentries havesub-entries; so,forexample,ustrójstuo has eighty-four, spreadoverpp. 666-68, and úxohas twenty-nine, spreadover pp. 669-70 (and witha finequotationto illustrate otkúda rastút usi).The translations are also extremely apt,withfullexplanationand expansionas necessary. One couldgo on and on; itwouldseemto thisreviewer thattogo on and on hereisnotat all appropriate ornecessary - ASRED 2 is a labouroflove, an astonishing treasure troveoflinguistic, literary, political,cultural, social and scientific information aboutRussiaand theRussianlanguage.ASRED 1 was alreadyquitea phenomenon; ASRED 2 notonlyconfirms thatphenomenon but demonstrates, fifteen yearslater,thatit remainssupremely useful and needed,and a realjoy through whichtoflit and inwhichto dwell. StAndrews University Ian Press Sovik, MargretheB. Support, Resistance and Pragmatism: An Examination of Motivation inLanguage Policy inKharkiv, Ukraine. ActaUniversitatis Stockholmiensis, Stockholm SlavicStudies,34. Department ofSlavicLanguages andLiteratures, Stockholm University, Stockholm, 2007.356pp. Figures. Tables.Illustrations. Bibliographical references. Appendices. SEK 343.00 (paperback). The 'languagequestion'in Ukraine- specifically, the role and statusof Ukrainian and Russian- isone ofthoseheatedand controversial issuesthat repeatedly inserts itself intodomesticpoliticaldebates,especially, it seems, during electoral campaigns, whenpresidential candidates and political parties competing forvotesfindit to be a ratherusefultool formobilizing their supporters. Not infrequently, italso servesas a sourceofcontention between Kyivand Moscow.Much oftheproblemresidesin thefactthatin Ukraine languagepreference and usage largelyoverlapwiththe country's regional structure and conflicting politicalvalues,thereby setting the stageforthe 'languagequestion'to becomehighly politicized. The book underreview, whichwas written as a doctoraldissertation at Stockholm University, addressestheproblemin a veryspecific and, indeed, unorthodoxmanner.Instead of examiningpolicies and legal norms or analysing statistical data suchas thelanguageofinstruction in schoolsand universities or the resultsof country-wide sociologicalsurveys, whichhas beenthenormin thescholarly literature, theauthorfocuses on howthepredominantly Russian-speaking residents ofKharkiv - Ukraine's secondlargest...
Résumé Dans cette étude, qui se base sur un petit corpus de textes français et suédois originaux et traduits, l’usage des formes lexicales et pronominales en fonction anaphorique est examiné. Comme prévu, les formes pronominales s’avèrent plus fréquentes dans les textes français que dans les textes suédois, où la répétition du SN lexical thématique prédomine. Cette différence est en général liée aux normes rhétoriques ou stylistiques des cultures respectives, telles que la plus haute fréquence de marqueurs explicites de cohérence textuelle et le besoin plus fort de variation lexicale dans les langues romanes comparées aux langues germaniques, mais l’absence systématique de pronoms anaphoriques dans les textes suédois demande une autre explication. Elle semble due à une répugnance générale des pronoms personnels suédois inanimés ( den, det ) d’assumer une fonction anaphorique.
This paper reports a sociolinguistic study of the state of Greek language in Australia as spoken by native-speaking Greek immigrants and their children. Emphasis is given to the analysis of the linguistic behaviour of these Greek Australians which are attributed to contact with English and to other environmental, social and linguistic influences. The paper discusses the non-standard phenomena in various types of inter-lingual transferences in terms of their incidence and causes and, in correlation with social, linguistic and psychological factors in order to determine the extent of language assimilation, attrition, and the content and context and medium of the language-event. The paper also discusses the transferences from English to Greek and vice- versa from a qualitative and quantitative perspective, of the phonemic, lexical, morphological, syntactic, semantic, pragmatic and prosodic deviations. During the last 170 years of settlement, Greek Australians know and use a new communicative norm with some degree of stability, the Ethnolect, (a non-standard variety of language used by an ethnic group in a static or dynamic bilingual situation) which serves their linguistic needs.