Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The paper studies Hausa film language through the analysis of three communication strategies, namely proverbs, imperatives and forms of address. It shows that Hausa film creates a new discourse by reflecting modern and traditional Hausa society. The films preserve some accepted cultural norms of behavior and norms of communication in order to please the more conservative public. On the other hand, combination of traditional and modern Hausa lifestyle evokes changes in the discourse. The paper shows that proverbs are commonly used as communication strategy for indirectness, rather than a specialized language. It also discovers that imperatives are used as communication strategy in close relations between interlocutors (no matter what their social status is) to express the direct message. As for forms of address, traditional and borrowed terms reflect the changing style of life. The examples extracted from the Hausa films are to show how the regular grammatical and lexical means change their discourse function in new social context.
Abstract The present study explores linguistic predictors and behavioural implications of the orthographic alternation between a spaced (bell tower), hyphenated (bell-tower), and concatenated (belltower) format observed in English compound words. On the basis of two English corpora, we model the evolution of spelling for compounds undergoing lexicalisation, as well as define the set of orthographic, distributional, and semantic properties of the compound's constituents that co-determine the preference for one of the available realisations. We explore iconicity and economy as competing motivations for both the diachronic change and synchronous preferences in spelling. Observed patterns of written production closely mirror the demands and strategies of recognition of compound words in reading. Orthographic choices that go against the reader's economy of effort come with a high recognition cost, as evidenced in inflated lexical decision and naming latencies to concatenated compounds that occur in other spelling formats. Keywords: CompoundingProductionComprehensionSpellingCorpus linguisticsCompeting motivations Acknowledgements Authors would like to thank Valentin Spitkovsky for extensive and insightful discussions of computational aspects of this work, and to Joan Bresnan, Barbara Juhasz and Emmanuel Keuleers for their comments on an earlier draft. Notes 1Throughout this article, we will use the plus sign in the compound spelling (e.g., girl+friend) to refer to the compound regardless of how it is spelled. 2While this procedure meets our goal of describing orthographic alternation, it cannot be used to identify nonalternating concatenated compounds. Unlike spaced and hyphenated compounds, concatenated compounds cannot be told apart from morphologically simple singular common nouns on the basis of part-of-speech tags or orthography in parsed Wikipedia, and the size of the corpus prohibits manual identification of concatenated compounds. To roughly estimate the number of nonalternating concatenated compounds, we extracted all concatenated compounds identified in the morphological coding of the lexical database CELEX (Baayen, Piepenbrock, & Gulikers, Citation1995). We found 500 compounds that were not part of our alternating set in Wikipedia. The sum of the alternating concatenated compounds and the nonalternating ones are 2,100 (=1,600+500) and are a more precise – though potentially still too low – estimate of the total type count of (alternating or nonalternating) concatenated compounds in the Wikipedia corpus. 3The list of alternating compounds comprises: audio+book, call+center, care+giver, chick+pea, coffee+house, copy+cat, cyber+cafe, data+base, die+cast, field+house, fund+raiser, help+line, house+cat, jump+start, race+day, road+show, salary+cap, shoe+box, show+biz, slide+show, soap+box, sound+board, steak+house, paint+ball. 4We are indebted to Emmanuel Keuleers for raising this possibility. 5We also considered entropy as the measure of uncertainty in the choice of one of several alternatives. Entropy is minimal (zero) when a compound occurs in only one of the available formats, and it is maximal when all alternatives are equiprobable (see Milin et al., Citation2009b). Entropy of orthographic choice was not a significant predictor (at the 0.05-level) of either lexical decision or word naming response times.
The importance of attribution is becoming evident due to its relevance in particular for Opinion Analysis and Information Extraction applications. Attribution would allow to identify different perspectives on a given topic or retrieve the statements of a specific source of interest, but also to select more relevant and reliable information. However, the scarce and partial resources available to date to conduct attribution studies have determined that only a portion of attribution structures has been identified and addressed. This paper presents the collection and further annotation of a database of attribution relations from the Penn Discourse TreeBank (PDTB) (Prasad et al., 2008). The aim is to build a large and complete resource that fills a key gap in the field and enables the training and testing of robust attribution extraction systems
Abstract Peace is arguably the problem of the 21st century. Peacefulness is not uniquely human, but a dearth of it among humans disproportionately threatens people and other animals around the globe. The urgent need for peace—if not immediately, everywhere, at any cost, then soon, as a pervasive norm—coincides with unprecedented scholarly attention to peace and to the implications of evolution for psychological functioning in the context of complex sociality. The time is ripe to integrate evolutionary perspectives into peace studies. Toward that end, this chapter describes potential impediments to an evolutionary peace project, provides a basic lexical and conceptual tool kit, and identifies some promising research directions.
The present article looks at a section of Estonian translation history, through the criticism and practice of translation of an extremely prolific and well-known Estonian translator and reviewer of translations Marta Sillaots (1887–1969). Touching upon the issue of the aging of translational texts as well as the norms of translational behavior in the Estonia of the first half of the 20th century, the article makes use of Peeter Torop’s concepts of the explicit poetics of translation (theoretical ideas about translation expressed in paratexts) and implicit poetics of translation (translational choices inside translations) in an attempt to map the translational thought of Marta Sillaots. Firstly, the explicit poetics of translation of Marta Sillaots is analyzed by focusing on the recurring topics and keywords in Sillaots’ translation reviews, published in a monthly literary journal Eesti Kirjandus during the 1920s and 1930s. Secondly, the implicit poetics of Marta Sillaots’ translations is analyzed based on a comparative corpus of two novels translated by Sillaots and their later editions: Romain Rolland’s Jean-Christophe 1936 and 1958 (edited by Henno Rajandi) and Charles Dickens’ David Copperfield 1937 and 1991 (edited by Lia Rajandi). Although the originals are held close at hand, the translations are approached as facts in Estonian cultural space, and thus the focus is on the translation rather than on the original. In comparison with the first edition, the changes made by the editors reveal a pattern of movement towards greater fluency by standardizing the word order, adding linking devices and making the constructions confirm to the rules of target language.The comparison of the explicit and implicit poetics of Sillaots reveals certain conformities as well as inconsistencies. In reviewing translations by other translators, Sillaots traditionally balances between the terms good and bad translation. Firstly, for her, a precondition for a good translation is the connection between the translator and the original that can be expressed by a theoretical interest of the translator in the original author or in a similarity of their literary styles. Secondly, Sillaots emphasizes the need to lead the reader to the translational text by a comprehensive foreword, something that she, as a translator, always tried to do. But most importantly, Sillaots finds it necessary for the translators to educate the Estonian readers, first by the choice of authors to translate and second by the language of translation, which has to be in accordance with the norms of grammar and style of the target language.However, the comparative analysis of her translations to their later editions shows that, although never adding any lexical items, Sillaots often used sentence structure, word order as well as minimal collocation in order to convey what, according to her, was the style of the original author. This practice distanced her translations somewhat from what is regarded as adherence to the rules of Estonian text formation. Sillaots remains an important part of Estonian translation history because of the number of her translations as well as for her poetics of translation.
This document describes a list of Modern Standard Arabic closed-class words, which can be used as a stop list for a variety of natural language processing applications. The list contains 740 inflected words and clitics in the Arabic Treebank (ATB) tokenization scheme (Maamouri et al., 2004; Habash, 2010). The inflected words are based on 309 lemmas from the Standard Arabic Morphological Analyzer, SAMA (Graff et al., 2009). To get a copy of the full list, please contact the authors.
This paper deals with an exploration of the poetry world of E. E. Cummings’ (Edward Estlin Cummings, October 14, 1894- September 3, 1962) from a linguistic perspective. For limitations of time and space, a couple of his representative poems are selected with the purpose of conducting a stylistic analysis of his poetry in the light of discourse analysis with particular reference to lexical and syntactic features that help make Cummings’ style a peculiar example of style as deviation from the norm.
Summary: This paper is devoted to the inter- and intra-linguistic comparison of comics on the basis of the volume Le trésor de Rackham le Rouge (engl. Red Rackham’s Treasure) of the Franco-Belgian Tintin series and its two Catalan versions published in 1964 and 2002, respectively. These versions are analyzed with respect to morpho-syntactic (forms of address, clitic pronoun combinations, clause-initial que and future-oriented temporal adverbial clauses introduced by quan) and lexical criteria (technical vocabulary). The study reveals that the models for the reference variety applied by the translators diverge and that the translations reflect, to a certain extent, the sociolinguistic situation of Catalan at the moment of their creation. [Keywords: Tintin; inter- and intra-linguistic translation; linguistic norm; sociolinguistics of normativization; Joaquim Ventalló]
Cohesion and coherence are inseparable features of any text and discourse. There are several grammatical and lexical devices by which cohesion can be reached. Conjunction is one of them and it is assigned to the intermediate — lexico-grammatical category. Usage of conjunctions and other connecting words for expressing the logic-semantic relations in different languages can vary due to different reasons. One of them is the common norms of a language and certain types of discourse that are not necessarily equally settled among languages. The analysis has been based on the purposeful selection of the scientific (linguistic) English (3articles — 28815 words) and Lithuanian (10 articles — 29421 word) articles. In order to count the average value of the units of the conjunctive relations in the articles, the quantitative measurements have been performed. The investigation of such issues requires performing an interlanguage analysis of these elements, which has shown such general trends: within the English academic discourse additive conjunction has had the largest expansion, and organizational conjunction — the smallest; conjunctive devices have been used less frequently in Lithuanian research articles; the largest index of variety of connectives in both languages belongs to additive conjunction, the lowest diversity has been demonstrated by adversative conjunctions. The analysis of connectives found in English and Lithuanian research articles has led to some certain conclusions about their use preferences in the field of linguistic academic discourse. It has been revealed that the constitutional conjunctions is the biggest (42 %) and the most variable group of all cohesion relations both in English and Lithuanian articles. The opposing conjunctions of the semantic subgroup in the English articles (27 %) is bigger than in the Lithuanian (14 %) ones. The conjunctions of reason comprise 33 % in the Lithuanian and 17 % in the English texts. The scientific discourse of both languages has expletive means of the sentence conjunctions that make the text coherent and logical.DOI: http://dx.doi.org/10.5755/j01.sal.0.20.1188
The study was carried out in the mainstream of linguistic phenomena in philology. The article is devoted to the study of the speech of the tatars, enduring in the cities of Urumqi and Kuldja of the People´s Republic of China. Studied the general characteristic of the tatar speech of China diaspora and considered the features of the use of the tatar language in the region Investigated some of the lexical phenomenon, an old vocabulary, drawing, and synonyms in the language of the China tatars. Discusses some of the phonetic phenomena in the field of substitution of vowels and changes consonants in a speech, which are directly connected with lexical norms of the tatar literary language and its dialects. The same examples as from the oral speech of the inhabitants of Kuldja and Urumqi, as well as from folklore material of the Tatar Diaspora in China. Identified preconditions of use of the tatar language, similarities and peculiarities of use of native Turkic tokens and borrowed words in the speech of the tatars, living in the PRC.
French Sign Language has a far smaller specialised lexicon than French, which poses regular problems to interpreters working between the two. Four management control sessions attended by a deaf student and interpreted for him by four professional interpreters were recorded, and the interpreters' tactics when encountering the problems of missing signs in French Sign Language ('lexical gaps') were identified, counted and analyzed. Lexical gaps were found to be numerous in the corpus. The tactics often used elements of French spoken language, in contradiction with a strong sociolinguistic norm in the French deaf community. This can be explained by the interpreters' wish to cater to the needs of the deaf student, who needed to know the French terms when taking exams, and is in line with skopos theory.
This paper explores the lexical semantic properties of five near-synonymous Chinese words expressing the emotion of SHAME. The concept of self-construal is vital in understanding emotions such as shame as it relies on the reflections of oneself. The interdependent self-construal is a view of the self through relationship with others and it is related to the characteristics of SHAME in Chinese context. The current study carried out an in-depth examination of how interdependent self-construal shapes the shame concept in Chinese, and how features concerning "self versus others" are encoded in Chinese shame words. The "self versus others" features that we look at include cause attribution (to self vs. others), probable relevant outcome (affect self vs. others), social relations (between self and others that cause shame), social norm (personal values vs. social norms), and presence (or absence) of audience. We examine whether and how these features play a crucial role in the Chinese SHAME concept and how they contribute to the differences between the shame words in Mandarin Chinese. The features can be described as a dimension that is implicit in the denotative meaning of these words.
The article focuses on some chosen axiological issues characteristic of homilies. The problems in question are considered from the philological point of view and with reference to a transdisciplinary perspective due to the status of the object of research. In homilies, which are considered by the believers as the Word of God present in the human word, the axiological horizon is outlined precisely. The axiological sphere exhibits permanence both in the respect of the set of values it comprises and in the respect of their hierarchy, the shape of this sphere being determined by the Gospel. The transcendent values are duly conceptualized and as such they remain the measure and the source of all other valuation procedures. It is in the light of these values, above all, that moral norms are interpreted. Particular values are given a more detailed aspectual reference in separate messages, or preaching units, yet it is always done in accordance with the doctrine. The process of more detailed value description involves the means of valuation which are to a large extent unchangeable (in particular the lexical ones). The hierarchy of values which is considered as model comprises significant revaluations, as it describes as good such vital values as old age, illness or death and such material ones as poverty, and in doing so invokes the authority of Christand his axiological judgments as they appear in the Gospel. While presenting various duties of the followers of Christ, preachers interpret particular elements of values typical of the individuals (functioning in particular conditions) within the axiological space laid out by God himself for the community of believers. Homilies as such happen to be better or worse as far as their stylistic and linguistic aspects are concerned. They are also subject to the norm of multi-sector propriety, so judgments pertaining to their correctness in the above-mentioned aspects need to be formulated carefully. Translated by Dorota Chabrajska
Cette thèse porte sur une étude de la variation et du changement lexicaux des mots référant aux notions de « véhicule automobile » et de « travail rémunéré » dans le français de l’Outaouais, une variété de français laurentien caractérisée par le bilinguisme équilibré et stable et le contact intense avec l’anglais. La thèse est réalisée dans le cadre de la sociolinguistique variationniste labovienne combinée avec des méthodes quantitatives et des techniques analytiques multivariationnelles des règles variables. Cette étude se base sur les données empiriques recueillies dans les communautés francophones de la région de la capitale canadienne parmi les locuteurs nés entre 1846 et 1994 (RFQ, Ottawa-Hull, FdO).\nLe chapitre 2 suit l’évolution sémantique des termes lexicaux, étudie un système d’interaction des facteurs historiques et ceux socialement motivés, et examine l’hypothèse du développement interne du vocabulaire du français canadien. Les chapitres 3 et 4 examinent la corrélation des facteurs liés au bilinguisme et au contact avec l’anglais avec la fréquence d’emploi des variables lexicales; et la marque sociale des variantes lexicales.\nCette thèse: i) met en valeur la méthodologie variationniste quantitative dans l’étude de la variation lexicale; ii) approfondit plusieurs réflexions théoriques et des patrons classiques sur la théorie variationniste; iii) caractérise le lien dynamique entre le parler des locuteurs et les normes de la communauté à laquelle ils se rattachent; iv) contribue à la meilleure compréhension de la dynamique lexicale en fonction du statut du français en situation de contact de langues.
This thesis investigates the changes in whether compound nouns were closed (written as one word), open (written as separate words) or hyphenated in Early New High German between 1550 and 1710. Due to the fact that there were no orthographic norms in the German of this time, graphematic phenomena in this period of the German language are very fruitful to examine. The study is based on a corpus of 249 sermons in 90 different postils. Since this thesis aims to show a diachronic development, the corpus texts originate from six time windows centred around the years 1550, 1570, 1600, 1620, 1660 and 1710. The results of the study show a general development from 1550, when around 80% of the occurrences of compound nouns were written as one word, to 1620, when this way of writing dominated almost entirely. In the texts from the last two time windows, the hyphenation spreads, and by 1710, nearly two thirds of the instances of compound nouns were written with a hyphen. The present study also shows that the geographical origin of a text is of lesser importance for the writing of compound nouns as one word, separate words or with a hyphen. However, the distinction between genuine compound nouns (a compound noun with the modifier in an unmarked case) and artificial ones (a compound noun with the modifier in an oblique case) seems to be of greater relevance. The artificial compound nouns are closed to a lesser extent in the period between 1550 and 1620 and hyphenated to a higher extent from 1660 onwards than the genuine compound nouns. In a second part of the study, the compound nouns of the different time windows are examined from a lexical point of view, showing that many compound noun lexemes were almost consistently written in the same way (either as one word, as separate words, or with a hyphen) in all occurrences within each time window.
The author analyzes the point of view of Vodjidali Mudjmali, the author of the well-known work “Matlat-ul-Ulum va Madjma-ul-Funun” dwelling on the ways of the right usage in regard to lexical units and grammatical constructions. In the tenth chapter of his hook Vodjidali Mudjmali registers numerous cases of violation of grammatical norms and proper combinability of words, he offers 11 factors which might promote smoothness and expressiveness of speech.
We present the ongoing development of MCG, a linguistically deep and precise grammar for Mandarin Chinese together with its accompanying treebank, both based on the linguistic framework of HPSG, and using MRS as the semantic representation. We highlight some key features of our grammar design, and review a number of challenging phenomena, with comparisons to alternative linguistic treatments and implementations. One of the distinguishing characteristics of our approach is the tight integration of grammar and treebank development. The two-step treebank annotation procedure benefits from the efficiency of the discriminant-based annotation approach, while giving the annotators full freedom of producing extra-grammatical structures. This not only allows the creation of a precise and full-coverage treebank with an imperfect grammar, but also provides prompt feedback for grammarians to identify the errors in the grammar design and implementation. Preliminary evaluation and error analysis shows that the grammar already covers most of the core phenomena for Mandarin Chinese, and the treebank annotation procedure reaches a stable speed of 35 sentences per hour with satisfying quality. Keywords:Grammar Engineering, Treebank Annotation, Syntax 1.
Reviewed by: Nova gramática do português brasileiro Gláucia Silva Castilho, Ataliba T. de. Nova gramática do português brasileiro. São Paulo: Contexto, 2010. Pp. 768. ISBN 978-85-7244-462-0. Innovative. The adjective, used in the preface (25), accurately describes this grammar, which does not resemble any other that I have seen. “Impressive” would be another good descriptor for this 768-page volume—not because of the number of pages, but because of its richness and its depth. This is a comprehensive volume dedicated exclusively to the grammar of Brazilian Portuguese (henceforth BP). Here, “grammar” does not mean a book or a discipline that dictates norms, but a set of natural rules found in spoken language. Castilho sets out to describe and explain these rules, based, in great part, on corpora of spoken BP. He does it successfully and clearly (more on this point below), but not before laying out, in chapter 1, what language and grammar are, from different points of view. In this chapter, besides explaining different theories and views of grammar, the author discusses linguistic policies for BP, not only as a first language, but also as a foreign language (calling for the Brazilian government to create an organization that could be in charge of implementing such a policy, as Instituto Camões does for Portugal or Instituto Cervantes for Spain). [End Page 181] As if an in-depth discussion of what constitutes grammar were not enough, Castilho dedicates several more chapters to different subdisciplines of linguistics. He delves into the language as a multisystem in chapter 2, in which he discusses the lexicon/lexicalization, semantics/semanticization, discourse/discursivization, and grammar/grammaticalization. Next we find a chapter on the history of BP, in which a discussion of the social history of the language is included. Chapter 4 tackles variation in BP, both geographical and social, while chapter 5 discusses conversational and textual analysis. All of these are elements not found in a traditional grammar. It is true that this is a “new grammar” as stated in the title, but it is much more than that: it is an introduction to linguistics, where the reader is not only presented with accounts based on theory, but also invited to conduct research. Castilho dedicates the last chapter (chapter 15) to a call for continuing research on BP, laying out several possible areas that might interest the reader, along with suggested texts that would jump start research in each area. The “more traditional” (28) topics in this grammar start with the sentence, which is the subject of four chapters (chapters 6, 7, 8, and 9). In these, the author examines just about everything at the sentence level: from grammatical, semantic, and discursive properties to modality and typology. Naturally, we find analyses of expected categories, such as the subject, complements, and adjuncts. However, these are presented in light of research, providing linguistic accounts of the facts related to each of these categories. The examination of the sentence is followed by chapters on the verb phrase, the noun phrase, the adjective phrase, the adverbial phrase, and the prepositional phrase. As with the other topics presented in the grammar, these chapters contain in-depth explorations carried out in light of linguistic theory. In this grammar, the solid theoretical underpinnings are illustrated with citations of relevant research that support each point discussed, always based on examples from spoken language. Throughout the book, Castilho invites the reader to continue studying each topic, providing lists of pertinent readings. All of this is done in a clear style that addresses the reader directly, as if the author were talking to her/him. Castilho’s style is truly refreshing and can make the reading quite fun—a quality not normally associated with grammar treatises. In spite of the lengthy discussions (this is not a quick reference grammar), the author’s writing style is key to the success of his explanations. Through questions that he supposes the reader might ask, he makes the reader a type of coauthor (33), thus engaging her/him in the quest for answers. Contributing to the clarity of the text are the tables that we find in most chapters...
The article covers analysis of lexical units, which represent the concept of the child in folk culture. At present dialectal cultural linguistics, called to model dialectal linguistic world-image, is quickly developing. The urgency of studying folk dialects and dialectal word's cultural meanings is caused by social aspiration for self-cognition, which among other things is achieved by means of traditional culture exploration. The present research has been based on the material of Middle Priobie dialects. The object of the research is a dialectal word, which contains a cultural component in its semantic structure. The approach to the lexical dialect system is realized through characterizing the concept of the child from the positions of cultural linguistics. The choice of this concept as the research object is caused by the fact that from the scientific point of view childhood is a particular phenomenon, which, when studied, shows the world of ''adult'' culture and makes it possible to remodel its world-view principles. In the peasants' world family is the basic community unit, which explains the village society's great attention to inter-family relationship. Families, in which parents have children together, are considered the standard. The deviation from the norm is registered by means of the language. The article covers lexical representation of three situations closely connected with the child's family status: belonging to one of the spouses only, orphanhood and illegitimacy. On traditional mind's mental level there is an opposition of related and unrelated children, which is realized in syntagmatic expansion. The ''related unrelated'' opposition is a special case of ''one's own somebody else's'' opposition, according to which everything that is not ''one's own'' or ''together'' is estranged. The idea of ''extraneity'' models stereotyped images of the unjustly oppressed stepson and stepdaughter. The denomination ''orphan'' marks a child who lost parents out of other children's mass. In folk mind the image of the orphan is closely connected with the notion of fate. Orphan's emotional deprivation is explained by his loneliness. Compassionate treatment of orphans is represented on the language level through the ability of this word to have a diminutive form and its semantic support. The natural child in traditional culture is a marginal creature by birth. Since he was born out of wedlock, outside the law, he is unprotected before the society. Negative attitude towards the woman who gave birth out of wedlock and her child is realized in abusive, insulting designations. There are also denominations formed from reputed loci of the natural child's conception or birth. Their semantics is opposed to the idea of the house as ''one's own'' space. It is remarkable that there are no particular appellations for legitimate children since legitimacy is considered as a norm and is not marked by linguistic means. Thus, the consideration of linguistic realisation of the notion of the child's family status allows us model a fragment of the native dialect speaker's value worldimage. It is possible to make a conclusion that family occupies one of the fundamental places in the world-view constants of traditional culture.
Lexical access is the process in which basic components of meaning in language, the lexical entries (words) are activated. This activation is based on the organization and representational structure of the lexical entries. Semantic features of words, which are the prominent semantic characteristics of a word concept, provide important information because they mediate semantic access to words. An experiment was conducted to examine the importance of semantic feature distinctiveness and feature frequency in accessing the lexical representations of young and older adults in an off-line task using features of animals. The McRae, Cree, Seidenberg, and McNorgan (2005) feature norm corpus is the basis for the selection of stimuli for the current research project. Semantic features were utilized to explore the structure of the lexicon. Stimuli varied in feature distinctiveness based on the study by McRae, et al. (2005) in 3 broad stimulus groups: Distinctive (D), Low Frequency Non-Distinctive (LFND), and Non-Distinctive High Frequency (NDHF). Participants were asked to list all of the concepts that came to mind for a given feature in an untimed task. Distinctiveness was examined between stimulus groups for the number of concepts and variety of first concepts given to the presented feature. It was found that fewer concepts were given and there was less variety in first concepts given for the distinctive features and the most concepts and greater variety of first concepts were given for the high-frequency non-distinctive features. Distinctiveness appears to vary along a continuum, supporting theories of lexical access based on activation and competition between concept words. Additionally, participant age groups were compared for the number of concepts given and the variety of first concepts given. The older adult group produced more concepts and more variety of first concepts than the younger group, in all three feature categories. These results indicate that greater (lifetime) language experience of the participants in the older group was reflected in their performance. A continued interest in semantic features is important to our understanding of the influence of features on the retrieval of semantic concepts and the changes in those retrieval processes over the lifespan.
This stylebook is an updated version of Telljohann et al. (2006). It describes the design principles and the annotation scheme for the German treebank TüBa-D/Z developed by the Division of Computational Linguistics (Lehrstuhl Prof. Hinrichs) at the Department of Linguistics (Seminar für Sprachwis-
The subject of the article is the actual state and functioning of modern Brasilian Portuguese language norm. The European Portuguese language norm, considered as a standard in Brasil, in the recent decades has been losing its social base, while the national (Brasilian) forms of speech are being actively recognized in grammatical and lexical descriptions.
This paper deals with Discourse Argument Identification (DAI) from both intra-sentence and inter-sentence perspectives. For intra-sentence cases, we approach it via a simplified shallow semantic parsing framework, which recasts the discourse connective as the predicate and its scope into several constituents as the argument of the predicate. Different from state-of-the-art chunking approaches, our parsing approach extends DAI from the chunking level to the parse tree level, where rich syntactic information is available, and focuses on determining whether a constituent, rather than a token, is an argument or not. For inter-sentence cases, we present a lightweight heuristic rule-based solution. Evaluation using Penn Discourse Treebank (PDTB) shows that the current research’s parsing approach significantly outperforms the state-of-the-art chunking alternatives.
This paper discusses hybridization in contemporary Russian language. In particular, the work focuses on Dina Rubina�s novel Here comes the Messiah! (?????????????!.....).Her prose is characterized by a wide range of linguistic and communicative means; a key role is played by linguistic hybridization which alienates the text from any contexts, notions of time, as well as the ordinary and accepted linguistic norms. In her work, Rubina - a migrant writer living in Israel - describes a carnival atmosphere (based on Bachtin�s idea of carnival) as one of the main leitmotivs of the Russian community in Jerusalem. Words - which are are the primary means of expression in a world considered as a stage - often undergo hybridization processes. Therefore, new meanings, neologisms, original lexemes decorate Rubina�s multicultural text.
This work describes how derivation tree fragments based on a variant of Tree Adjoining Grammar (TAG) can be used to check treebank consistency. Annotation of word sequences are compared both for their internal structural consistency, and their external relation to the rest of the tree. We expand on earlier work in this area in three ways. First, we provide a more complete description of the system, showing how a naive use of TAG structures will not work, leading to a necessary refinement. We also provide a more complete account of the processing pipeline, including the grouping together of structurally similar errors and their elimination of duplicates. Second, we include the new experimental external relation check to find an additional class of errors. Third, we broaden the evaluation to include both the internal and external relation checks, and evaluate the system on both an Arabic and English treebank. The evaluation has been successful enough that the internal check has been integrated into the standard pipeline for current English treebank construction at the
The social and cultural ‘turn’ in language education of recent years has helped move language teaching and curriculum design away from many of the more rigid dogmas of earlier generations, but the issue of the roles of the learners’ first language (L1) in language pedagogy and classroom interaction is far from settled. Some follow a strict ‘exclusive target language’ pedagogy, while others ‘resort to’ the use of the L1 for a variety of purposes (see ACTFL 2008). Underlying these competing views is the perspective of the L1 as an impediment to second language learning. Following sociocultural theory and ecological perspectives of language and learning and based on the findings of research on classroom code-switching and code choice, this paper lays out an approach to the language classroom as a multilingual social space in which learners and teacher study, negotiate, and co-construct code choice norms toward the dynamic, creative, and pedagogically effective use of both the target language and the learners’ L1(s). Learner use of the L1 for the purpose of grammatical or lexical learning is also considered, and some examples for instruction are offered.
This paper presents an ongoing project whose goal is to create a freely available dependency treebank for Persian. The data is taken from the Bijankhan corpus, which is already annotated for parts of speech, and a syntactic dependency annotation based on the Stanford Typed Dependencies is added through a bootstrapping procedure involving the open-source dependency parser MaltParser. We report preliminary parsing experiments with promising results after training the parser on a manually annotated seed data set of 215 sentences.
Texts The Prague Czech-English Dependency Treebank 2.0 (PCEDT 2.0) is a major update of the Prague Czech-English Dependency Treebank 1.0 (LDC2004T25). It is a manually parsed Czech-English parallel corpus sized over 1.2 million running words in almost 50,000 sentences for each part. Data The English part contains the entire Penn Treebank - Wall Street Journal Section (LDC99T42). The Czech part consists of Czech translations of all of the Penn Treebank-WSJ texts. The corpus is 1:1 sentence-aligned. An additional automatic alignment on the node level (different for each annotation layer) is part of this release, too. The original Penn Treebank-like file structure (25 sections, each containing up to one hundred files) has been preserved. Only those PTB documents which have both POS and structural annotation (total of 2312 documents) have been translated to Czech and made part of this release. Each language part is enhanced with a comprehensive manual linguistic annotation in the PDT 2.0 style (LDC2006T01, Prague Dependency Treebank 2.0). The main features of this annotation style are: dependency structure of the content words and coordinating and similar structures (function words are attached as their attribute values) semantic labeling of content words and types of coordinating structures argument structure, including an argument structure ("valency") lexicon for both languages ellipsis and anaphora resolution. This annotation style is called tectogrammatical annotation and it constitutes the tectogrammatical layer in the corpus. For more details see below and documentation. Annotation of the Czech part Sentences of the Czech translation were automatically morphologically annotated and parsed into surface-syntax dependency trees in the PDT 2.0 annotation style. This annotation style is sometimes called analytical annotation; it constitutes the analytical layer of the corpus. The manual tectogrammatical (deep-syntax) annotation was built as a separate layer above the automatic analytical (surface-syntax) parse. A sample of 2,000 sentences was manually annotated on the analytical layer. Annotation of the English part The resulting manual tectogrammatical annotation was built above an automatic transformation of the original phrase-structure annotation of the Penn Treebank into surface dependency (analytical) representations, using the following additional linguistic information from other sources: PropBank (LDC2004T14) VerbNet NomBank (LDC2008T23) flat noun phrase structures (by courtesy of D. Vadas and J.R. Curran) For each sentence, the original Penn Treebank phrase structure trees are preserved in this corpus together with their links to the analytical and tectogrammatical annotation.
The present study focuses on the characteristics of parental child-directed communication and its relationship with child language development. For this purpose, thirty-six toddlers (18 males and 18 females) and their parents were observed in a laboratory during triadic free play at ages 1; 3 and 1; 9. The characteristics of the maternal and paternal child-directed language (characteristics of communicative functions and lexicon as reported in psycholinguistic norms for Italian language) were coded during free play. Child language development was assessed during free play and at ages 2; 6 and 3; 0 using the Italian version of the MacArthur-Bates Communicative Development Inventory (2; 6) and the revised Peabody Picture Vocabulary Test (PPVT-R) (3; 0). Data analysis indicated differences between mothers and fathers in the quantitative characteristics of communicative functions and language, such as the mean length of utterances (MLU), and the number of tokens and types. Mothers also produced the more frequent nouns in the child lexicon. There emerged a relation between the characteristics of parental child-directed language and child language development.
Annotation of discourse relations is a project related to the Prague Dependency Treebank 2.5. It represents a new manually annotated layer of language description, above the existing layers of the PDT, and it portrays linguistic phenomena from the perspective of discourse structure and coherence.
We present the Prague Dependency Treebank 2.5, the newest version of PDT and the first to be released under a free license. We show the benefits of PDT 2.5 in comparison to other state-of-the-art treebanks. We present the new features of the 2.5 release, how they were obtained and how reliably they are annotated. We also show how they can be used in queries and how they are visualised with tools released alongside the treebank.
Annotated corpora such as treebanks are important for the development of parsers, language applications as well as understanding of the\nlanguage itself. Only very few languages possess these scarce resources. In this paper, we describe our efforts in syntactically annotating\na small corpora (600 sentences) of Tamil language. Our annotation is similar to Prague Dependency Treebank (PDT) and consists of\nannotation at 2 levels or layers: (i) morphological layer (m-layer) and (ii) analytical layer (a-layer). For both the layers, we introduce\nannotation schemes i.e. positional tagging for m-layer and dependency relations for a-layers. Finally, we discuss some of the issues in\ntreebank development for Tamil.
The annotation of large corpora is usually restricted to syntactic structure and word class. Pure lexical information and information on the structure of words are stored in specialized dictionaries (Baayen et al., 1995). Both data structures ‐ dictionary and text corpus ‐ can be matched to get e.g. a distribution of certain (restricted) lexical information from a text. This procedure works fine for synchronic corpora. What is missing, however, is either a special mark-up in texts linking each of the items to a certain time or a diachronic lexical database that allows for the matching of the items over time. In what follows, we take the latter approach and present a tool set (MoreXtractor, Morphilizer, MorQuery), a database (Morphilo-DB) and the architecture of a platform (Morphorm) for a sustainable use of diachronic linguistic data for Middle English, Early Modern English and Modern English.
The central problems that this paper addresses are (i) the lack of large and rich formalised lexicons for multi-word expressions for use in Natural Language Processing (NLP); (ii) the lack of proper methods and tools to extend the lexicon of an NLP-system for multi-word expressions given a text corpus in a maximally automated manner. The paper describes innovative methods and tools for the automatic identification and lexical representation of multi-word expressions. In addition, it describes a 5.000 entry corpus-based multi-word expression lexical database for Dutch developed using these methods. The database has been externally validated, and its usability has been evaluated in NLP-systems for Dutch. The MWE database developed fills a gap in existing lexical resources for Dutch. The generic methods and tools for MWE identification and lexical representation focus on Dutch, but they are largely language-independent and can also be used for other languages, new domains, and beyond this project. The research results and data described in this paper contribute directly to strengthening the digital infrastructure for Dutch.
The study presented in this article is dedicated to a syntactic parser for Romanian. The central goal of the presented technique is to learn a model which is able to discriminate between probability for a word to be head of another word in a dependency structure corresponding to a sentence in the considered language. The model described in this paper was trained on a dependency treebank linguistic resource and is intended to be used in order to develop a dependency syntactic parser.
International audience
Cet article constitue une version réduite de l'article "The French Social Media Bank: a Treebank of Noisy User Generated Content" (mêmes auteurs)
Après un bref rĂŠsumĂŠ de la thĂŠorie de Topic-Focus Articulation (TFA), la prĂŠsente ĂŠtude dĂŠmontre, à l'aide de plusieurs exemples illustrant l'annotation de principaux traits de TFA sur un large corpus (the Prague Dependency Treebank), que l'annotation du corpus apporte une valeur ajoutĂŠe au corpus, si deux conditions sont rĂŠunies: (i) le schĂŠma de l'annotation est basĂŠ sur une thĂŠorie linguistique solide, (ii) le procĂŠdĂŠ d'annotation est ĂŠtabli avec soin (c'est-à-dire de façon systĂŠmatique et cohĂŠrente). Une telle annotation est importante non seulement pour la structure de surface de la phrase mais encore davantage pour la structure phrastique sous-jacente, car elle est susceptible de mettre en ĂŠvidence les phĂŠnomènes cachĂŠs au niveau de la structure de surface, mais incontournables lors de la reprĂŠsentation du sens et du fonctionnement de la phrase.
State-of-the-art dependency representations such as the Stanford Typed Dependencies may represent the grammatical relations in a sentence as directed, possibly cyclic graphs. Querying a syntactically annotated corpus for grammatical structures that are represented as graphs requires graph matching, which is a non-trivial task. In this paper, we present an algorithm for graph matching that is tailored to the properties of large, syntactically annotated corpora. The implementation of the algorithm is built on top of the popular IMS Open Corpus Workbench, allowing corpus linguists to re-use existing infrastructure. An evaluation of the resulting software, CWB-treebank, shows that its performance in real world applications, such as a web query interface, compares favourably to implementations that rely on a relational database or a dedicated graph database while at the same time offering a greater expres-sive power for queries. An intuitive graphical interface for building the query graphs is available via the Treebank.info project.
This paper presents the work of the Hong Kong Polytechnic University (PolyUCOMP) team which has participated in the Semantic Textual Similarity task of SemEval-2012. The PolyUCOMP system combines semantic vectors with skip bigrams to determine sentence similarity. The semantic vector is used to compute similarities between sentence pairs using the lexical database WordNet and the Wikipedia corpus. The use of skip bigram is to introduce the order of words in measuring sentence similarity. 1
This paper describes the support for mouth activity annotation provided by the iLex annotation workbench on a holistic level connected to the lexical database, on a feature level, as well as in the context of semi-automatic annotation.