Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
ABSTRACT. The present article has as its purpose to present way in which vocabulary of Romanian language from last two decades has known very important modifications, as result of changes present in political, economic, social, cultural plan as well as at level of mentalities. These modifications should be focus of attention for those concerned with communication issues, but especially for those who work or are in process of training to work in specialized translations departments or in mass-media.Keywords: connection between language and society; predominance of technical-scientific language; metaphorization of specialized terms; demarcation literary-non-literary; Romanian language of nowadays: mixture of Anglo-Romanian jargon, pretentious expressions, slang and familiar termsWe all agree that there is very tight connection between language and society, periods of major changes in life of society (as period our society has been going through since 1990) leading to important modifications of verbal means that we use to communicate. At level of Romanian vocabulary, some words and some meanings associated to words fall into disuse or acquire new connotations, whereas other words or meanings of words emerge - many of which are not yet to be found as dictionary entries. Moreover, stylistic modifications regarding relation between language variants are also obvious. Under these circumstances, in order to consolidate and disseminate cultivated variant of Romanian language, in 2005, Romanian Academy, more precisely Iorgu Iordan - Al. Rosetti Institute of Linguistics succeeded in publishing two fundamental works: Dicfionarul ortografie, ortoepic §i morfologic al limbii române (The Orthographic, Orthoepie and Morphological Dictionary of Romanian, second edition, also called DOOM2) and Gramatica limbii române (Romanian Grammar). On occasion of second edition of DOOM, president of Romanian Academy at that time, Eugen Simion, defined it as follows: a necessary work, that large public has been waiting for, work of national interest, which will from now on be used as unique source in correctly applying academic norms in domain of Romanian orthography. Our DOOM appears at time it is highly needed. Needless to say that this is work that translators cannot do without.Hosting television show at România de Mâine (Romania Tomorrow) television, linguistic culture show - Speak, write Romanian - I asked academician Eugen Simion to briefly characterize way in which Romanian language is used. His answer was that the Romanian language has become ugly in sense that way most of us speak is characterized by relaxing exigencies, allowing unliterary linguistic elements such as: regionalisms, popular expressions and even slang expressions to penetrate neat language. The explanation is simple: Romanian society has been confronting with an acute crisis of values in these last two decades, which has led to horizontal positioning, to accepting both value and non-value, to socially imposing model of an individual who, as concerns norms of Romanian language, adopts principle of character mam-mare (big mama) from famous Romanian sketch Mr. Goe, by Caragiale: may speak as you know, I speak as I speak. And when half-leamedness shakes hands with preciosity and snobbery, one may commit mistakes that are harder and harder to eliminate (Theodor Hristea). According to specialists' estimations, most sensible 'seismograph' that register changes that appear in political, economic, social, cultural plan as well as at level of mentalities is publicistic language, always in permanent search for expressive innovation and lexical pictoresqueness.Although language, seen from historical perspective, appears as an element of stability in people's life, it is nevertheless subject to variation and change. …
The article deals with the linguistic issues of composing a reference book of regional toponyms – a genre that requires special consideration in national lexicography. The assortment of these issues gave the possibility to carry out complex description of regional toponyms on the basis of semantic, functional, and orpthologuos criteria that let unify the names of Volgograd region settlements that are registered in various documents. The significance of the composed reference book is determined by several factors – the presence of local subsystems of geographical names in Russian toponymy; the inconsistency of current orthography norms on using capital letter in compound proprius names and fused-with-hyphen spelling of toponyms and off-toponym derivations; the lack of linguistically justified explanation of peculiarities of grammatical norms in the field of proper names use. The reference book of regional toponyms is based on the object description (toponymic vocabulary), principles of lexical units selection (description of spelling and grammatical properties of toponyms, encyclopedic information), the glossary (full list of toponyms of Volgograd region), typical article. The articles in the reference book are arranged in lexicographical zones with grammatical and semantic markers, lexicographical illustrations, other lexicographical labels, word etymology including. The reference book on Volgograd region toponymy is addressed to executive and administration authorities, journalists, regional ethnographers.
This study presents first-year findings of a 25-week longitudinal project derived from a two-year longitudinal randomized trial study at the elementary school level in Costa Rica on effective computer-assisted language learning (CALL) approaches in an English as a foreign language (EFL) setting. A pre-test–post-test experimental group design was implemented to evaluate two varying types of CALL curriculum (Treatment A and Treatment B, both technology based to assist English language learning) with a difference in students’ time-on-task, as opposed to the Control/Comparison group (that received no treatment and was a typical practice within the regular English teaching time period). Four subtests from Woodcock Munoz Language Survey-Revised (WMLS-R), a norm-referenced, standardized instrument, were selected to monitor participants’ oral English development. A total of 76 urban, rural, and urban/marginal schools with 816 third graders were included in the analysis through multilevel modeling. Results suggested that (1) students held very limited oral English proficiency at the beginning of the third grade; (2) although students significantly improved their oral English proficiency during the 25-week intervention, they were still significantly below the typical native English-speaking norm at the end of the third-grade level; (3) those who were exposed to CALL modules in Treatment A developed at a faster rate than did students in Treatment B and in Control classrooms in lexical knowledge and listening skills when statistically controlling for student-level variables, including initial level and time-on-task; (4) although students in CALL intervention (especially in Treatment B) started with a lower level of oral English proficiency, their gain was numerically higher than that in the Control condition; and (5) time-on-task demonstrated to be an irrelevant variable in the study. These findings imply that it is not just exposure to English that matters for significant gains in the language; rather, it is the type of instruction a student receives, or the quality of instruction in which the software engages the students interactively in a hands-on, minds-on, scaffolded manner, that matters the most in developing steep gains. Finally, we recommend that additional research be conducted with groups that move through kindergarten, first grade, and second grade longitudinally to determine cohort effects in learning English via CALL instruction in EFL countries.
How does the presence of a categorically related word influence picture naming latencies? In order to test competitive and noncompetitive accounts of lexical selection in spoken word production, we employed the picture-word interference (PWI) paradigm to investigate how conceptual feature overlap influences naming latencies when distractors are category coordinates of the target picture. Mahon et al. (2007. Lexical selection is not by competition: A reinterpretation of semantic interference and facilitation effects in the picture-word interference paradigm. Journal of Experimental Psychology. Learning, Memory, and Cognition, 33(3), 503-535. doi:10.1037/0278-7393.33.3.503 ) reported that semantically close distractors (e.g., zebra) facilitated target picture naming latencies (e.g., HORSE) compared to far distractors (e.g., whale). We failed to replicate a facilitation effect for within-category close versus far target-distractor pairings using near-identical materials based on feature production norms, instead obtaining reliably larger interference effects (Experiments 1 and 2). The interference effect did not show a monotonic increase across multiple levels of within-category semantic distance, although there was evidence of a linear trend when unrelated distractors were included in analyses (Experiment 2). Our results show that semantic interference in PWI is greater for semantically close than for far category coordinate relations, reflecting the extent of conceptual feature overlap between target and distractor. These findings are consistent with the assumptions of prominent competitive lexical selection models of speech production.
Although laboratory phonology techniques have been widely employed to discover the interplay between the acoustic correlates of English Lexical Stress (ELS)–fundamental frequency, duration, and intensity - studies on ELS in polysyllabic words are rare, and cross-linguistic acoustic studies in this area are even rarer. Consequently, the effects of language experience on L2 lexical stress acquisition are not clear. This investigation of adult Arabic (Saudi Arabian) and Mandarin (Mainland Chinese) speakers analyzes their ELS production in tokens with seven different stress-shifting suffixes; i.e., Level 1 [+cyclic] derivations to phonologists. Stress productions are then systematically analyzed and compared with those of speakers of Midwest American English using the acoustic phonetic software, Praat. In total, one hundred subjects participated in the study, spread evenly across the three language groups, and 2,125 vowels in 800 spectrograms were analyzed (excluding stress placement and pronunciation errors). Nonnative speakers completed a sociometric survey prior to recording so that statistical sampling techniques could be used to evaluate acquisition of accurate ELS production. The speech samples of native speakers were analyzed to provide norm values for cross-reference and to provide insights into the proposed Salience Hierarchy of the Acoustic Correlates of Stress (SHACS). The results support the notion that a SHACS does exist in the L1 sound system, and that native-like command of this system through accurate ELS production can be acquired by proficient L2 learners via increased L2 input. Other findings raise questions as to the accuracy of standard American English dictionary pronunciations as well as the generalizability of claims made about the acoustic properties of tonic accent shift.
Social interaction deficits in drug users likely impede treatment, increase the burden of the affected families, and consequently contribute to the high costs for society associated with addiction. Despite its significance, the neural basis of altered social interaction in drug users is currently unknown. Therefore, we investigated basal social gaze behavior in cocaine users by applying behavioral, psychophysiological, and functional brain-imaging methods. In study I, 80 regular cocaine users and 63 healthy controls completed an interactive paradigm in which the participants' gaze was recorded by an eye-tracking device that controlled the gaze of an anthropomorphic virtual character. Valence ratings of different eye-contact conditions revealed that cocaine users show diminished emotional engagement in social interaction, which was also supported by reduced pupil responses. Study II investigated the neural underpinnings of changes in social reward processing observed in study I. Sixteen cocaine users and 16 controls completed a similar interaction paradigm as used in study I while undergoing functional magnetic resonance imaging. In response to social interaction, cocaine users displayed decreased activation of the medial orbitofrontal cortex, a key region of reward processing. Moreover, blunted activation of the medial orbitofrontal cortex was significantly correlated with a decreased social network size, reflecting problems in real-life social behavior because of reduced social reward. In conclusion, basic social interaction deficits in cocaine users as observed here may arise from altered social reward processing. Consequently, these results point to the importance of reinstatement of social reward in the treatment of stimulant addiction.
The aim of this paper is to illustrate the potential of a parallel corpus in the context of (computer-assisted) language learning. In order to do so, we propose to answer two main questions (1) what corpus (data) to use and (2) how to use the corpus (data). We provide an answer to the what-question by describing the importance and particularities of compiling and processing a corpus for pedagogical purposes. In order to answer the how-question, we first investigate the central concepts of the interactionist theory of second language acquisition: comprehensible input, input enhancement, comprehensible output and output enhancement. By means of two case studies, we illustrate how the abovementioned concepts can be realized in concrete corpus-based language learning activities. We propose a design for a receptive and productive language task and describe how a parallel corpus can be at the basis of powerful language learning activities. The Dutch Parallel Corpus, a ten-million word sentence aligned and annotated parallel corpus, is used to develop these language tasks.
У статті проаналізовано лексико-семантичні, морфологічні та словотвірні процеси в українській мові у зв’язку з тенденціями динаміки літературної норми. Особливу увагу звернено на функціювання варіантів, узагальнено основні напрямки лексичних і граматичних змін, зокрема взаємодії власне української і запозиченої лексики, розкриття внутрішнього словотвірного потенціалу. На основі зіставно-порівняльного аналізу даних історичних та сучасних граматик і функційних особливостей різнорівневих мовних одиниць простежено зміни в нормативному статусі, визначено загальні принципи їх кодифікації. (The research presents an analysis of lexical, semantical word-formative and morphological processes in Ukrainian language in accordance with the tendencies of the literary norm dynamics. Special attention is paid to the functioning of variants. The main trends of lexical and grammatical changes, and the interaction of originally and borrowed grammatical, the disclosing of inherent laws of development of word-building, have been generalized. The juxtapositive and comparative data analysis of functional peculiarities of neutral and stylistically marked words show the changes in the normative status of lexical units and determine the general principles of their codification.)
OBJECTIVE: In alcohol-dependent patients, alcohol cues evoke increased activation in mesolimbic brain areas, such as the nucleus accumbens and the amygdala. Moreover, patients show an alcohol approach bias, a tendency to more quickly approach than avoid alcohol cues. Cognitive bias modification training, which aims to retrain approach biases, has been shown to reduce alcohol craving and relapse rates. The authors investigated effects of this training on cue reactivity in alcohol-dependent patients. METHOD: In a double-blind randomized design, 32 abstinent alcohol-dependent patients received either bias modification training or sham training. Both trainings consisted of six sessions of the joystick approach-avoidance task; the bias modification training entailed pushing away 90% of alcohol cues and 10% of soft drink cues, whereas this ratio was 50/50 in the sham training. Alcohol cue reactivity was measured with functional MRI before and after training. RESULTS: Before training, alcohol cue-evoked activation was observed in the amygdala bilaterally, as well as in the right nucleus accumbens, although here it fell short of significance. Activation in the amygdala correlated with craving and arousal ratings of alcohol stimuli; correlations in the nucleus accumbens again fell short of significance. After training, the bias modification group showed greater reductions in cue-evoked activation in the amygdala bilaterally and in behavioral arousal ratings of alcohol pictures, compared with the sham training group. Decreases in right amygdala activity correlated with decreases in craving in the bias modification but not the sham training group. CONCLUSIONS: These findings provide evidence that cognitive bias modification affects alcohol cue-induced mesolimbic brain activity. Reductions in neural reactivity may be a key underlying mechanism of the therapeutic effectiveness of this training.
In adults, patterns of neural activation associated with perhaps the most basic language skill--overt object naming--are extensively modulated by the psycholinguistic and visual complexity of the stimuli. Do children's brains react similarly when confronted with increasing processing demands, or they solve this problem in a different way? Here we scanned 37 children aged 7-13 and 19 young adults who performed a well-normed picture-naming task with 3 levels of difficulty. While neural organization for naming was largely similar in childhood and adulthood, adults had greater activation in all naming conditions over inferior temporal gyri and superior temporal gyri/supramarginal gyri. Manipulating naming complexity affected adults and children quite differently: neural activation, especially over the dorsolateral prefrontal cortex, showed complexity-dependent increases in adults, but complexity-dependent decreases in children. These represent fundamentally different responses to the linguistic and conceptual challenges of a simple naming task that makes no demands on literacy or metalinguistics. We discuss how these neural differences might result from different cognitive strategies used by adults and children during lexical retrieval/production as well as developmental changes in brain structure and functional connectivity.
Proof theory began in the 1920’s as a part of Hilbert’s program. That program aimed to secure the foundations of mathematics by modeling infinitary mathematics with formal axiomatic systems, and proving those systems consistent using restricted, “finitary” means. The program thus viewed mathematics as a system of reasoning with precise linguistic norms, governed by rules that can be described and studied in concrete terms. Such a viewpoint, today, has applications in mathematics, computer science, and the philosophy of mathematics.
The article is devoted to the study of the specificity of phatic speech on the radio. Phatics in this communication is conditioned by the process of democratization which is the main strategy of the media in general and, as a consequence, by the increased personal element in the sphere of oral public communication. There are two extralinguistic elements of identifying phatics: a) functional, where phatics is defined either as a simple exchange of words contributing to people uniting or as an implementation of the contact-making language function; b) semantic, where the specificity of phatics is determined in contrast to the informative content. 0riginally used in the sphere of everyday communication, phatic speech becomes an integral part of the mass media space today, as it enables to establish the contact with the audience and create an illusion of friendly communication between interlocutors. A radio listener enters the artificial world the presenters create by their specifically organized speech. The phatics in the speech makes public communication similar to the every day one, thus reducing the distance between the presenter and the audience. Approaching the level of their public speech to the language usage of prospective listeners and their language competence, radio presenters create a sense of home atmosphere suitable for informal communication. Colloquial speech is actively used. It adapts elements of different communicative spheres and facilitates their interaction. Colloquial elements express phatic communication on the radio. Infotainment as the principle of informing in the media leads to the activation of the game as a form of communication. The game is a deliberate breaking of the language norms with a pragmatic aim to establish and maintain the expressive phatic contact with the audience through the comic effect. Language game in a variety of techniques functions at different levels of radio broadcasts: program names, presenters' nicknames, presenters' speech. Pun is a common form of a language game; it is based on the use in one text of words with identical or similar sound forms (in precedent texts, idioms). The comic effect is created by a deliberate breach of lexical compatibility laws, constructing unusual lexical units and using circumlocutory nominations. Thus, phatic speech actively penetrating the area of public communication determines the tendency of the modern radio to ambivalence, which combines different phenomena a natural vivid dialogue with the listener and an informative message. Imitation of colloquial speech and language game on the radio bring an element of live human communication and make it more unpredictable and interesting for the audience.
In addition to content and presentation, user-friendliness is gaining attention as an issue in the production of contemporary dictionaries. Electronic dictionaries, especially as they become accessible on the web and mobile devices, often capitalise on interactive user interfaces to enhance word access. Inspired by psychological models of human lexical access, one important direction is to expand navigational means by allowing word searches from a variety of associative relations. Although in practice it may only be feasible to extract various kinds of semantic links from large corpora, word association norms obtained from human subjects nevertheless remain a credible and useful resource for informing such corpus-based methods. This paper reports on a word association test with native Hong Kong Cantonese speakers and discusses the potential application of the resulting psycholinguistic evidence for enhancing the design of Chinese dictionaries, with particular focus on the relative significance of different association types with respect to the properties of individual words.
Some studies have shown that bilinguals gesture more than monolinguals. One possible reason for the high gesture frequency is that bilinguals rely on gestures even more than monolinguals in constructing their message. To test this, we asked French–English bilingual adults and English monolingual adults to tell a story twice; on one occasion they could move their hands and on the other they could not. If gestures aid bilinguals in information packaging and/or lexical access, bilinguals should tell shorter stories with fewer word types than monolinguals when their gestures are restricted. In fact, we found that gesture restriction affected bilinguals’ stories only in French, the language in which they used more gestures. These findings challenge the interpretation that bilinguals gesture frequently as an aid in constructing their message. We argue that cultural norms in gesture frequency interact with gesture use in message construction.
Interlinear Glossed Text (IGT) is a well established data format within philology and the structural and generative fields of linguistics. The best known format for an IGT is the one found in linguistic publications, where one line of text is followed by one line of glosses and one line of free translation. Although used in different functions, IGTs are ubiquitous in linguistic research and publications. Yet they also have been criticised for being fabricated and unreliable in some of their uses. However that might be, IGTs represent linguistic knowledge, and in particular for less-resourced languages, they are not rarely the only structured data available. Under the auspices of the Digital Humanities, linguists increasingly focus on the advantages of Semantic Web technologies. Presenting the modules and procedures of the web-based linguistic application TypeCraft (TC), we outline how the creation of IGTs can become an integral part of a shared linguistic methodology. Linguistic services have the potential of allowing efficient data management, and their strength lies in facilitating new forms of collaboration beyond social networking. They pave the way towards what one might call shared methodologies. In this paper we would like to discuss the linguistic value of web-based technology. By presenting the functionalities of TC and giving a detailed summary of online linguistic data creation and retrieval, we will present external and internal criteria for a single system evaluation of TC centred on usage objectives.
WordNet semantic classes to improve dependency parsing. We study the effect of semantic classes in three dependency parsers, using two types of constituencyto-dependency conversions of the English Penn Treebank. Overall, we can say that the improvements are small and not significant using automatic POS tags, contrary to previously published results using gold POS tags In addition, we explore parser combinations, showing that the semantically enhanced parsers yield a small significant gain only on the more semantically oriented LTH treebank conversion.
Michaił KotinUniwersytet Zielonogórksi, Zielona Góra, Polandmichailkotin1@gmail.com W artykule są rozpatrywane dwa aspekty semantyki jednostek leksykalnych zawierających znaczenie prawdy w języku rosyjskim w zestawieniu z innymi ję-zykami indoeuropejskimi. Pierwszy aspekt badawczy to pojęciowe źródła dwóch rosyjskich rzeczowników: prawda oraz istina, a także pochodnych od nich. Po-kazano, że leksemy te bazują na różnych wyjściowych konceptach. „Norma” (w innych językach jej odpowiednik — postrzegana rzeczywistość) — stanowiąca podstawę konceptu „prawda” — jest niewyspecyfikowana, pozwala na wiele inter-pretacji, ma rozmyte granice, podczas gdy leżąca u podstaw wyrazu istina „istota”, odwrotnie, jest wyspecyfikowana, monosemantyczna i niepodatna na przesunięcia semantyczne. Ta różnica warunkuje, że wyraz prawda w większym stopniu staje się obiektem procesów gramatykalizacji, tracąc swoje znaczenie podstawowe, choć wyraz istina nie traci autonomii semantycznej i posiada swego rodzaju immunitet na gramatykalizację. The paper deals with two aspects of the semantics and the categorial potential of language entities encoding various concepts of TRuTH in Russian in comparison with several other I.-E. languages. The first one is the conceptual origins of both Russian key nouns in question, namely правда denoting, among others, true sentences, right solutions, an honest behavior, etc., and истина referring to the truth in its scientific, philosophical, religious etc. sense. It is shown that each lexical item is based on different source concepts, respectively NORM and ESSENCE, so that the two different target concepts of the ‘truth’ are to a huge degree influenced by them. The second approach is devoted to the grammaticalization potential of Russian and other lexical entities with “veritative” semantics. Since the source concept of NORM (as well as, in other languages, one of an observed REALITY) is non-specific and widely poly-semantic, it is per se predestinated to semantic change including the loss of semantic autonomy and grammaticalization. On the contrary, the ESSENCE concept prohibits or at least strongly restricts both semantic change and grammaticalization of the items based on it.
Linguistic competences are of foremost importance to students in ESP (English for Specific Purposes) programs where they receive training in areas such as vocabulary, sentence structure and rhetoric norms common to science engineering fields in order to function professionally in English in an increasingly international environment. To achieve this goal, students have to be nurtured in field-specific language contexts - an aim which is more focused than General English or English for Academic Purposes approaches. To this end, ESP instructors attempt to find effective methods to analyze student abilities and tailor materials suited to their needs and level of proficiency. This paper describes the first part of a longitudinal project which aims to improve vocabulary development of third-year science engineering students in a Japanese university. In this pilot study, one third of the classes of undergraduate Technical English were given the vocabulary size test developed by Paul Nation and David Beglar (2007) to gauge students' approximate vocabulary size for general English reading. The 30-minute test was administered in class via the university's e-learning system (WebClass UEC) to expedite the compiling of results. This presentation reports the background and rationale for using this particular measurement, the holistic results of the pilot assessment, the analysis of the correlations and deviations between departments, and the implications of these similarities and differences. In addition, these results will be used for deciding on the level from which to develop teaching materials to bridge the gaps that may appear in students' semi-technical and technical lexical repertoire. In future studies, these results will be correlated with TOEIC scores of the same students, as well as the results of other academic and specialized technical word lists.
BSL SignBank is an online, usage-based dictionary of British Sign Language, based on signs from the BSL Corpus (http://www.bslcorpusproject.org).
The purpose of our work is to explore the possibility of using sentence diagrams produced by schoolchildren as training data for automatic syntactic analysis. We have implemented a sentence diagram editor that schoolchildren can use to practice morphology and syntax. We collect their diagrams, combine them into a single diagram for each sentence and transform them into a form suitable for training a particular syntactic parser. In this study, the object language is Czech, where sentence diagrams are part of elementary school curriculum, and the target format is the annotation scheme of the Prague Dependency Treebank. We mainly focus on the evaluation of individual diagrams and on their combination into a merged better version.
The present dissertation is a comparative study of the intonation of yes-no questions in Hungarian and Spanish. Based especially on my own corpora, I examine the realization of the main accent in utterances, pitch range, and the intonational patterns applied. First, these aspects will be investigated in a Spanish corpus (Corpus 1) then in a Hungarian corpus (Corpus 2) and after that, I will make hypotheses about the ways Hungarians pronounce Spanish yes-no questions. These predictions then will be validated by means of a corpus containing Spanish yes-no questions produced by Hungarian learners of Spanish (Corpus 3). My predictions were the following: (a) As the place of main accent in an utterance depends on lexical stress, and lexical stress placement obeys different rules in the two languages, it is predictable that Hungarian learners of Spanish will not produce Spanish main accents according to the Spanish norms. (b) Hungarian uses a narrower pitch range than Spanish, thus, the Spanish yes-no interrogatives produced by Hungarian learners are expected to have a narrower pitch range. (c) The intonation contours applied will be investigated in 3 subgroups of yes-no questions: ordinary yes-no questions, echo yes-no questions and yes-no questions followed by a vocative. Ordinary yes-no questions in Hungarian are typically accompanied by rising-falling contours, whereas in Spanish, by rising ones; Hungarian echo yes-no questions have several main accents, each triggering a rise-fall contour, while in their Spanish counterparts there is one main accent in these cases, with a characteristically rising pattern. Yes-no question + vocative sequences contain two intonation units in both languages, but in Hungarian the yes-no interrogative conserves its rising-falling melody, and the vocative is accompanied by a fall, unlike in Spanish, where both contours are rising, and the final vocative is given the higher rise. Based on these observations, the prediction is that Hungarians will transfer their Hungarian intonational patterns to Spanish yes-no questions, which may be found unacceptable by Spanish listeners. My hypotheses will be validated by the analysis of the Spanish yes-no interrogatives of Hungarian students, which will cast light on those areas of intonation which should be given more attention in Spanish language teaching in Hungary.
Syllable structure constitutes the component of phonological word division focused on pronounceable segments of words and how they are composed, divided, and distributed. Syllable structure is also a subdivision of the study of phonotactics, or the rules of sound distribution, the specific sequences of sound that occur in a language. And, third, the study of syllables in Arabic involves the analysis of lexical stress. Although syllables themselves are linear and segmental in nature, word stress (the loudness or emphasis placed on a syllable) is suprasegmental; that is, it occurs at the same time as the pronunciation of the segment, adding a dimension of complexity to the syllable itself. MSA has explicit structural restrictions on syllables, as well as predictable rule-based stress based on syllable strength. Although not a spontaneous spoken register of Arabic, MSA is nonetheless spoken on formal occasions (usually scripted) and in broadcast news and information formats, and adheres to established norms of stress placement. Recent published work on the stress system of MSA has largely been done within the theoretical framework of prosodic morphology. The discussion set forth here uses a basic descriptive approach similar to the one used in Ryding 2005 (36–39), Mitchell 1990 (19–21), and McCarus and Rammuny (1974: 7–8, 23).
We propose a new simple but effective method for building Tibetan-Chinese machine Translation corpus and a novel Tibetan-Chinese Machine Translation model integrating Tibetan syntactic cues which is based on the Treebank, this model can be used on the system of Tibetan-Chinese Machine Translation successfully. Keywords: syntactic Treebank; Tibetan syntactic cues; Machine Translation;
The article deals with the linguistic issues of composing a reference book of regional toponyms a genre that requires special consideration in national lexicography. The assortment of these issues gave the possibility to carry out complex description of regional toponyms on the basis of semantic, functional, and orpthologuos criteria that let unify the names of Volgograd region settlements that are registered in various documents. The significance of the composed reference book is determined by several factors the presence of local subsystems of geographical names in Russian toponymy; the inconsistency of current orthography norms on using capital letter in compound proprius names and fused-with-hyphen spelling of toponyms and off-toponym derivations; the lack of linguistically justified explanation of peculiarities of grammatical norms in the field of proper names use. The reference book of regional toponyms is based on the object description (toponymic vocabulary), principles of lexical units selection (description of spelling and grammatical properties of toponyms, encyclopedic information), the glossary (full list of toponyms of Volgograd region), typical article. The articles in the reference book are arranged in lexicographical zones with grammatical and semantic markers, lexicographical illustrations, other lexicographical labels, word etymology including. The reference book on Volgograd region toponymy is addressed to executive and administration authorities, journalists, regional ethnographers.
The article deals with the violations of language norms in press. The object of investigation is regional newspapers: morphological syntactical, lexical and spelling mistakes occuring on the pages of print media are analysed.
The objective of this paper is to provide an overview of the CDT annotation design with special emphasis on the modelling of the interface between the syntactic level and two other linguistic levels, viz. morphology and discourse. In connection with the description of NP annotation we present the fundamentals of how CDT is marked up with semantic relations in accordance with the dependency principles governing the annotation on the other levels of CDT. Specifically, focus will be on how Generative Lexicon (GL) theory has been incorporated into the unitary theoretical dependency framework of CDT. An annotation scheme for lexical semantics has been designed so as to account for the lexico-semantic structure of complex NPs, and the four GL qualia also appear in some of the CDT discourse relation labels as a description of parallel semantic relations at this level.
In spite of the development of content-based data management, text-based searching remains the primary means of multimedia retrieval in many areas. Automatic creation of text metadata is thus a crucial tool for increasing the findability of multimedia objects. Search-based annotation tools try to provide content-descriptive keywords by exploiting web data, which are easily available but unstructured and noisy. Such data need to be analyzed with the help of semantic resources that provide knowledge about objects and relationships in a given domain. In this paper, we focus on the task of general-purpose image annotation and present the VCO, a new ontology of visual concepts developed as a part of image annotation framework. The ontology is linked with the WordNet lexical database, so the annotation tools can easily integrate information from both these resources.
Recently, we reported on our efforts to build the first prototype of KurdNet. In this proposal, we highlight the shortcomings of the current prototype and put forward a detailed plan to transform this prototype to a full-fledged lexical database for the Kurdish language.
Language provides a finite set of labels (words) for an infinite set of possible objects that a speaker may encounter. Previous picture naming studies of language production have focused on highly familiar object stimuli that elicit uniform responses among native-speaker participants. By contrast, this dissertation explores picture naming responses to a broad range of object stimuli, from typical prototypes to unusual or unfamiliar examples of the same name. Drawing on research in lexical categorization, I document variation in native-speaker picture naming behavior and examine how this naming behavior and the underlying neural responses change as a result of language interaction in bilinguals. Behavioral studies (Chapter 2) measure differences in lexical categorization patterns (picture naming responses over many different examples of an object name) among monolingual and bilingual speakers of Chinese and English. Further, this inter-personal variation is explained in terms of speakers' unique language histories and norms of each linguistic community (native speakers of English and Chinese). I propose a statistical model which describes the role of these variables in each language in predicting categorization patterns in both the native language (L1) and second language (L2), identifying significant effects of cross-language interaction for bilinguals in both languages. Next, I introduce a new stimulus set of 407 objects sampled from several semantic domains (e.g., clothing and vehicles) and normed by native, monolingual speakers of English and Chinese (Chapter 3). This stimulus set demonstrates the extensive variation in picture naming responses among native speakers of each language and between languages. I use these norms to examine a few specific variables relating to native speaker norms identified in the previous chapter. In the final set of experiments (Chapters 4 and 5) I select a subset of 183 objects from the new stimulus set to test functional neural correlates of these categorization variables in native, monolingual speakers of each language and in Chinese-English bilinguals. Each categorization variable is associated with brain regions that uniquely respond to its variation, and activity in these regions confirms functional involvement of both L1 and L2 variables in bilinguals' L1 picture naming behavior. These findings are situated in the broader context of neurocognitive models of language production and offer a refined view of lexical semantic retrieval and selection, accounting for the variety of potential objects that speakers may encounter and variation among native speakers of a language.
Due to globalization there is an increase in the appearances of languages in the multilingual linguistic landscape in urban spaces. Commentators have described this state of affairs as super-, mega- or complex diversity. Mainstream sociolinguists have argued that languages have no fixed boundaries and that they are "fluid" in fact. The output of speech production and language use is actually referred to as "languaging". The terms implies that languages are rather resources but not fixed tools for communication. In this paper, I will argue that this theory to which I will refer as the superdiversity/languaging theory cannot cover multilingual data in terms of resources only, if phenomena of multilingual linguistic landscape are studied more carefully. It turns out that constructions that look like "languaging" are from a linguistic point of view in fact well-known cases of code-switching (or -mixing) with separate languages involved, a dominant language and clearly targeted messages for the speakers of the "underlying" language. Hence, I will conclude that linguistic data in multilingual urban spaces are not necessarily arranged in terms of resources but rather in terms of Fishmanian diglossia, triglossia, and so on. This implies that even in these cases of languaging there is no reason to operate with concepts of language other than recognizable languages that are characterized by a prototypical grammatical and lexical basic core. Hence, languages in this sense and not code-switched variants, like "English as a Lingua Franca" feed into strategies of transnational communication, although the output of transnational communication can<br/>be a code-switched variant of English as well. However, I agree with the proponents of the<br/>superdiversity/languaging theory that it is highly relevant to study the proliferation of all sorts of multilingualism in the context of complex linguistic diversity. This reveals not only the structures and rules of language and languages that I will define as linguistic categories in accordance with Chomskyan grammar but also provides insight into the quickly changing semantic and world view concepts due to globalization. However, the code-switched variants appearing in multilingual complex spaces are not suitable for linguistic diversity management that includes institutions. Institutions are by definition the outcome of norm-based governance strategies and will implement norm-based entities, like languages that are recognizable and make possible contextualized, sophisticated language use. This rules out highly individual, spontaneous production of language, like languaging-phenomena.
The last decade has seen an explosion in the number of people learning English as a second language (ESL). In China alone, it is estimated to be over 300 million (Yang in Engl Today 22, 2006). Even in predominantly English-speaking countries, the proportion of non-native speakers can be very substantial. For example, the US National Center for Educational Statistics reported that nearly 10 % of the students in the US public school population speak a language other than English and have limited English proficiency (National Center for Educational Statistics (NCES) in Public school student counts, staff, and graduate counts by state: school year 2000–2001, 2002). As a result, the last few years have seen a rapid increase in the development of NLP tools to detect and correct grammatical errors so that appropriate feedback can be given to ESL writers, a large and growing segment of the world’s population. As a byproduct of this surge in interest, there have been many NLP research papers on the topic, a Synthesis Series book (Leacock et al. in Automated grammatical error detection for language learners. Synthesis lectures on human language technologies. Morgan Claypool, Waterloo 2010), a recurring workshop (Tetreault et al. in Proceedings of the NAACL workshop on innovative use of NLP for building educational applications (BEA), 2012), and a shared task competition (Dale et al. in Proceedings of the seventh workshop on building educational applications using NLP (BEA), pp 54–62, 2012; Dale and Kilgarriff in Proceedings of the European workshop on natural language generation (ENLG), pp 242–249, 2011). Despite this growing body of work, several issues affecting the annotation for and evaluation of ESL error detection systems have received little attention. In this paper, we describe these issues in detail and present our research on alleviating their effects.
Ferdydurke by Witold Gombrowicz: reflections on the translation of proper names. The aim of the paper is to analyse selected proper names in Italian translations of the first novel written by Witold Gombrowicz (1904-1969), a Polish writer, entitled Ferdydurke: (1) translated by S. Miniussi in 1961 from a French version of the novel, and (2) by V. Verdiani in 1991 from the original Polish. The attempt at interpretation involves anthroponyms referring to the characters in the novel; these have been divided into two groups depending on the translation strategy employed: translation or foreignization. The translated names in both versions comprise only five lexical items: the name of the main character – J.zio (Momo /Gingio), the surname of the middle-class family – Mlodziak (Giovincelli / Giovanotti), the nickname of professor Bladaczka (Stecchino / Pallore), the nickname of one of the students – Syfon (Sifone) and the dog’s name, marked with dialectal features – Burecek (Bubi / Medoro). Among the translated anthroponyms, the diminutives of the characters’ names, which are characteristic of the original version, are particularly interesting. As for untranslated names, there are both original names without any changes and names that have been adjusted to Italian phonetic norms with the elimination of the Polish diacritics.
This article aims at a more satisfying explanation of differential object case marking (DOM), and demonstrates that a group of mass nouns displays properties that are preserved in derivation. The central tenet of all accounts relates the Finnish type accusative-partitive DOM to the distinction between mass and count noun objects. I challenge this established view by introducing new data from Estonian: deadjectival mass nouns that unexpectedly behave like count nouns in DOM. I propose an account that has a wider coverage of data and is based on the scalar and boundedness-related properties of the base adjectives of the derived abstract nouns. Typically, the unexpected count-like behavior occurs with abstract nouns that are derived from adjectives that cannot denote open scales for various lexical-semantic and pragmatic reasons. Since the semantic properties of scales as well as the pragmatic standards determining boundedness are preserved in the course of derivation, they are cross-categorial properties. These findings are also relevant in understanding of the role of lexical aspect and aspectual composition as well as the links between morphosyntax in language and norms and standards in cognition.
Assuming that collaboration between theoretical and computational linguistics is essential in projects aimed at developing language resources like annotated corpora, this paper presents the first steps of the semantic annotation of the Index Thomisticus Treebank, a dependency-based treebank of Medieval Latin. The semantic layer of annotation of the treebank is detailed and the theoretical framework supporting the annotation style is explained and motivated.
A recent dramatic increase in the number and scope of chronometric and norming lexical megastudies offers the ability to conduct virtual experiments-that is, to draw samples of items with properties that vary in critical linguistic dimensions. This paper introduces a bootstrapping approach, which enables testing of research hypotheses against a range of samples selected in a uniform, principled manner and evaluates how likely a theoretically motivated pattern is in a broad distribution of possible outcome patterns. We apply this approach to conflicting theoretical and empirical accounts of the relationship between the psychological valence (positivity) of a word and its speed of recognition. To this end, we conduct three sets of multiple virtual experiments with a factorial and a regression design, drawing data from two lexical decision megastudies. We discuss the influence that criteria for stimuli selection, statistical power, collinearity, and the choice of dataset have on the efficacy and outcomes of the bootstrapping procedure.
This paper introduces a pilot study on incorporating effective teaching methods in computer-aided pronunciation training (CAPT) programs to help English-speaking learners acquire Mandarin lexical tones by using speech analysis software. It is proved that CAPT programs help learners identify relevant acoustic cues and discern the four tones in Chinese language, which they found hard to differentiate and imitate. Acoustic analyses of the pitch track comparisons between pre- and post-training productions in the form of visual display of speakers' pitch curves (tracks) reveal the nature of the improving process for the learner. Acoustic images also indicate that post-training tone curves (tracks) approximate native norms to a greater degree than pre-training tone tracks. The methodology developed hereby may provide a platform for more efficient Chinese learning.
SECOND LANGUAGE LEARNERS FACE A DUAL CHALLENGE IN VOCABULARY LEARNING: First, they must learn new names for the 100s of common objects that they encounter every day. Second, after some time, they discover that these names do not generalize according to the same rules used in their first language. Lexical categories frequently differ between languages (Malt et al., 1999), and successful language learning requires that bilinguals learn not just new words but new patterns for labeling objects. In the present study, Chinese learners of English with varying language histories and resident in two different language settings (Beijing, China and State College, PA, USA) named 67 photographs of common serving dishes (e.g., cups, plates, and bowls) in both Chinese and English. Participants' response patterns were quantified in terms of similarity to the responses of functionally monolingual native speakers of Chinese and English and showed the cross-language convergence previously observed in simultaneous bilinguals (Ameel et al., 2005). For English, bilinguals' names for each individual stimulus were also compared to the dominant name generated by the native speakers for the object. Using two statistical models, we disentangle the effects of several highly interactive variables from bilinguals' language histories and the naming norms of the native speaker community to predict inter-personal and inter-item variation in L2 (English) native-likeness. We find only a modest age of earliest exposure effect on L2 category native-likeness, but importantly, we find that classroom instruction in L2 negatively impacts L2 category native-likeness, even after significant immersion experience. We also identify a significant role of both L1 and L2 norms in bilinguals' L2 picture naming responses.
We present HamleDT 2.0 (HArmonized Multi-LanguagE Dependency Treebank). HamleDT 2.0 is a collection of 30 existing treebanks harmonized into a common annotation style, the Prague Dependencies, and further transformed into Stanford Dependencies, a treebank annotation style that became popular recently.\n\nWe use the newest basic Universal Stanford Dependencies, without added language-specific subtypes. We describe both of the annotation styles, including adjustments that were necessary to make, and provide details about the conversion process. We also discuss the differences between the two styles, evaluating their advantages and disadvantages, and note the effects of the differences on the conversion.\n\nWe regard the stanfordization as generally successful, although we admit several shortcomings, especially in the distinction between direct and indirect objects, that have to be addressed in future.\n\nWe release part of HamleDT 2.0 freely; we are not allowed to redistribute the whole dataset, but we do provide the conversion pipeline.
Dans le but de présenté notre projet de licence nous avons réalisé un moteur de recherche sémantique afin de trouver les documents pertinent pour l’utilisateur pour cella on a utilisé la base de donnée lexical WordNet Comme principe afin d’apporter les sacs de mots. L’implémentation et la conception sont faites à l’aide du langage java en utilisant l’IDE NetBeans6.8. Abstract In order to present our project of licence we conducted a semantic search engine to find the relevant documents for the user that’s why we used the lexical database WordNet as a principle to provide bags of words. The implementation and design are made with the java language using the IDE NetBeans6.8.
See http://clld.org
This paper presents the ITU Turkish Dependency Validation Set firstly introduced in 2007 [36] in order to serve as the test set of the CoNLL-XI shared task (shared task of the Conference on Computational Natural Language Learning 2007 [28] ). The dataset is available from http://web.itu.edu.tr/gulsenc/treebanks.html and is used by several academic studies so far.
To solve the problem of lower precision caused by traditional query expansion technology, a new query expansion technique based on semantic context was proposed. The semantic context is constructed by WordNet knowledge base and related feedback documents. Firstly, the query words senses are confirmed by disambiguation with WordNet lexical database. Secondly, the initial expansion words are obtained according to the WordNet semantic hierarchy structure. Finally, the weight of the expansion terms is determined according to the overall correlation of candidate expansion terms and all the query words. These words whose weight is higher than weight threshold will be chosen as the final query expansion word. The experimental results show that the proposed method obviously improves the retrieval precision while preserving higher recall.
We describe our efforts to scale up a syntactic search engine from a 1 million word treebank of written Dutch text to a treebank of 500 million words, without increasing the query time by a factor of 500. This is not a trivial task. We have adapted the architecture of the database in order to allow querying the syntactic annotation layer of the SoNaR corpus in reasonable time. We reduce the search space by splitting the data in many small databases, which each link similar syntactic patterns with sentence identifiers. By knowing on which databases we have to apply the XPath query we aim to reduce the query times.
CARMA is a media annotation program that collects continuous ratings while displaying audio and video files. It is designed to be highly user-friendly and easily customizable. Based on Gottman and Levenson's affect rating dial, CARMA enables researchers and study participants to provide moment-by-moment ratings of multimedia files using a computer mouse or keyboard. The rating scale can be configured on a number of parameters including the labels for its upper and lower bounds, its numerical range, and its visual representation. Annotations can be displayed alongside the multimedia file and saved for easy import into statistical analysis software. CARMA provides a tool for researchers in affective computing, human-computer interaction, and the social sciences who need to capture the unfolding of subjective experience and observable behavior over time.
We present the Uppsala Persian Dependency Treebank (UPDT) with a syntactic annotation scheme based on Stanford Typed Dependencies. The treebank consists of 6,000 sentences and 151,671 tokens with an average sentence length of 25 words. The data is from different genres, including newspaper articles and fiction, as well as technical descriptions and texts about culture and art, taken from the open source Uppsala Persian Corpus (UPC). The syntactic annotation scheme is extended for Persian to include all syntactic relations that could not be covered by the primary scheme developed for English. In addition, we present open source tools for automatic analysis of Persian containing a text normalizer, a sentence segmenter and tokenizer, a part-of-speech tagger, and a parser. The treebank and the parser have been developed simultaneously in a bootstrapping procedure. The result of a parsing experiment shows an overall labeled attachment score of 82.05% and an unlabeled attachment score of 85.29%. The treebank is freely available as an open source resource.