Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Treebanks have become crucial for the development of data-driven approaches to natural language processing, human language technologies, grammar extraction, and linguistic research in general. Manifold projects aim at compiling representative treebanks for specific languages. Other projects focus on the development of tools for exploration of annotated treebanks, or explore annotation beyond syntactic structure and beyond single languages. The Seventh International Workshop on Treebanks and Linguistic Theories (TLT7) provides a forum for researchers in the field of Computational Linguistics who are experts in the design, creation and exploitation of treebanks and their relation to linguistic theories. A selection of 16 workshop papers is published in these proceedings. Together, they cover a wide range of topics, including the building, querying, exploring, exploiting and evaluating of treebanks.
This presentation will describe Floresta Sintactica, a syntactic Treebank for Portuguese. Some new linguistic features will be presented, as well as examples of how Floresta can be used to explore aspects of Portuguese syntax. The new interface of Floresta, Milhafre, work in progress, will also be shown.
In the paper we describe the results in the development of the tools for handling lexical resources of Czech and possibly other languages. The tools are based on the client/server DEBII platform, they are language independent and use standard XML formats. In this direction we strive to a standardization of the lexical resources and also their interoperability. The goal is to have a collection of the Web based tools for browsing and editing machine readable dictionaries, semantic networks (WordNets),ontologies and lexical databases in general. In particular, we pay attention to the following tools: DEBDict, DEBTerm, PRALED and Visual Browser.
In sound perception the focus often lies in the cognitive aspect of the sound. We argue that the emotional aspect has to be added to get a fuller picture of sound perception. By using emotions as parameter in design of auditory alerts, one can reach a more accurate reaction to the alert. In this paper we studied the emotional connection to some attributes, common in music psychology, that are possible to describe by simple parameters. Short stimuli were created from these parameters in a factorial test design. The sounds were presented over headphones, with same signal fed to both ears, to 30 participants. The participants were asked to rate level of valence and activation, using a pictorial scale (SAM). Statistical differences was mostly found in ratings of activation, but differences were also shown in valence ratings. Results will be discussed in relation to theories of sound perception as well as music psychology.
878 Reviews focus the subject as clearly as possible, the results are unpredictable and heterogene ity isunavoidable. That notwithstanding, thevolume fulfils its simply stated aims of seeing theproblems and possible solutions better incontext. KING'S COLLEGE LONDON D. N. YEANDLE Sprachnormenwandel imgeschriebenenDeutsch an der Schwelle zum 2I. Jahrhundert. By VIT DOVALIL. (Duisburger Arbeiten zur Sprach- und Kulturwissenschaft, 63) Frankfurt a.M.: Peter Lang. 2006. vi+236 pp.?42.50. ISBN 978-3-631 53425-0. The questions surrounding linguistic norms, standardization, and language change have recently become a major focus of research interest, particularly in respect of German, and thiscontribution by a young Czech scholar is a very timelycontribution to thedebates on these issues.Other studies have tended toconcentrate on the spoken language, examining the degree towhich spoken norms may deviate from standard codifications, but thiswork centres on thewritten language, specifically the language of the supraregional press inGermany, and examines the extent towhich thenorms commonly taken to be standard are actually adhered to in practice and accepted by those usually regarded as language authorities. After a short account of the aims of thework in the firstchapter, the author presents inChapter 2 an extensive attempt to define thebasic termswhich are relevant tohis study, i.e. 'norms', 'language variety', 'standard language', 'nonstandard/substandard language', and ' Umgangssprache'. For each of these he takes thedefinitions given in earlier work, summarizes and classifies them,and tries toarrive at a defensible definitionwhich will serve as a basis forhis own study.This isa useful presentation in itsown rightgiven the range ofoften very varied and sometimes contradictory definitions which have been attempted in the literature over the past century (forty-nine are listed here for 'norms'), and if the author does not always finda sureway through theseminefields (very fewwill be capable of this), theattempt is immenselyworthwhile and provides a helpful set of references for those wishing to inform themselves of thedebates surrounding these terms. The main body of the book (Chapters 3 to 7) contains the empirical investigation and a thorough analysis of the findings.The author's method was firstto select ten variable features (e.g. theuse of preposition plus pronoun rather than thepronominal adverb-i.e. zu was? or wozu?-or the use of the genitive or dative case with statt, wdhrend, and wegen), and establish whether the formsconventionally regarded as non standard are used in thenational press. Using principally the corpora of the Institut fur Deutsche Sprache in Mannheim, he finds this is so forall thevariables he selected, with, forexample over a hundred attestations for theuse of preposition plus pronoun where thenorm requires thepronominal adverb. The sheer volume of thedata found and presented here is quite remarkable. In a second stage he conducted an investi gation bymeans of a questionnaire to fifty-three professors ofGerman linguistics in Germany, asking themwhether they considered the specified forms to be standard or non-standard, and acceptable inwriting. The result of this survey is ifanything even more astonishing, in that a significant number of these variants (notably, by a huge majority, the past subjunctive form brduchte, the past participle gewunken, or the use of brauchenwithout a following zu) were considered to be standard German and wholly acceptable by thisgroup of informants. This latter result isparticularly interesting because itparallels findings fromother recent researchwhich has shown thatGerman schoolteachers and German Lektoren in theUK may be quite uncertain of theprescriptive codification incases where usage varies. However, rather than seeing thisas 'destandardization', as theauthor is inclined MLR, I03.3, 20o8 879 to, it isperhaps the case that even inGerman, where formallyprescribed norms have longbeen unchallenged and therehas been a somewhat defensive attitude towards the codified standard, usage norms are beginning to establish themselves as acceptable competitors. But whatever the interpretation which one might wish to put on the development, the author deserves much credit for such enlightening documentation ofwhat iscurrently taking place inGerman. UNIVERSITY OF MANCHESTER MARTIN DURRELL VonMythen undMdren: Mittelalterliche Kulturgeschichte imSpiegel einerWissen schaftler-Biographie. Festschriftfiir Otfrid Ehrismann zum 65. Geburtstag. Ed. by GUDRUN MARCI-BOEHNCKE and JORGRIECKE. Hildesheim: OIms. 2oo6. 68i pp.?84. ISBN 978-3-487-I3179-5 It isa lovely custom togive senior academics a special birthday present in the formof a book towhich colleagues and formerstudents have contributed...
Functional Arabic Morphology is a formulation of the Arabic inflectional system seeking the working interface between morphology and syntax. ElixirFM is its high-level implementation that reuses and extends the Functional Morphology library for Haskell. Inflection and derivation are modeled in terms of paradigms, grammatical categories, lexemes and word classes. The computation of analysis or generation is conceptually distinguished from the general-purpose linguistic model. The lexicon of ElixirFM is designed with respect to abstraction, yet is no more complicated than printed dictionaries. It is derived from the open-source Buckwalter lexicon and is enhanced with information sourcing from the syntactic annotations of the Prague Arabic Dependency Treebank. MorphoTrees is the idea of building effective and intuitive hierarchies over the information provided by computational morphological systems. MorphoTrees are implemented for Arabic as an extension to the TrEd annotation environment based on Perl. Encode Arabic libraries for Haskell and Perl serve for processing the non-trivial and multi-purpose ArabTEX notation that encodes Arabic orthographies and phonetic transcriptions in parallel.
Ideas concerning several areas are still fruitful: (1) The concept of markedness as underlying the difference between the centre of the language system, patterned in a relatively simple way, and its vast and complex periphery. (2) Dependency (valency) based syntax, tested in detail and enriched in the Prague Dependency Treebank. (3) The description of the topic-focus articulation, connected with the scope of negation, presupposition and allegation; the fundamental nature of the articulation and a relatively perspicuous description of its interplay with syntactic dependency within the underlying sentence structure. (4) In linguistic typology, a single basic structural property of every type: the manner of expression of grammatical values is favourable to the other properties. (5) The stratification of Czech as a national language, with colloquial speech exhibiting an oscillation of forms of the “literary” norm and of Common Czech. This attitude is only slowly finding its way into school education. (6) In the field of word formation, attention is focused on the transition zone towards morphemics. (7) In the study of the process of communication, the degrees of activation in the stock of shared knowledge are especially significant.
RESUMO: Este trabalho, fruto de pesquisa de cunho etnografico e colaborativo na area da LinguisticaAplicada, descreve algumas representacoes sobre o “erro” no processo de aprendizagem da escrita numaclasse de alfabetizacao de jovens e adultos em uma escola da rede municipal de Vitoria da Conquista, Bahia.Tais representacoes, como se percebeu na analise, simbolizam, para os atores sociais envolvidos(alfabetizadora e alfabetizandos), a (re) producao de uma imagem de “escrita correta”, sob a qual se formauma rede de sentidos em torno de uma “lingua escrita correta” nos eventos de letramento escolar.Palavras chave: erro; letramento; norma linguistica.ABSTRACT: This paper is based on an ethnographic and collaborative research conducted in the area ofApplied Linguistics. It describes some representations of “error” in the writing learning process of theyoung and adult literacy students at a school in Vitoria da Conquista, Bahia. Data results show that therepresentations for both the teacher and the young and adult learners are reproductions of a “correctwriting” image produced in the literacy events of the classroom.Keywords: error; literacy; linguistic norm.
In this paper, we give a description of the machine translation (MT) system developed at DCU that was used for our third participation in the evaluation campaign of the International Workshop on Spoken Language Translation (IWSLT 2008). In this participation, we focus on various techniques for word and phrase alignment to improve system quality. Specifically, we try out our word packing and syntax-enhanced word alignment techniques for the Chinese–English task and for the English–Chinese task for the first time. For all translation tasks except Arabic–English, we exploit linguistically motivated bilingual phrase pairs extracted from parallel treebanks. We smooth our translation tables with out-of-domain word translations for the Arabic–English and Chinese–English tasks in order to solve the problem of the high number of out of vocabulary items. We also carried out experiments combining both in-domain and out-of-domain data to improve system performance and, finally, we deploy a majority voting procedure combining a language modelbased method and a translation-based method for case and punctuation restoration. We participated in all the translation tasks and translated both the single-best ASR hypotheses and the correct recognition results. The translation results confirm that our new word and phrase alignment techniques are often helpful in improving translation quality, and the data combination method we proposed can significantly improve system performance.
Studies examining factors that influence when words are learned typically investigate one lexical category or a small set of words. We provide the first evaluation of the relation between input frequency and age of acquisition for a large sample of words. The MacArthur-Bates Communicative Development Inventory provides norming data on age of acquisition for 562 individual words collected from the parents of children aged 0; 8 to 2; 6. The CHILDES database provides estimates of frequency with which parents use these words with their children (age: 0; 7-7; 5; mean age: 36 months). For production, across all words higher parental frequency is associated with later acquisition. Within lexical categories, however, higher frequency is related to earlier acquisition. For comprehension, parental frequency correlates significantly with the age of acquisition only for common nouns. Frequency effects change with development. Thus, frequency impacts vocabulary acquisition in a complex interaction with category, modality and developmental stage.
\n Dans cette proposition, nous plaidons pour une meilleure mutualisation des résultats de recherche sur le lexique à travers loutil informatique que représente le Web. Après avoir analysé les conditions de réussite dune telle mutualisation et limportance des normes et standards en ce domaine, nous montrons quelques exemples de réussite dune telle mutualisation tant en lexicographie contemporaine, à travers le Trésor de la langue française informatisé, quen lexicographie historique. Ainsi à travers le DMF (Dictionnaire du Moyen Français), nous explicitons le concept nouveau de lexicographie évolutive et montrons quelques exemples de résultats de recherche qui nauraient pas vu le jour sans sappuyer sur la richesse dexploitation inégalée, rendue possible grâce à son informatisation: \n - en lexicologie, par exemple sur la datation dapparition de sens nouveaux dun lexème dans la langue,\n - en pragmatique, à travers lexemple dune anté-datation de près de deux siècles de lusage de enfin énumératif,\n - ou en morphologie constructionnelle, à travers létude des formations en inr- qui pour certaines furent ensuite abandonnées au profit de formation en irr- (tel inrégulier versus irrégulier).\nToujours dans le domaine de la lexicographie historique, nous montrons lintérêt de mutualiser nos connaissances sur létymologie, tel quil se pratique dans le projet TLF-Etym, ou sur les « mots fantômes », pseudo lexèmes disposant à tort dun statut lexicographique (« ces mots qui nexistent pas »), et les lemmatisations erronées qui se trouvent encore trop souvent dans les dictionnaires historiques et étymologiques français de référence.\nNous terminons enfin par la présentation dun exemple dintégration et de valorisation de données lexicographiques et lexicales au sein du portail lexical du Centre National de Ressources Textuelles et Lexicales (CNRTL, www.cnrtl.fr) qui à travers les quelque 300 000 requêtes quil sert par jour est aujourdhui une magnifique vitrine des résultats de recherche en lexicographie, morpho-syntaxe, étymologie, synonymie, antonymie.\n\n
The studies on lexicographic definitions connected with the French tradition take charge eminently of typology and leave aside the question of metalanguage. So, in lexicography, the metalinguistic definition is often considered in the typological frame. This is because the above-mentioned studies are mostly based upon definitions either of nouns or verbs. In my presentation I shall attempt to demonstrate, from defining statements of the syncategorematic words drawn from the Tresor de la langue francaise, that the metalinguistic definition is indeed a category of the definitions but that, when compared to the other categories, it requires a different criteria of analysi, due to its nature. In order to do this, I shall present, first, the different nature of this issue from a typological approach on one side and a metalinguistic approach on the other. I shall expose, then, the main typological studies-in particular the unpublished document which is stored in the archives of the Laboratory ATILF [.Pour un nouveau cahier de normes...., 1979] as well as Martin (1983) and Rey-Debove (1998)-in which the question of the metalanguage is dealt with inside and following the example of typology to demonstrate that, if a definition such as aiguillette-nom populaire de l.orphie-is metalinguistic and a definition such as chaise - siege a dossier sans bras - is perifrastic, nom et siege are both hyperonyms, so that the typological criteria are not enough to distinguish between mealinguistic and perifrastic definition. Thus, I will establish, in accordance with Rey-Debove (1997), in which the definition is considered from a metalinguistic point of view-according to the sintactic relation between a lexical entry and its lexicographical definition, the principles which govern the metalinguistic analysis. The results will lead to three different categories of metalinguistic definitions of the syncategorematic words: 1. the definition refers to both infralinguistic and extralinguistic reality-in this case two sub-categories are possible: a) the hyperonym refers to the infralinguistic reality while the specific semes refer to the extralinguistic reality; b) the hyperonym refers to the infralinguistic reality while the specific semes, among which there is at least an autonym with 'schize' (cf. Rey-Debove 1997: 116-118), refer to the extralinguistic reality; 2. the definition refers to the only infralinguistic reality; 3. the definition refers to the only extralinguistic reality.
Abstract This paper offers a model to explain the general observation that lexical items are more often borrowed from a higher status language into a lower status one, than visa versa. Material from Lahore, Pakistan, shows that in casual speech among plurilinguals codeswitching is the norm. In formal contexts, in which there is attention to proper language, educated speakers filter out features which are not part of the standard language. Constraints on language and education in the hierarchical social structure withhold from most speakers of the lower status languages the knowledge necessary to evaluate their own speech in this way, thus allowing features of other languages to become established in their language.
We present the STYX system, which is designed as an electronic corpus-based exercise book of Czech morphology and syntax with sentences directly selected from the Prague Dependency Treebank, the largest annotated corpus of the Czech language. The exercise book offers complex sentence processing with respect to both morphological and syntactic phenomena, i. e. the exercises allow students of basic and secondary schools to practice classifying parts of speech and particular morphological categories of words and in the parsing of sentences and classifying the syntactic functions of words. The corpus-based exercise book presents a novel usage of annotated corpora outside their original context.
The lexical information of verbal lexemes, such as verbs and adjectives, plays an important role in syntactic parsing, because the structure of a sentence mainly hinges on the type of verbal lexemes. The question we address in this research is how to acquire the argument structure (henceforth ARG-ST) of verbal lexemes in Korean. It is well known that manual build-up of type hierarchy usually cost too much time and resources, so an alternative method, namely automatic collection of relevant information is much more preferred. This paper proposes a procedure to automatically collect ARG-ST of Korean verbal lexemes from a Korean Treebank. Specifically, the system we develop in this paper first extracts lexical information of ARG-ST of verbal lexemes from a 0.8 million graphic word Korean Treebank in an unsupervised way, checks the hierarchical relationship among them, and builds up the type hierarchy automatically. The result is written in an HPSG-style annotation, thus making it possible to readily implement the result in an HPSG-based parser for Korean. Finally, the result is evaluated with reference to two Korean dictionaries and also with respect to a manually constructed type hierarchy.
Organizations working in a multilingual environment demand multilingual ontologies. To solve this problem we propose LabelTranslator, a system that automatically localizes ontologies. Ontology localization consists of adapting an ontology to a concrete language and cultural community. LabelTranslator takes as input an ontology whose labels are described in a source natural language and obtains the most probable translation into a target natural language of each ontology label. Our main contribution is the automatization of this process which reduces human efforts to localize an ontology manually. First, our system uses a translation service which obtains automatic translations of each ontology label (name of an ontology term) from/into English, German, or Spanish by consulting different linguistic resources such as lexical databases, bilingual dictionaries, and terminologies. Second, a ranking method is used to sort each ontology label according to similarity with its lexical and semantic context. The experiments performed in order to evaluate the quality of translation show that our approach is a good approximation to automatically enrich an ontology with multilingual information.
Sometreebanks, such as German TIGER/NeGra, represent discontinuous elements directly, i.e. trees contain crossing edges, but the context-free grammars that are extracted from them, fail to make any use of this information. In this paper, we present amethod for extracting mildly context-sensitive grammars, i.e. simple range concatenation grammars (RCGs), from such treebanks. A measure for the degree of a treebank’s mild contextsensitivity is presented and compared to similar measures used in non-projective dependency parsing. Our work is also compared to discontinuous phrase structure grammar (DPSG).
We previously observed robust activation in the hippocampal region in response to novel valenced stimuli during an fMRI recognition memory paradigm in healthy individuals across the lifespan. In this study, we compared activation during the memory task in elderly controls (EC) and individuals with mild cognitive impairment (MCI). In 23 right-handed participants (15 EC, 8 MCI; 16M/7F; mean age=71.6), neural activity was compared for novel and previously learned (familiar) items. During encoding, participants viewed 10 B/W pictures of baby and elderly faces with happy/sad expressions repeated in a block 6x, alternating with blocks of 10 circles. Participants indicated whether each face was happy or sad. After a 20-minute consolidation period, the 10 encoded faces were presented 4x each intermixed with 40 new faces (half happy, half sad). Participants identified “new” or previously learned (“old”) faces. Random effects group analysis was performed using SPM5 to compare activity in response to novel vs familiar faces. Outside the scanner, participants viewed 40 new faces (half sad, half happy), 10 faces from the encoding scan, and 80 novel faces from the recognition scan. They indicated whether each face was “new” or “old” and rated valence and arousal of each face. EC showed significant activity in right fusiform and hippocampal regions in response to novel faces compared to previously learned faces (MNI coordinates FF: 40, -58, -14, T=8.07; pcorrected=.0001; HC: 22, -12, -16; T=5.60, puncorrected=.048). Significant activations were also observed in occipital and frontal cortices. A similar pattern was found in response to baby faces but did not hold for elderly faces, suggesting that arousal is important. MCI did not show increased activity in response to novel versus familiar faces in predicted regions. Groups did not significantly differ in overall reaction time/accuracy during encoding, valence/arousal ratings or accuracy during the post-scan task, or on the Florida Affect Battery, suggesting MCIs' affective perception was intact. MCIs were slower and less accurate during recognition scans. Lack of encoding-associated activity in MCIs suggests that continued studies of the functional correlates of the emotional-memory enhancement effect in the early stage of Alzheimer's disease are warranted.
We present a robust parser which is trained on a treebank of ungrammatical sentences. The treebank is created automatically by modifying Penn treebank sentences so that they contain one or more syntactic errors. We evaluate an existing Penn-treebank-trained parser on the ungrammatical treebank to see how it reacts to noise in the form of grammatical errors. We re-train this parser on the training section of the ungrammatical treebank, leading to an significantly improved performance on the ungrammatical test sets. We show how a classifier can be used to prevent performance degradation on the original grammatical data.
We present the second version of the Penn Discourse Treebank, PDTB-2.0, describing its lexically-grounded annotations of discourse relations and their two abstract object arguments over the 1 million word Wall Street Journal corpus. We describe all aspects of the annotation, including (a) the argument structure of discourse relations, (b) the sense annotation of the relations, and (c) the attribution of discourse relations and each of their arguments. We list the differences between PDTB-1.0 and PDTB-2.0. We present representative statistics for several aspects of the annotation in the corpus. 1.
Objective: To carry out the native assessment of International Affective Picture System(IAPS) among Chinese older adults.Methods:Altogether 116 Chinese older adults,including 51 male and 65 female,from three communities in Dalian City,aged from 60 to 80 years,rated 60 pictures(positive:25,neutral:12,negative:23) selected from the IAPS in terms of valence,arousal and dominance with Self-Assessment Manikin(SAM).The mean affective ratings were compared to the normative ratings of USA National Institute of Mental Health(NIMH).Result: Reliability analysis indicated that the affective ratings of our sample were stable and highly internally consistent.The affective ratings of Chinese older participants were strongly correlated with the normative ratings of NIMH(r=0.92,0.54 and 0.88 respectively for valence,arousal and dominance,P0.001).But paired t test showed there were still significant differences between the two samples.Chinese aged reported relatively higher arousal and dominance than NIMH sample for all pictures [(5.33±0.93)vs.(4.83±1.25),(5.60±1.20)vs.(5.19±1.21),P0.001],but lower valence than NIMH sample[(4.99±2.28)vs.(5.28±1.85),P=0.020].Male and female Chinese older participants showed similar emotional responses to most pictures.But female Chinese older participants reported higher valence than male ones(5.05±2.33/4.93±2.24,P0.05).The 60 pictures were distributed as shape in the two-dimensional affective space(valence-arousal).The association between valence and arousal was pronounced and linear for positive pictures(r=0.71,P0.001),but unpronounced for negative pictures,(r=-0.35,P0.05).Conclusion: IAPS is highly internationally accessible just as the expectation of its designers.However,considering about great differences in many aspects such as culture,social living and age between Chinese aged and NIMH sample,they may have different affective experiences to the same emotional stimuli.Therefore it is necessary to do some revisal before the IAPS is applied to Chinese aged.
espanolEste articulo presenta un metodo de identificacion y clasificacion de la valencia y las emociones presentes en un texto. Para ello, se introduce un nuevo concepto denominado disparador de emocion. Inicialmente, se construye de forma incremental una base de datos lexica de disparadores de emocion asociados a la cultura con la que se quiere trabajar, basandose en tres teorias diferentes: la Teoria de la Relevancia de Pragmatica, la Teoria de la Motivacion de Maslow de Psicologia y la Teoria de Necesidades de Neef de Economia. La base de datos creada parte de un conjunto inicial de terminos y es ampliada con la informacion de otros recursos lexicos, como WordNet, NomLex y dominios relevantes. El enlace entre idiomas se hace por medio de EuroWordNet y se completa y adapta a diversas culturas con bases de conocimiento especificas para cada lengua. Tambien, se demuestra como la base de datos construida puede ser utilizada para buscar en textos la valencia (polaridad) y el significado afectivo. Finalmente, se evalua el metodo utilizando los datos de prueba de la tarea no 14 de Semeval Texto afectivo y su traduccion al espanol. Los resultados y las mejoras se presentan junto con una discusion en la que se tratan los puntos fuertes y debiles del metodo y las directrices para el trabajo futuro. EnglishThis paper presents a method to automatically spot and classify the valence and emotions present in written text, based on a concept we introduced-of emotion triggers. The first step consists of incrementally building a culture dependent lexical database of emotion triggers, emerging from the theory of relevance from pragmatics, Maslow's theory of human needs from psychology and Neef's theory of human needs in economics. We start from a core of terms and expand them using lexical resources such as WordNet, completed by NomLex, sense number disambiguated using the Relevant Domains concept. The mapping among languages is accomplished using EuroWordNet and the completion and projection to different cultures is done through language-specific commonsense knowledge bases. Subsequently, we show the manner in which the constructed database can be used to mine texts for valence (polarity) and affective meaning. An evaluation is performed on the Semeval Task No. 14: Affective Text test data and their corresponding translation to Spanish. The results and improvements are presented together with an argument on the strong and weak points of the method and the directions for future work.
Sciendo provides publishing services and solutions to academic and professional organizations and individual authors. We publish journals, books, conference proceedings and a variety of other publications.
In this paper, we describe our work on building a parallel treebank for a less studied and typologically dissimilar language pair, namely Swedish and Turkish. The treebank is a balanced syntactically annotated corpus containing both fiction and technical documents. In total, it consists of approximately 160,000 tokens in Swedish and 145,000 in Turkish. The texts are linguistically annotated using different layers from part of speech tags and morphological features to dependency annotation. Each layer is automatically processed by using basic language resources for the involved languages. The sentences and words are aligned, and partly manually corrected. We create the treebank by reusing and adjusting existing tools for the automatic annotation, alignment, and their correction and visualization. The treebank was developed within the project Supporting research environment for minor languages aiming at to create representative language resources for language pairs dissimilar in language structure. Therefore, efforts are put on developing a general method for formatting and annotation procedure, as well as using tools that can be applied to other language pairs easily. 1.
One problem facing the extraction of treebank grammars is that of ad hoc rules, rules used for constructions specific to one data set and unlikely to be used on new data (Dickinson, 2008). These rules can be erroneous, cover ungrammatical text, or reveal issues with the treebank’s annotation scheme.
Abstract Affective ratings of multiple religious (sub)groups (Muslims, Christians, Jews and non-believers, as well as Sunni, Alevi and Sjiit Muslims), the endorsement of Islamic minority rights and religious group identification were examined among Sunni and Alevi Turkish-Dutch participants. The findings show that both groups differ in important ways. Some Alevi participants considered themselves Muslims but others interpreted Alevi identity in a secular way. The Sunnis were quite negative towards Jews and non-believers, they more strongly endorsed Islamic minority rights and they had very high Muslim group identification. Furthermore, the Sunnis were negative towards Alevis and the Alevis were negative towards the Sunnis. Muslim group identification was positively and strongly related to feelings towards Muslims and to the endorsement of Islamic group rights.
The cycle of lexicographic and linguistic work involved in compiling a computational phraseological database is divided into three phases and described in relation to the specific challenges multi-word expressions (MWEs) pose for a lexical database. Data collection is a process that is far from complete for the MWEs found in English, with the variability of some phrases making identification of all occurrences in large corpora a major challenge. Formalization of the form and variability ofMWEs is an interrelated process which can improve tools for data collection and other applications. Increased use of the phraseological lexical database in NLP applications can ultimately lead to further insights into the nature of MWEs and to improvements in the database. Due to the volume of lexicographic data on MWEs that still needs to be collected, analysed and formalized, and the cyclical nature of the work, the resulting lexical database should be reusable in as many applications as possible. WordManager-PhraseManager, the lexical resource described in the second part of the chapter, can capture the variability ofMWEs in a way that allows for maximum reusability of lexical data.
We address corpus building situations, where complete annotations to the whole corpus is time consuming and unrealistic. Thus, annotation is done only on crucial part of sentences, or contains unresolved label ambiguities. We propose a parameter estimation method for Conditional Random Fields (CRFs), which enables us to use such incomplete annotations. We show promising results of our method as applied to two types of NLP tasks: a domain adaptation task of a Japanese word segmentation using partial annotations, and a part-of-speech tagging task using ambiguous tags in the Penn treebank corpus.
Melatonin in elderly patients Rixt F. Riemersma-van der Lek, Dick F. Swab, Jos Twisk, Elly M. Hol, Witte J.G. Hoogendijk and Eus Van Someren* Sleep and Cognition Laboratory, Netherlands Institute for Neuroscience, Amsterdam, The Netherlands (Received 10 March 2007; final version received 29 January 2008) Long-term combined light and melatonin treatment improve sleep, cognition and mood in demented elderly patients. A good night’s sleep sustains cognitive performance. Disturbed sleep, a frequent decisive factor in caregiver burden and institutionalisation of Alzheimer patients, may thus augment their characteristic impairments. We hypothesized that a possible reversible lack of activation of the circadian clock could contribute to sleep problems, and performed the first controlled human study on the effect of prolonged combined stimulation with light and melatonin. During a 3.5 year double-blind placebo-controlled randomized follow-up study, 189 elderly patients received daily supplementation of the circadian synchronisers light (+ lux, whole-day), and/or melatonin (2.5 mg). Half-yearly assessments were made of actigraphic sleep–wake rhythm estimates, cognition (MMSE), and non-cognitive symptoms. Combined light and melatonin treatment improved nocturnal restlessness by 8 + 3% per year (p 5 0.01), resulting in increased sleep duration and efficiency and a more pronounced 24-hour amplitude. Light improved cognition by 0.9 + 0.4 MMSE points or 5% (p 1⁄4 0.04). Light ameliorated depressive symptoms by 19% (Cornell Scale for Depression in Dementia). Melatonin ameliorated the worsening of psychiatric symptoms that occurred in subjects about to drop out of the study due to nursing home placement or death (Questionnaire format of the Neuropsychiatric Inventory). A negative effect of melatonin on affect (Philadelphia Geriatric Centre Affect Rating Scale) was counteracted in combination with bright light. Combined treatment also attenuated aggressive behaviour (Cohen-Mansfield Agitation Index). Both light and melatonin treatment enhanced the nocturnal rise of the 24-hour saliva melatonin rhythm, as measured in the absence of melatonin gifts. Melatonin treatment however also resulted in an increased daytime melatonin level, which was associated with an attenuation of diurnal activity. This first study on long-term stimulation of the human circadian timing system showed that improvement of the sleep–wake rhythm contributed to attenuation of cognitive The complete description of this study is available in: Riemersma-van der Lek et al. 2008. JAMA 299:2642–2655. *Corresponding author. Email: e.van.someren@nin.knaw.nl Biological Rhythm Research Vol. 40, No. 1, February 2009, 83–84 ISSN 0929-1016 print/ISSN 1744-4179 online! 2009 Taylor & Francis DOI: 10.1080/09291010802067155 http://www.informaworld.com D ow nl oa de d by [V rij e U ni ve rs ite it A m ste rd am ] a t 0 3: 36 2 0 A pr il 20 15 decline, with an affect exceeding that achieved with acetylcholinesterase inhibitors. Light improved non-cognitive symptoms, whereas melatonin may aggravate withdrawal behaviour and should – for long-term treatment in demented elderly – preferably be given in a lower dosage and in combination with bright light. 84 R.F. Riemersma-van der Lek et al. D ow nl oa de d by [V rij e U ni ve rs ite it A m ste rd am ] a t 0 3: 36 2 0 A pr il 20 15
This paper investigates transforms of split dependency grammars into unlexicalised context-free grammars annotated with hidden symbols. Our best unlexicalised grammar achieves an accuracy of 88% on the Penn Treebank data set, that represents a 50% reduction in error over previously published results on unlexicalised dependency parsing.
We describe our initial efforts towards developing a large-scale corpus of Hindi texts annotated with discourse relations. Adopting the lexically grounded approach of the Penn Discourse Treebank (PDTB), we present a preliminary analysis of discourse connectives in a small corpus. We describe how discourse connectives are represented in the sentence-level dependency annotation in Hindi, and discuss how the discourse annotation can enrich this level for research and applications. The ultimate goal of our work is to build a Hindi Discourse Relation Bank along the lines of the PDTB. Our work will also contribute to the cross-linguistic understanding of discourse connectives. 1
Conversation is one of the poetical and cultural key concepts in the writings of Mme de Stael. Departing from the contemporary understanding of the term, she outlines the conversational act as an essentially transgressive praxis that surpasses a classical canon of social, aesthetic, and linguistic norms. Highlighting a broad variety of musical, sensual, and pre-semantic aspects of language, she insists on the informal dynamics of conversation. To a certain extent, Mme de Stael claims, all conversation is improvisation, and it is within improvisation that she locates (and welcomes) the trigger of unsuspected twists and turns, the spirited, the emphatic, and even drunken quality of verbal and non-verbal exchange. The paper discusses the development of this concept in Mme de Stael's major novel, Corinne ou l'Italie. It argues that this novel is one of the first to delineate conversation as transference, shifting the focus of attention from the realm of intentionality to non-thematic qualities of communication.
Abstract Large linguistic databases, especially databases having a global coverage, such as the World Atlas of Language Structures, the Automated Similarity Judgment Program, and Ethnologue, are making it possible to systematically investigate many aspects of how languages change and compete for viability. Agent‐based computer simulations supplement such empirical data by analyzing the necessary and sufficient parameters for the current global distributions of languages or linguistic features. By combining empirical datasets with simulations and applying quantitative methods, it is now possible to address fundamental questions, such as ‘what are the relative rates of change in different parts of languages?’, ‘why are there a few large language families, many intermediate ones, and even more small ones?’, ‘do small languages change faster or slower than large ones?’, or ‘how does the borrowing of words relate to the borrowing of structural features?’
In this paper the role of concept characteristics in lexical dialectometric research is examined in three consecutive logical steps. First, a regression analysis of data taken from a large lexical database of Limburgish dialects in Belgium and The Netherlands is conducted to illustrate that concept characteristics such as concept salience, concept vagueness and negative affect contribute to the lexical heterogeneity in the dialect data. Next, it is shown that the relationship between concept characteristics and lexical heterogeneity influences the results of conventional lexical dialectometric measurements. Finally, a dialectometric procedure is proposed which downplays this undesired influence, thus making it possible to obtain a clearer picture of the ‘truly’ regional variation. More specifically, a lexical dialectometric method is proposed in which concept characteristics form the basis of a weighting schema that determines to which extent concept specific dissimilarities can contribute to the aggregate dissimilarities between locations.
Conventional n-best reranking techniques often suffer from the limited scope of the n-best list, which rules out many potentially good alternatives. We instead propose forest reranking, a method that reranks a packed forest of exponentially many parses. Since exact inference is intractable with non-local features, we present an approximate algorithm inspired by forest rescoring that makes discriminative training practical over the whole Treebank. Our final result, an F-score of 91.7, outperforms both 50-best and 100-best reranking baselines, and is better than any previously reported systems trained on the Treebank. 1
The Portal da Lingua Portuguesa is a website containing information about the Portuguese language oriented towards the general public. The largest part of the information on the Portal is lexical information concerning formal characteristics of words, such as orthography, derivations, loanwords and gentiles. The lexical information comes from a lexical database called MorDebe-or more precisely, a network of lexical databases called the Open Source Lexical Information Network (OSLIN). This abstract shows the general set-up of and major functions of MorDebe Admin, which is the lexicon management system for OSLIN. MorDebe Admin provides an easy and secure way of updating and editing the content of the different databases of OSLIN. Furthermore, much of the data on the Portal are organised as mini-dictionaries and MorDebe Admin provides an integrated collection of tools dedicated to the maintenance of these mini-dictionaries, as well as a built-in neologism tracking system. The software demonstration will illustrate these functions from a user perspective, and how easy ir is to maintain the data behind the Portal.
Morphological processes in Semitic languages deliver space-delimited words which introduce multiple, distinct, syntactic units into the structure of the input sentence. These words are in turn highly ambiguous, breaking the assumption underlying most parsers that the yield of a tree for a given sentence is known in advance. Here we propose a single joint model for performing both morphological segmentation and syntactic disambiguation which bypasses the associated circularity. Using a treebank grammar, a data-driven lexicon, and a linguistically motivated unknown-tokens handling technique our model outperforms previous pipelined, integrated or factorized systems for Hebrew morphological and syntactic processing, yielding an error reduction of 12% over the best published results so far. 1
XARA is a rule-based PropBank labeler for Alpino XML files, written in Java. I used XARA in my research on semantic role labeling in a Dutch corpus to bootstrap a dependency treebank with semantic roles. Rules in XARA are based on XPath expressions, which makes it a versatile tool that is applicable to other treebanks as well.
This paper presents a semi-automatic approach for extraction of collocations from corpora which uses the results of Conceptual Vectors as a semantic filter. First, this method estimates the ability of each co-occurrence to be a collocation, using a statistical measure based on the fact that it occurs more often than by chance. Then the results are automatically filtered (with conceptual vectors) to retain only one given semantic kind of collocations. Finally we perform a new filtering based on manually entered data. Our evaluation on monolingual and bilingual experiments shows the interest to combine automatic extraction and manual intervention to extract collocations (to fill multilingual lexical databases). It proves especially that the use of conceptual vectors to filter the candidates allows us to increase the precision noticeably.
The paper aims at the complexity of syntactic network and the feasibility that the complex network work as a means of linguistic studies.The paper proposes the method how to build a syntactic based on dependency treebank and investigates the complexity of Chinese syntactic dependency network based on two Chinese treebanks with different genres.The results show that syntactic networks have similar average path length and diameter with the random networks,but cluster coefficients of syntactic networks are much greater than that of random networks,and degree distributions of syntactic networks also obey the power law.The paper reveals that two syntactic networks have the same diameter,but with different average degree,path length,cluster coefficients and power exponent.
This paper presents the first steps towards a statistical syntactic analyzer for Basque. The system is based on a syntactically dependency annotated treebank and an adaptation of the deterministic syntactic analyzer of Nivre et al. (2007), which relies on a shift/reduce deterministic analyzer together with a machine learning module that determines which one of 4 analysis options to take, giving a unique syntactic dependency analysis of an input sentence. The results are near to those obtained by similar systems.