Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Morphology and phonology can in many cases be used to figure out which words correspond to which in Scandinavian. For instance, it is rather easy to figure out which Norwegian personal pronoun corresponds to which in Danish, and even Icelandic or Faroese. However, when it comes to prepositions and modal verbs we cannot rely on morphology or phonology alone. For example, Norwegian and Danish måtte do not always have the same meaning, and similarly, Icelandic vilja is not used as the future modal as Norwegian and Danish ville is. Instead of relying on morphology or phonology, we can use parallel corpora. Unfortunately, there are not many parallel corpora that include all of the Scandinavian languages, and those that exist are maybe not large enough to give reliable results. Nevertheless, to get a picture of what it could look like, the Danish, Faroese, Icelandic, Norwegian, Swedish, and English parts of a small treebank, The Sophie Treebank, were used to find out which modal verbs correspond to which in the various languages.
Proceedings of the 16th Nordic Conference \nof Computational Linguistics NODALIDA-2007. \nEditors: Joakim Nivre, Heiki-Jaan Kaalep, Kadri Muischnek and Mare Koit. \nUniversity of Tartu, Tartu, 2007. \nISBN 978-9985-4-0513-0 (online) \nISBN 978-9985-4-0514-7 (CD-ROM) \npp. 270-273.
The purpose of this study was to attain a deeper understanding of youth coaches’ attitudes toward the display of moral character (e.g., the values they try to teach their players, the concrete means they use to teach game rules, and prosocial norms) and to examine how they make rule abidance compatible with intensive efforts to achieve success. Semistructured interviews were conducted with 16 coaches of adolescent rugby teams. The interviews dealt with how values are taught to players and how rule following is enforced during practice and competition. A lexical analysis (Alceste software) and a thematic analysis were performed on the interview answers. The findings illustrate the complexity of the coaching role—coaches must impart a certain number of rules and ways of acting to their athletes while simultaneously inciting them to a high performance level that can lead players to go overboard in competitive situations.
07–484 Aceto, Michael (East Carolina U, USA; acetom@ecu.edu ), Statian Creole English: An English-derived language emerges in the Dutch Antilles. World Englishes (Blackwell) 25.3 & 4 (2006), 411–435. 07–485 Anchimbe, Eric A. (U Munich, Germany), World Englishes and the American tongue. English Today (Cambridge University Press) 22.4 (2006), 3–9. 07–486 Bartha, Csilla & Anna Borbély (Hungarian Academy of Sciences, Budapest, Hungary; bartha@nytud.hu ), Dimensions of linguistic otherness: Prospects of minority language maintenance in Hungary. Language Policy (Springer) 5.3 (2006), 337–365. 07–487 Coetzee-Van Rooy, Susan (North-West U, Potchefstroom, South Africa; basascvr@puk.ac.za ), Integrativeness: Untenable for world Englishes learners? World Englishes (Blackwell) 25.3 & 4 (2006), 437–450. 07–488 Gooskens, Charlotte (U Groningen, The Netherlands; c.s.gooskens@rug.nl ) & Renée van Bezooijen, Mutual comprehensibility of written Afrikaans and Dutch: Symmetrical or asymmetrical? Literary and Linguistic Computing (Oxford University Press) 21.4 (2006), 543–557. 07–489 Gooskens, Charlotte & Wilbert Heeringa (U Groningen, The Netherlands; c.s.gooskens@rug.nl ), The relative contribution of pronunciational, lexical, and prosodic differences to the perceived distances between Norwegian dialects. Literary and Linguistic Computing (Oxford University Press) 21.4 (2006), 477–492. 07–490 Guilherme, Manuela (U De Coimbra, Portgual), English as a Global language and education for cosmopolitan citizenship. Language and International Communication (Multilingual Matters) 7.1 (2007), 72–90. 07–491 Koscielecki, Marek (The Open U, Hongk Kong, China). Japanized English, its context and socio-historical background. English Today (Cambridge University Press) 22.4 (2006), 25–31. 07–492 Meilin, Chen (Three Gorges University, China) & Hu Xiaoqiong, Towards the acceptability of China English at home and abroad. English Today (Cambridge University Press) 22.4 (2006), 44–52. 07–493 Mesthrie, Rajend (U Cape Town, South Africa; raj@humanities.uct.ac.za ), World Englishes and the multilingual history of English. World Englishes (Blackwell) 25.3 & 4 (2006), 381–390. 07–494 Poole, Brian (Ministry of Manpower, Muscat, the Sultanate of Oman), Some effects of Indian English on the language as it is used in Oman. English Today (Cambridge University Press) 22.4 (2006), 21–24. 07–495 Robinson, Ian (U Calabria, Italy), Genre and loans: English words in an Italian newspaper. English Today (Cambridge University Press) 22.4 (2006), 9–20. 07–496 Ross, Kathryn (U Oxford, UK; kathryn.ross@trinity.ox.ac.uk ), Status of women in highly literate societies: The case of Kerala and Finland. Literacy (Blackwell) 40.3 (2006), 171–178. 07–497 Sala, Bonaventure M. (Cameroon), Does Cameroonian English have grammatical norms? English Today (Cambridge University Press) 22.4 (2006), 59–64. 07–498 Wei-Yu Chen, Cheryl (National Taiwan Normal U, Taiwan; wychen66@hotmail.com ), The mixing of English in magazine advertisements in Taiwan. World Englishes (Blackwell) 25.3 & 4 (2006), 467–478. 07–499 Wong, Jock (National U Singapore, Singapore; jockonn@hotmail.com ), Contextualizing aunty in Singaporean English. World Englishes (Blackwell) 25.3 & 4 (2006), 451–466. 07–500 Xiaoxia, Cui (Yunnan U, China), An understanding of ‘China English’ and the learning and use of the English language in China. English Today (Cambridge University Press) 22.4 (2006), 40–43. 07–501 Young, Ming Yee Carissa (Macao U Science & Technology, Macau; myyoung@must.edu.mo ), Macao students' attitudes toward English: A post-1999 survey. World Englishes (Blackwell) 25.3 & 4 (2006), 479–490.
In this chapter I argue for a key premise in the argument I will present (in chapter 4) for anti-individualism regarding the individuation of speech content, linguistic meaning, and (ultimately) the attitudes. The premise itself asserts the existence of public linguistic norms. As developed in the argument to follow, such norms are normative in that they provide standards for the correct usage and interpretation of lexical items (that is, they provide the semantic standards of these items); and such norms are public, first, in that they derive from the shared, public language that is common property to all members of a speech community, and second, in that participants in speech exchanges – speakers and hearers alike – are positively but defeasibly presumed to be answerable to these norms on each particular speech exchange. In this chapter I will be pursuing the idea that the practice whereby knowledge is spread through a speaker's use of language itself depends on the existence of semantic standards provided by the shared public language. Or rather: such standards are required if this practice is to be, as we take it to be, a pervasive and efficient way to spread very specific pieces of knowledge, under conditions in which speaker and hearer may know nothing about one another's speech and interpretative dispositions beyond what is manifested in the brief speech exchange itself.
MLR, I02.3, 2007 849 them to a traditional published form as was done inEngler's critical edition. This gives readers amuch more immediate sense of connection with Saussure's thought, as well as providing some inkling ofwhat it is like towork with theoriginal manuscript materials. The textproper ends on page 240, with 87 pages given over to the bibliography of secondary literatureon Saussure since I970. It isa very fulland useful listingcovering awide range ofworks coming from linguists and literarycritics, and while some omis sions are inevitable, nothing can detract from the fact that this volume marks a true watershed in thedevelopment of Saussurean studies in theEnglish-speaking world. UNIVERSITY OF EDINBURGH JOHN E. JOSEPH La Langue, lestyle,lesens: etudesoffertes aAnne-Marie Garagnon. By CLAIRE BADIOU MONFERRAN, FRiED1ERIC CALAS, JULIEN PIAT, and CHRISTELLE REGGIANI. Paris: L'Improviste. 2005. 383pp. E28. ISBN978-2-913764-27-9. The titleof thevolume, a collection of articles inhonour ofAnne-Marie Garagnon, indicates the threemain themes explored. There are fourmain sections which reflect the interestsboth of the dedicatee and of her colleagues and formerpupils. The first section, entitled 'La langue entre histoire et systeme', isparticularly rich and inter esting. It examines what the editors call 'l'heritage conceptuel etmethodologique du xxe siecle' (p. 7), opening up interesting questions in the history of French linguis tic thought and of the French language, and in linguistic theory and ideologies of language. The second part, 'Faits de langue et faitsde style', includes articles which discuss and analyse particular lexical and semantic features of the French language from the Renaissance to the present day. The third part, 'Effets de style et effets de discours', offers insights into current trends in stylistic analysis, which inFrance is particularly marked by theories of 'enonciation'. Notable in the final part, en titled 'Stylistique et hermeneutique des formes', are the articles byGeorges Molinie and Thomas Clerc, which aim to present the epistemological basis ofAnne-Marie Garagnon's stylistic approach; other contributors explore certain stylisticmotifs in literary texts ranging fromLa Princesse de Cleves to Proust. It is impossible in the space of a review to do justice to all the articles and I propose to focus on a num ber of papers which seem to reflectwell the overall quality of the volume. Gilles Siouffi's excellent article, for example, reviews the notion of 'langue classique' and considers how a stylistic and grammatical analysis of its features can enable us better tounderstand it.He challenges the paradoxical notion of classical French as at once forming the norm forcontemporary French and being considered as 'radicalement autre' (p. i). A number of contributions treat the analysis of proper names fromdif ferentperspectives. Delphine Denis provides an interesting discussion of the status of the proper name in the seventeenth century. She shows how in this period there gradually emerges a theoryof proper nouns 'plus ou moins rapprochee du nom com mun' (p. 33). She notes how thegrammars and volumes ofobservations on theFrench language elaborate an opposition between common and proper nouns, but, with the possible exception of thePort-Royal logic,never discuss thecomplex nature ofproper nouns or the nature of theirmeaning. Adopting a stylistic approach, Van Dung Le Flanchec uses an analysis ofMaurice Sceve's 'Mon Orphee' toquestion modern ana lyses of proper names in terms of antonomasia, which should in her view rather be considered a stylisticquestion. Adopting a similar stylistic approach, Claire Badiou Monferran looks at different types of anaphora in a corpus of seventeenth-century nouvelles galantes; once again she sees the repeated use of anaphoric demonstrative pronouns as being a stylistic traitrather than an inherent feature of theFrench of the 85o Reviews period. Notable among the articleswhich adopt an approach associated with theories of 'enonciation' is the article by Frederic Calas. Starting froman analysis of an extract from Marivaux's Yournaux and fromLesage's Gil Blas de Santillane, he demonstrates thatepanorthosis isnot simply a rhetorical figurebut also gives thediscourse an ironic effect,therebycriticizing 'l'hypocrisiemondaine' (p. 247). In short, thevolume offers insights into a wide range of perspectives which are currently being applied to the study of French language and literature.What theyperhaps all have in common is the desire to reflecton 'legeste hermeneutique' (p. i). UNIVERSITY...
A method of data collection is presented that unites the efficiency of mass testing with the ease of instant electronic data collection that is typical of computer-based experiments run on individual participants. A wireless response system (WRS), originally designed as a teaching tool, is used to replicate three classic and robust effects from the memory literature (effects of false memory, levels of processing, and word frequency). It is shown that for these types of experimental designs, data can be collected more efficiently (in both time and effort) with the WRS method than through traditional mass- and individual-testing methods alone. The advantages and limitations of WRSs for use in mass electronic data collection are discussed.
The task we are investigating is unsupervised learning of natural language morphology for inflectional languages. The target morphological grammar consists of a lexicon of morphological base forms and transforms. A base form represents all inflections of a lexeme, and all base forms of the same POS category share the same fine-grained morphosyntactic type. Transforms are morphophonemic rewrite rules that convert base forms to derived forms, and whose context of application is limited to a specific set of base forms. We have developed a greedy algorithm to induce such a grammar. At each iteration, suffixal transforms to convert between base and derived forms of lexemes are hypothesized. The algorithm chooses the transform that maximizes vocabulary coverage, while minimizing the number of conflicts resulting from proposing as base forms words previously found to be derived forms. After base forms and transforms have been learned, a distributional clustering step assigns the base forms to POS classes. In future work, the transforms will be converted to generalized rewrite rules by inducing phonological characteristics common to the base forms. We have tested this algorithm on a version of the Penn Treebank annotated for inflectional morphology. The algorithm achieves 71.7% recall and 92.9% precision on inflectional relations, where both a base and derived form occur in the corpus. We are currently testing the algorithm on other languages, and will present results on the Morphochallenge gold standards.
This article investigates the use of the lexical items belonging to the complementizer phrase of relative clauses (relative CP): overt complementizer that, null complementizer o, and relative pronoun which. The research purpose was to understand whether language learners are moving toward the target norm for nativeness, in terms of the selection of lexical items in the relative CP. Two hypotheses were tested: (1) Learners prefer the [?wh] relative CP, as L2 proficiency improves; and (2) learners prefer the [?wh] relative CP more for DO RCs than for SUB RCs. Data were elicited from Korean learners, teachers, and native speakers by way of a cross-sectional production method. A picture-description instrument was employed in which the participants were presented with sets of two pictures containing the same item. The statistical results show that learners, not teachers, show a closer correspondence to native speakers, in terms of the selection of lexical items in the relative CP, and it is teachers who seem to be moving away from the target norm for nativeness. Teachers’ feature preference is more deviated from native speakers’ than learners’ in the production of English relative CPs.
Journal Article Optimizing Procedures for the Making of Bilingual Dictionaries and the Concept of Linking Contrastive Lexical Databases Get access Godelieve Laureys Godelieve Laureys University of Gent (godelieve.laureys@Ugent.be) Search for other works by this author on: Oxford Academic Google Scholar International Journal of Lexicography, Volume 20, Issue 3, September 2007, Pages 295–311, https://doi.org/10.1093/ijl/ecm024 Published: 24 August 2007
Fanny Rinck s’intéresse au genre de l’article de recherche en linguistique et aux styles au sens d’usages singuliers qu’un auteur fait du genre. Elle se base sur une approche stylométrique et sur les méthodes de la linguistique de corpus pour mettre en évidence les caractéristiques spécifiques des textes de 15 auteurs au niveau morpho-syntaxique et lexical. L’analyse montre que certaines s’assimilent à un usage idiomatique de la langue et du genre, aux plans syntaxique, énonciatif et argumentatif. D’autres sont liées aux concepts abordés dans les articles et au type d’études qui y sont présentés. En comparant en quoi le genre varie avec l’auteur et indépendamment de lui, on rend compte de la tension entre normes individuelles et collectives du genre et la singularité des textes en termes de profils qui structurent de manière relativement stable le genre et le champ considéré.
In this paper we propose an algorithm for converting dependency structures to phrase structures. This algorithm mainly concerns the characteristics of non-configurational languages. We review current works in the field and on the basis of these works we try to adopt a more flexible approach to the problem. 1
The article deals with the lexical changes that occurred in the Belarusian language within the 1930s. Books many times printed in the 1930s were the source of the study. Let alone that a certain stability in the lexicon of the Belarusian standard language was achieved by the end of the 1920s, one can observe a landslide of many norms previously used in editions of 1933 and later on. These changes happened against the background of the party-state activities in the field of the corpus planning of Belarusian (the “Results of the Discussion on Linguistics” resolution of the bureau of the Central Committee of the Belarusian Communist Party of 1930, the “About changes and simplification of the Belarusian orthography” resolution of the Council of People’s Commissars of the Belarusian Soviet Socialist Republic of 1933, language cultivation campaign etc.). Changes mostly concerned Polonisms and words identical to Polish, as well as unique neologisms and colloquial Belarusian words. Mainly Russianisms and words with Russian cognates replaced them.
The present research addresses the induction of emotion during music listening in adults using categorical and dimensional theories of emotion as background. It further explores the influences of musical preference and absorption trait on induced emotion. Twenty-five excerpts of classical music representing `happiness', `sadness', `fear', `anger' and `peace' were presented individually to 99 adult participants. Participants rated the intensity of felt emotions as well as the pleasantness and arousal induced by each excerpt. Mean intensity ratings of target emotions were highest for 20 out of 25 excerpts. Pleasantness and arousal ratings led to three main clusters within the two-dimensional circumplex space. Preference for classical music significantly influenced specificity and intensity ratings across categories. Absorption trait significantly correlated with arousal ratings only. In sum, instrumental music appears effective for the induction of basic emotions in adult listeners. However, careful screening of participants in terms of their musical preferences should be mandatory.
A parallel treebank consists of syntactically annotated sentences in two or more languages, taken from translated documents. These parallel sentences are linked through alignment. This paper explores the use of word n-gram alignment, computed for statistical machine translation, to create syntactic phrase alignment. We achieve a weighted F0.5 -score of over 65%.
Collins’ widely-used parsing models treat noun phrases (NPs) in a different manner to other constituents. We investigate these differences, using the recently released internal NP bracketing data (Vadas and Curran, 2007a). Altering the structure of the Treebank, as this data does, has a number of consequences, as parsers built using Collins’ models assume that their training and test data will have structure similar to the Penn Treebank’s. Our results demonstrate that it is difficult for Collins’ models to adapt to this new NP structure, and that parsers using these models make mistakes as a result. This emphasises how important treebank structure itself is, and the large amount of influence it can have.
The lexicon of the modern Norwegian bokmal standard needs a better description and documentation than what is the situation today. The article presents a plan for building a modern lexical database based on a balanced corpus of 40 million words of modern bokmal. This base should serve as a source for a traditional scientific dictionary as well as a dictionary for language technological applications.
This paper presents a new research and development project called Papillon [Planas00]. It is a French-Japanese cooperation between laboratories GETA/CLIPS (Grenoble, France) and NII (Tokyo, Japan). Its goal is to build a French-English-Japanese multilingual lexical database by using interlingual links and to extract from it digital bilingual French-Japanese and Japanese-French dictionaries. These dictionaries will be available under the terms of an open source license. This project, initiated by some computational linguists, aims at being useful and open to all those who are interested in Japanese and French. A seminar was organized on the 10-12 of August 2000 in Tokyo [Planas00]. It was devoted to discussions aiming at reaching a general consensus on the structure and content of the database, and to decide some technical aspects of database development, i.e. database configuration, contents of the entries and link between the entries. Introduction There are few French-Japanese usa...
The article foregrounds some basic explanations about syntactic corpus annotation and the methodology of building treebanks. It draws attention to the linguistic theoretical models on the basis of which the majority of treebank annotation approaches are formed, and to the applicability of syntactic corpus annotation for the studies of descriptive and theoretical linguistics, and for natural language processing. In addition the article presents the results of three studies connected with multi-lingual inductive dependency parsing on the first treebank for the Slovene language, the Slovene Dependency Treebank.
Abstract In this chapter, we discuss the development and use of picture stimuli incorporated in the International Affective Picture System (IAPS, pronounced “eye-aps”; Lang, Bradley, & Cuthbert, 2005), a large set of emotionally evocative color photographs that includes pleasure, arousal, and dominance ratings made by men and women. The IAPS is currently used in experimental investigations of emotion and attention worldwide, providing experimental control in the selection of emotional stimuli, facilitating the comparison of results across different studies, and encouraging replication within and across psychological and neuroscience research laboratories. Numerous studies in our laboratory over the past 15 years have explored subjective, psychophysiological, behavioral, and neurophysiological reactions when viewing these affective stimuli. Basic findings from these studies, which will be informative for researchers considering or using the IAPS stimuli, are briefly summarized in this chapter.
Parsing unrestricted text is useful for many language technology applications but requires parsing methods that are both robust and efficient. MaltParser is a language-independent system for data-driven dependency parsing that can be used to induce a parser for a new language from a treebank sample in a simple yet flexible manner. Experimental evaluation confirms that MaltParser can achieve robust, efficient and accurate parsing for a wide range of languages without language-specific enhancements and with rather limited amounts of training data.
The PARC 700 dependency bank is a potentially very useful resource for parser evaluation that has, so to speak, a high barrier to entry, because of tokenisation that is quite different from the source of the data, the Penn Treebank, and because there is no representation of word order, producing an uncertainty factor of some 15%. There is also a small, but perhaps not insignificant, number of errors. When using the dependency bank for evaluation, it seems likely that these things will cause inflated counts for mismatches, so to obtain more accurate measurements, it is desirable to eliminate them. The work reported here consists of an automatic conversion of the dependency bank into a Prolog representation where the word order is explicit, as well as graphical representations of the dependency trees for all 700 sentences, automatically generated from the Prolog data. As a side effect of the transformation, errors were detected and corrected. It is hoped that this work will lead to more widespread use of the PARC 700 dependency bank for parser evaluation.
Query expansion(QE) has been proved to be one of effective methods for improving the performance of the information retrieval(IR) system.Therefore,a new fuzzy QE method based on synonymy thesaurus is proposed,and the synonymy thesaurus is built based on the famous lexical database WordNet.In the synonymy thesaurus,the similarity between the synonyms is,which is obtained by Tanimoto coefficient.By using this synonymy thesaurus,query expansion can be done well.Then the fuzzy QE method is introduced into the document information retrieval system together with the modified vector space model.The experimental results show that the developed information retrieval system has got more effective performance than before by using the fuzzy query expansion method.One feature of the proposed information retrieval model is that it can be treated as one of simple semantic models.Another feature is that the expansion degree is controllable based on different thresholds.
This paper proposes a novel Chinese syntactic parsing model based on semantic class, which is a variant of normal lexicalized statistical model. It attempts to make use of the syntactic and semantic similarity between Chinese words and then produces a more knowledgeable estimate of the probability of grammar rules. A simple but effective unsupervised method is designed to determine the proper semantic class of given words. Semantic class is used to improve the performance of parsing model. We evaluate our methods on the widely used Penn Chinese Treebank. Experimental results show that it outperforms a famous lexicalized model significantly on appropriate semantic class levels.
Collocation is of great importance in dictionary compilation and natural language processing.Collocation extraction is one of the principal applications of corpus linguistics.Automatic extraction of bi-grams as candidate collocations is studied on Penn Treebank using the criteria of log likelihood,chi square and mutual information as association measure.The experimental results show the feasibility of the statistical methods.On the other hand,collocations extracted show different characteristics because of the different distribution assumptions by the three criteria.
Many problems in NLP require solving a cascade of subtasks. Traditional pipeline approaches yield to error propagation and prohibit joint training/ decoding between subtasks. Existing solutions to this problem do not guarantee nonviolation of hard-constraints imposed by subtasks and thus give rise to inconsistent results, especially in cases where segmentation task precedes labeling task. We present a method that performs joint decoding of separately trained Conditional Random Field (CRF) models, while guarding against violations of hard-constraints. Evaluated on Chinese word segmentation and part-of-speech (POS) tagging tasks, our proposed method achieved state-of-the-art performance on both the Penn Chinese Treebank and First SIGHAN Bakeoff datasets. On both segmentation and POS tagging tasks, the proposed method consistently improves over baseline methods that do not perform joint decoding.
In this paper, we address the issue of improving a Chinese chunking system with rich lexicalized information. A method that incorporates statistical information based on distributional similarity between words obtained from large unlabeled corpus and morphological knowledge into a state-of-the-art CRF-based chunking model is proposed to tackle the data sparseness problem given limited amount of labeled training data. Evaluations are performed on the latest release of Chinese Treebank, and experimental results show that our method outperforms the chunking models based on features over word and automatically assigned POS tags when using the same amount of training data.
We report work in progress on a complex system generating Czech sentences expressing the meaning of input syntactic-semantic structures. Such component is usually referred to as a realizer in the domain of Natural Language Generation. Existing realizers usually take advantage of a background linguistic theory. We introduce the Functional Generative Description, a framework of our choice conceived in 1960's by Petr Sgall. This language theory lays out foundations of the formalism in which our input syntactic-semantic structures are specified. The structure definition was further elaborated and refined during the annotation of the Prague Dependency Treebank, now available in its second version. A section of the paper is devoted to description of another theoretical framework suitable for the task of Natural Language Generation - the Meaning-Text Theory. We explore state-of-the-art realizers deployed in real life applications, describe common architecture of a generation system and highlight the strengths and weaknesses of our approach. Finally, preliminary output of our surface realizer is compared against a baseline solution.
In this paper we present a quantitative analysis of a bilingual lexical database which has been produced with OMBI, a tool for creating and editing bilingual dictionaries. OMBI has proven to be a valuable tool in the creation of rich bilingual multi-purpose lexical databases. One of the most distinctive features of the tool is reversal of source language and target language in order to create bilingual dictionaries in an economic and accurate way. We will focus on OMBI's reversal function, its initial concept and its results in practice. © 2007 Oxford University Press. All rights reserved.
We present an unsupervised linguistically-based approach to discourse relations recognition,\nwhich uses publicly available resources like manually annotated corpora (Discourse Graph\nBank, Penn Discourse TreeBank, RST-DT), as well as empirically derived data from “causally”\nannotated lexica like LCS, to produce a rule-based algorithm. In our approach we use\nthe subdivision of Discourse Relations into four subsets – CONTRAST, CAUSE, CONDITION,\nELABORATION, proposed by [1] in their paper where they report results obtained with a\nmachine-learning approach from a similar experiment against which we compare our results.\nOur approach is fully symbolic and is partially derived from the system called GETARUNS,\nfor text understanding, adapted to a specific task: recognition of Discourse Causal Relations\nin free text. We show that in order to achieve better accuracy both in the general task and in\nthe specific one, semantic information needs to be used besides syntactic structural information.\nOur approach outperforms results reported in previous papers
Typically, personalized information recommendation services automatically infer a user profile, a structured model of the user interests, from documents the user already deemed as relevant. Traditional keyword-based approaches are unable to capture the semantics of the user interests. This work proposes a strategy consisting of two steps. The first one is a semantic indexing procedure based on a word sense disambiguation strategy which exploits the WordNet lexical database to select, among all the possible meanings (senses) of a polysemous word, the correct one. In the second step, semantically indexed documents are mined by a naive Bayes learning algorithm that infer semantic, sense-based user profiles. Two experimental sessions were carried out to compare the performance of keyword-based profiles to that of sense-based profiles. We measured both the classification accuracy and the effectiveness of the ranking imposed by the two different kinds of profile on the documents to be recommended. The main outcome of both experiments is that the classification accuracy is improved without improving the ranking. Personalized systems adapt their behavior to individual users by learning their preferences during the interaction in order to construct a user profile that can be later exploited in the search process. Traditional keyword-based approaches are primarily driven by a string-matching operation: If a string, or some morphological variant, is found in both the profile and the document, a match is made and the document is considered relevant. String matching suffers from problems of polysemy, the presence of multiple meanings for one word, and synonymy, multiple words having the same meaning. The result is that, due to synonymy, relevant information can be missed if the profile does not contain the exact
Models of social evaluation aim to capture the information people use to form first impressions of unfamiliar others. However, little is currently known about the relationship between perceived traits across gender. In Study 1, we asked viewers to provide ratings of key social dimensions (dominance, trustworthiness, etc.) for multiple images of 40 unfamiliar identities. We observed clear sex differences in the perception of dominance-with negative evaluations of high dominance in unfamiliar females but not males. In Study 2, we used the social evaluation context to investigate the key predictions about the importance of pictorial information in familiar and unfamiliar face processing. We compared the consistency of ratings attributed to different images of the same identities and demonstrated that ratings of images depicting the same familiar identity are more tightly clustered than those of unfamiliar identities. Such results imply a shift from image rating to person rating with increased familiarity, a finding which generalises results previously observed in studies of identification.
Proceedings of the Sixth International Workshop on Treebanks and \nLinguistic Theories. \nEditors: Koenraad De Smedt, Jan Hajič and Sandra Kübler. \nNEALT Proceedings Series, Vol. 1 (2007), 85-96. \n© 2007 The editors and contributors. \nPublished by \nNorthern European Association for Language \nTechnology (NEALT) \nhttp://omilia.uio.no/nealt. \nElectronically published at \nTartu University Library (Estonia) \nhttp://hdl.handle.net/10062/4476.
Semantic differential techniques are a useful, well-validated tool to assess affective processing of stimuli and determine how that processing is impacted by various demographic factors, such as gender. In this paper, we explore differences in connotative word processing between men and women as measured by Osgood's semantic differential and what those differences imply about affective processing in the two genders. We recruited 94 young participants (47 men, 47 women, ages 18-39) using an online survey and collected their affective ratings of 120 words on three rating tasks: Evaluation (E), Potency (P), and Activity (A). With these data, we explored the theoretical and mathematical overlap between Osgood's affective meaning factor structure and other models of emotional processing commonly used in gender analyses. We then used Osgood's three-dimensional structure to assess gender-related differences in three affective classes of words (words with connotation that is Positive, Neutral, or Negative for each task) and found that there was no significant difference between the genders when rating Positive words and Neutral words on each of the three rating tasks. However, young women consistently rated Negative words more negatively than young men did on all three of the independent dimensions. This confirms the importance of taking gender effects into account when measuring emotional processing. Our results further indicate there may be differences between Osgood's structure and other models of affective processing that should be further explored.
An interface to WordNet using the Jawbone Java API to WordNet. WordNet (<<a href="https://wordnet.princeton.edu/" target="_top">https://wordnet.princeton.edu/</a>>) is a large lexical database of English. Nouns, verbs, adjectives and adverbs are grouped into sets of cognitive synonyms (synsets), each expressing a distinct concept. Synsets are interlinked by means of conceptual-semantic and lexical relations. Please note that WordNet(R) is a registered tradename. Princeton University makes WordNet available to research and commercial users free of charge provided the terms of their license (<<a href="https://wordnet.princeton.edu/license-and-commercial-use" target="_top">https://wordnet.princeton.edu/license-and-commercial-use</a>>) are followed, and proper reference is made to the project using an appropriate citation (<<a href="https://wordnet.princeton.edu/citing-wordnet" target="_top">https://wordnet.princeton.edu/citing-wordnet</a>>). The WordNet database files need to be made available separately, either via package 'wordnetDicts' from <<a href="https://datacube.wu.ac.at" target="_top">https://datacube.wu.ac.at</a>>, installing system packages where available, or direct download from <<a href="https://wordnetcode.princeton.edu/3.0/WNdb-3.0.tar.gz" target="_top">https://wordnetcode.princeton.edu/3.0/WNdb-3.0.tar.gz</a>>.
This paper reports on experiments in frame-semantic annotation of a parallel treebank. Selected English and Swedish sentences that contained verbs of motion and communication were annotated independently by two annotators. We found that they assigned the same frame to corresponding sentences in 52% of the cases. This leads us to the conclusion that parallel treebanks can save considerable effort when building semantically annotated resources.
Objective To validate the translated Chinese version of Pain Assessment in Advanced Dementia Scale (C-PAINAD) for its clinical application in assessing the discomfort level of severely cognitive impaired patients. Methods In developing the C-PAINAD, both clinical and academic experts were engaged in an iterative translation process for ensuring its semantic equivalence with its original language as well as it comprehensibility for clinical application. In establishing C-PAINAD's inter-rater reliability, its applicability for clinical use is further assessed in 11 severely cognitive impaired patients. The assessment tools included C-PAINAD, Discomfort Visual Analog Scale (DVA), and Philadelphia Geriatric Center Affect Rating Scale (PGCAR). Correlations, ANOVA and factor analysis were undertaken to examine the reliability and validity of C-PAINAD. Results The scores of C-PAINAD were not in normal distribution but clustered around Zero. C-PAINAD was positively correlated with DVA and negative affect but mildly and negatively correlated with positive affect. It was able to detect different pain level under different condition. One factor was extracted and the percentage of variance was 51.20%. C-PAINAD was proved to have satisfactory reliability and validity(Cronbach's α=0.66). Conclusions C-PAINAD was a simple, reliable and effective pain assessment instrument for measuring pain level of non-communicative patients like advanced dementia. Further research is deemed necessary to conduct with pre - and post-test of analgesic prescriptions prior to wider clinical application.
Reviewed by: Indian and British English: A handbook of usage and pronunciationby Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther Niladri Sekhar Dash Indian and British English: A handbook of usage and pronunciation. 2ndedn. By Paroo Nihalni, R. K. Tongue, Priya Hosali, and Jonathan Crowther. New Delhi: Oxford University Press, 2004. Pp. x, 260. ISBN 0195666569. $15.95. The present handbook is divided into two main parts. The first part (‘Lexicon of usage’) is designed to provide English users with information about the way in which certain words, idioms, collocations, phrases, and similar expressions of English used in India differ from British Standard English (BSE)—a model that has the closest affinity to Indian English. This part includes a thousand English words, which are used in a distinctive manner by large numbers of educated Indian speakers of English irrespective of their place, profession, education, gender, or other sociolinguistic factors. The words included in the handbook are selected from the speech or writing samples of the persons (such as university and school teachers, journalists, and radio commentators) who are likely to influence the English use of Indian learners. The handbook also contains many European words that have been Indianized over the years. Thus, it serves as a handy resource for Indian speakers of English, illustrating the many, often quite subtle, ways in which Indian English differs from standard British English usage, and where these differences are regarded as acceptable or substandard in the subcontinent. Examples in the handbook, which supplement the texts, are helpful to Indian users of English who are uncertain about the ‘correctness’ of their speech and writing, and serve those scholars who want to explore the differences between Indian and British uses of English. The book also has the potential to address special problems faced by learners of English, who are often impeded by the difficulties of recognizing finer nuances of meaning and usage. The second part of the handbook includes a brief report on the development of the pronunciation dictionary in India and abroad, followed by insightful discussions on standards of pronunciation in second/ foreign language teaching, the phonological systems of the British Received Pronunciation (BRP) and Educated Indian English (EIE), and the role of supraseg-mental properties (i.e. word stress, sentence stress, rhythm, intonation, etc.) in Indian English. The introduction contains a list of keywords for phonetic symbols used in the following part, ‘Dictionary of pronunciation’. Two types of pronunciation (Indian Recommended Pronunciation and the BRP) are supplied for more than two thousand words collected from the original lexical database of Michael West’s General service list of English wordstogether with a few additions compiled from the language resources available to the compilers. Each entry of the dictionary is tagged with relevant phonological information. This second edition also includes additional information on lexical collocation (the tendency of words to be used together in fixed phrases). In essence, the handbook not only serves as an invaluable reference guide for students and teachers of English, but also makes a valuable contribution for applied linguists, lexicographers, journalists, and scholars who write in Indian English. [End Page 465] Niladri Sekhar Dash Indian Statistical Institute, Kolkata Copyright © 2007 Linguistic Society of America
We propose a complex rule-based system for generating Czech sentences out of tectogrammatical trees, as introduced in Functional Generative Description (FGD) and implemented in the Prague Dependency Treebank 2.0 (PDT 2.0). Linguistically relevant phenomena including valency, diathesis, condensation, agreement, word order, punctuation and vocalization have been studied and implemented in Perl using software tools shipped with PDT 2.0. Parallels between generation from the tectogrammatical layer in FGD and deep syntactic representation in Meaning-Text Theory are also briefly sketched.
As the Introduction to this volume observes, sixteenth-century France is marked by ‘une vaste réflexion sur le bien dire’. This not only impacted upon theory and practice across the different literary genres using French but also promoted considerable debate on the form and basis of the emerging standard form of the vernacular. Given the wide-ranging nature of this réflexion, covering its different manifestations in one volume poses an almost insuperable challenge, but the twenty-eight contributions to the colloquium collected here certainly address an ambitiously broad span of topics and add usefully to our understanding of cultural developments in a period of major change. The papers are organized into three general sub-sections, ‘Interroger la norme’, ‘Évolutions de la norme’ and ‘Normes et société’. However, such is the fluid nature of the subject matter treated in certain papers that their classification under one or other of these headings can sometimes seem of doubtful appropriateness. The focus of the first sub-section is predominantly literary. The various contributions address the creation or adaptation of norms across a considerable number of different genres some of which are perhaps rather less familiar, for instance, oracular writings (Dubois), accounts of pilgrimages (Gomez-Géraud) and Jesuit letter-writing (Laborie). Particularly interesting is the close study by Duché of the influential approach which Nicolas Herberay adopted for translation. Herberay, an acknowledged master of French prose (‘un vray Cicero françois’, according to Jean Martin), wrote with a female as well as a male readership in mind, developing a prose style that was eloquent and natural that would set an example for bien dire in this area. The second sub-section begins with a cogent overview (Baddeley) of a familiar field, developments in orthography and the interplay between orthography and spelling, and is followed by a series of studies which explore revealingly topics such as the evolving relationship between poetics and grammar (Monferran), the increasing limitation on the use of metaphor in literary works (Cernogora) and developments in historiography (Dumontet). Perhaps the most interesting paper is the examination of the fortunes of the alexandrine in the early part of the century (Halévy). Particular attention is given to the writings of Jean Lemaire de Belges and Geoffroy Tory both of whom, on the basis of fanciful argumentation, sought to invest the alexandrine with special prestige and nationalistic symbolism matching the terza rima in Italian. Their exercises in myth-making were to contribute indirectly, it is argued, to the rapid rise in the alexandrine's use from around 1555. The final sub-section of the volume contains contributions that more particularly address linguistic issues. Notable amongst these are two items: a re-evaluation of the system of vers mesurés devised by Baïf which is seen as an attempt not only to reproduce the metrical patterns of ancient Greek but also to contribute towards the norms of spoken French by reflecting the élite ‘usage des Bons’ (Vignes); and a meticulous examination by Morin of change in the pronunciation norms presented by Peletier du Mans in his earlier works (1550, 1555) as against his 1581 Euvres poëtiques, the new norm correlating with that presented later in the works of Lanoue (1596) and La Touche (1696). Alongside these are a number of other attractive essays including a study of the linguistic norms in the speeches made at the formal opening of the Paris Parlement, with eloquence and high rhetoric dominating over practicality and clarity between 1560 and 1600 before a reversal occurred in the early seventeenth century (Petey-Girard), and an investigation of sixteenth-century liminaires (any text preceding a written work) composed by women where a complex set of norms operate involving humility, simplicity of style, the practice of dedicating the work to another woman and, in the light of the lack of image for the female writer, an attempt to ‘socialiser l'auteur’ (Gauthier). Completing the text is an Index Nominum and a table of contents. The diversity and scholarly depth of the volume should ensure that all seiziémistes will derive benefit from a close reading.
This paper presents recent extensions to Poliqarp, an open source tool for indexing and searching morphosyntactically annotated corpora, which turn it into a tool for indexing and searching certain kinds of treebanks, complementary to existing treebank search engines. In particular, the paper discusses the motivation for such a new tool, the extended query syntax of Poliqarp and implementation and efficiency issues.
Language deviation is a linguistic device of purposeful violation of language norms,which serves not only as the necessity but also as the inevitability of language development.There are various forms of language deviation,including phonological deviation,lexical deviation,grammatical deviation,semantic deviation,graphological deviation,deviation of register and figurative deviation.It is the reflection of variety of language and its social nature.Language deviation does bring richer cultural connotation to every language.
The current research explores the effects of exemplars on the stereotype representation of one's ingroup. Previous research demonstrated that exposure to an ingroup exemplar affects the stereotype one holds of one's ingroup (Coats & Smith, 1999). The primary purpose of the present study was to examine whether this effect is moderated by relative ingroup size. Participants were placed into either a minority or majority group situation and exposed to 1 of 2 dissimilar exemplars of their ingroup. Later, they rated their ingroup. Ratings of the ingroup differed between exemplar conditions in unexpected ways, indicating that the exemplar affected participants' stereotype of their ingroup. Furthermore, exemplars had a stronger effect on participants in the minority group than those in the majority group. Finally, relative ingroup size and, to a marginal extent, exemplars were found to affect ratings of ingroup variability.
Discourse segmentation is the task of determining minimal non-overlapping units of discourse called elementary discourse units (EDUs). It can be further subdivided into sentence segmentation and sentence-level discourse segmentation. This paper addresses the latter, more challenging subtask, which takes a sentence and outputs the EDUs for that particular sentence. (1) Saturday, he amended his remarks to say that he would continue to abide by the cease-fire if the U.S. ends its financial support for the Contras. (1a) Saturday, he amended his remarks (1b) to say (1c) that he would continue to abide by the cease-fire (1d) if the U.S. ends its financial support for the Contras. In example (1), a sentence from a Wall Street Journal article taken from the Penn TreeBank corpus is further segmented into four EDUs, (1a), (1b), (1c) and (1d) (RST, 2002). Discourse segmentation, clearly, is not as easy as sentence boundary detection. The lack of consensus with regards to what constitutes an elementary discourse unit adds to the difficulty. Building a rule based discourse segmenter can be a tedious task since these rules would have to be based on the underlying grammar of the particular parser that is to be used. Therefore, we adopted a neural network model for automatically building a discourse segmenter from an underlying corpus of segmented text. We chose to use part-of-speech tags, syntactic information, discourse cues and punctuation. Our ultimate goal is to build a discourse parser that uses this discourse segmenter.