Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
A new method to recognize the Chinese verb-object collocation is proposed on the basis of the conditional random fields(CRFs) model.The CRFs based model is examined with verb subcategorization features,context features,and features of their combination.The experiments are carried on two different Chinese word segmentation and part-of-speech tagging settings,with part-of-speech filtering rules to optimize the experiment.The results show that the best performance is 87.40% in F-score over Tsinghua Chinese Treebank,and 74.70% in F-score over the segmentation and part-of-speech tagging scheme of Peking University.Experimental results show that CRF model is effective in recognizing Chinese verb-object collocation automatically.
The Columbia Arabic Treebank (CATiB) is a database of syntactic analyses of Arabic sentences. CATiB contrasts with previous approaches to Arabic treebanking in its emphasis on speed with some constraints on linguistic richness. Two basic ideas inspire the CATiB approach: no annotation of redundant information and using representations and terminology inspired by traditional Arabic syntax. We describe CATiB's representation and annotation procedure, and report on inter-annotator agreement and speed.
A text is the carrier of language,and the foundation of cognizing and distinguishing its styles and functions.Text typology offers TQA objective and theoretical underpinnings.Classification and cognition of text types is a fundamental and cognitive approach to text types and has the original guiding effects on TQA.Different text types are meant not only to make up different writing forms including lexical features,styles and norms of writing,rhetorical devices,etc,but also to form different language functions,different text focuses,different TT purposes and translation methods.Obviously,these differences require us to establish different assessment criteria and principles,which provide us with TQA references,and lend themselves to further explore and study TQA model so as to make it maneuverable and practical.
Based on a structural approach and semi-directive interviews, this study analyzes athletic rules and legal consciousness of adolescents in relation to their contexts of athletic practice (institutionalized versus self-organized context). The lexical analysis (with alceste software) demonstrates that the institutionalized context leads mainly to a civilized consciousness where legal norms transcend individuals; on the other hand, the self-regulated context is the basis for a responsible and moral consciousness. These results contradict common beliefs associated with these two sport practices.
We present the first data-driven dependency parser for Romanian, which has been developed using the MaltParser system and trained and evaluated on a dependency treebank for Romanian developed within the RORIC-LING project. The parser achieves a labeled attachment score of 88.6 % (unlabeled 92.0%) when evaluated on held-out data from the treebank. We present a partial error analysis, focusing on accuracy for different parts of speech and dependencies of different length.
SEER, Vol.87,M. 3, July2009 Reviews Marder,Stephen. A Supplementary Russian-English Dictionary (ASRED2). Second edition. SlavicaPublishers, Bloomington, IN, 2007.xxv+ 736pp. $44.95. The first editionofthisdictionary came out in 1992.A corrected reprint of 1994was followed by a Moscow mirror editionin 1995.For some readers, oftennativeRussianspecialists, thedictionary was something ofa shock,a shockforrecovery from whichtheMoscowmirror edition ismorethanample evidence.Peoplelikethepresent reviewer, somewhat takenaback(something regrettably reflected in his reviewof 1994: SEER, 72, 1, pp. 161-62),but nonetheless massively impressed, wenton to use the dictionary morethan regularly overthenext, well,fifteen years.His first-edition copymaywellstill be inpretty goodcondition, butthere canbe no doubtthatthissecondedition will,whilethefirst onewillforall sorts ofreasonscontinue tobe used,replace it. Time has passedand,within thattime,numerous dictionaries ofRussian have appeared,ofall sorts- in thefirst place,ASRED 1 anticipated them, openedmanyeyesinthesociety whoselanguageitwasrevealing, andinspired them.AndASRED 2, quitedifferent from them, providing invaluable, lexical and grammatical information and employing toolsto facilitate easyuse and accessibility, reappears, considerably expanded,toteachthemand takethem further. The first edition as producedbySlavicawasextremely durable;thissecond looksindestructible. It is a hardback, beautifully producedby thepublisher and,particularly, bytheauthor(though theauthorofa dictionary has to be a 'compiler', itseemsthatinthiscase we really aremuchclosertohavingan 'author').There are otherthings to do in life,alas, and theymustrender it put-downable, but it is mostcertainly all but unput-downable. And the wonderful listof such adjectivesin Russian on p. 79, under vnusabel'nyj, reinforces sucha conviction. Thereis a setofpreliminary pages:after theContents (no page number) comesthe Introduction to the Second Edition(ASRED 2), pp. i-iii,where theauthor buildsonASRED 1and further justifies it.Therefollows theIntroductionto theFirstEdition,pp. v-x, giving invaluableinformation on how besttousethedictionary and reminding us thatthedictionary's original guidingprinciple was 'to fillan alarming - and exasperating gap between whathasbeenrecorded andwhatitispossibletorecord', something inwhich it succeededmostnotably.On pp. xi-xxwe have a SelectedBibliography (Updated),withRussiansources(pp. xi-xviii)and Englishsources(pp. xviiixx ).We have thenAcknowledgments (SecondEdition), p. xxi,Acknowledgments (First Edition), p. xxii,ListofAbbreviations and Conventional Symbols (Updated),pp. xxiii-xxiv, and a RussianAlphabet, p. xxv.The bodyofthe dictionary is in thefollowing 736pages. Importantly, thereareextremely clearand fullentries: theheadwordsand other wordsandphrases within entries areall inboldcharacters and stressed; REVIEWS 527 morphological juncturebetweenstemand endingis marked;thereis exhaustivecross -referring to relatedentries; and an enormousamountofcultural information ofall sorts is given.One can open anypage to giveexamplesof thelast:autizm andAFE on p. 23,pénsija pò [vózrast]u on p. 83,kosój on p. 263, mitëk, mitrofánuska, andMít'kaonp. 323,ocepjátka onp. 406,Pjatèrocka onp. 506, slivát' on p. 579,urjük on p. 665, utjug on p. 669, and cetvërka on p. 706,each page opened withoutsearchingand additionalexamplesbeing citableon almost eachpage.Manyentries havesub-entries; so,forexample,ustrójstuo has eighty-four, spreadoverpp. 666-68, and úxohas twenty-nine, spreadover pp. 669-70 (and witha finequotationto illustrate otkúda rastút usi).The translations are also extremely apt,withfullexplanationand expansionas necessary. One couldgo on and on; itwouldseemto thisreviewer thattogo on and on hereisnotat all appropriate ornecessary - ASRED 2 is a labouroflove, an astonishing treasure troveoflinguistic, literary, political,cultural, social and scientific information aboutRussiaand theRussianlanguage.ASRED 1 was alreadyquitea phenomenon; ASRED 2 notonlyconfirms thatphenomenon but demonstrates, fifteen yearslater,thatit remainssupremely useful and needed,and a realjoy through whichtoflit and inwhichto dwell. StAndrews University Ian Press Sovik, MargretheB. Support, Resistance and Pragmatism: An Examination of Motivation inLanguage Policy inKharkiv, Ukraine. ActaUniversitatis Stockholmiensis, Stockholm SlavicStudies,34. Department ofSlavicLanguages andLiteratures, Stockholm University, Stockholm, 2007.356pp. Figures. Tables.Illustrations. Bibliographical references. Appendices. SEK 343.00 (paperback). The 'languagequestion'in Ukraine- specifically, the role and statusof Ukrainian and Russian- isone ofthoseheatedand controversial issuesthat repeatedly inserts itself intodomesticpoliticaldebates,especially, it seems, during electoral campaigns, whenpresidential candidates and political parties competing forvotesfindit to be a ratherusefultool formobilizing their supporters. Not infrequently, italso servesas a sourceofcontention between Kyivand Moscow.Much oftheproblemresidesin thefactthatin Ukraine languagepreference and usage largelyoverlapwiththe country's regional structure and conflicting politicalvalues,thereby setting the stageforthe 'languagequestion'to becomehighly politicized. The book underreview, whichwas written as a doctoraldissertation at Stockholm University, addressestheproblemin a veryspecific and, indeed, unorthodoxmanner.Instead of examiningpolicies and legal norms or analysing statistical data suchas thelanguageofinstruction in schoolsand universities or the resultsof country-wide sociologicalsurveys, whichhas beenthenormin thescholarly literature, theauthorfocuses on howthepredominantly Russian-speaking residents ofKharkiv - Ukraine's secondlargest...
Methods for discriminant analysis were compared with respect to classification accuracy under nonnormality through Monte Carlo simulation. The methods compared were linear discriminant analyses based both on raw scores and on ranks; linear logistic discrimination; and mixture discriminant analysis. Linear discriminant analysis and linear logistic discrimination were suboptimal in a number of scenarios with skewed predictors. Linear discriminant analysis based on ranks yielded the highest rates of classification accuracy in only a limited number of situations and did not produce a practically important advantage over competing methods. Mixture discriminant analysis, with a relatively small number of components in each group, attained relatively high rates of classification accuracy and was most useful for conditions in which skewed predictors had relatively small values of kurtosis.
This paper examines the ideologies and practices surrounding respect at a Korean American heritage language school in California. It illustrates the interaction between locally circulating metadiscourses about children’s dispositions, intentions, and identities and the enforcement of classroom norms of respect. In some cases, teachers accommodated to children’s linguistic norms though a metadiscourse that reframed the indexicality of potentially disrespectful behavior. In other cases, forms of bodily demeanor were naturalized as indexical of children’s deliberate communication of disrespect. Teachers’ classroom narratives presented theories of affective accommodation and affective display, where respect for a teacher’s feelings was supposed to be given priority over respect for a child’s feelings, but children did not always comply with these theories. By illustrating how teachers’ metapragmatic ideologies about children’s identities as Korean Americans, contexts of language acquisition, and linguistic needs mediate the interactional construction of (dis)respect, this paper demonstrates the hybrid/multidirectional nature of language socialization. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Background: Studies of experts' problem-solving abilities have shown that experts can attend to the deep structure of a problem whereas novices attend to the surface structure. Although this effect has been replicated in many domains, there has been little investigation into such effects in medicine in general or patient management in particular. Methodology/Principal Findings: We designed a 10-item forced-choice triad task in which subjects chose which one of two hypothetical patients best matched a target patient. The target and its potential matches were related in terms of surface features (e.g., two patients of a similar age and gender) and deep features (e.g., two diabetic patients with similar management strategies: a patient with arthritis and a blind patient would both have difficulty with self-injected insulin). We hypothesized that experts would have greater knowledge of management categories and would be more likely to choose deep matches. We contacted 130 novices (medical)
This article introduces the topic of “Multilingual language resources and interoperability”. We start with a taxonomy and parameters for classifying language resources. Later we provide examples and issues of interoperatability, and resource architectures to solve such issues. Finally we discuss aspects of linguistic formalisms and interoperability.
Lexical fluency tests are frequently used in clinical practice to assess language and executive function. As part of the Spanish multicenter normative studies (NEURONORMA project), we provide age- and education-adjusted norms for three semantic fluency tasks (animals, fruit and vegetables, and kitchen tools), three formal lexical tasks (words beginning with P, M, and R), and three excluded letter fluency tasks (excluded A, E, and S). The sample consists of 346 participants who are cognitively normal, community dwelling, and ranging in age from 50 to 94 years. Tables are provided to convert raw scores to age-adjusted scaled scores. These were further converted into education-adjusted scaled scores by applying regression-based adjustments. The current norms should provide clinically useful data for evaluating elderly Spanish people. These data may also be of considerable use for comparisons with other international normative studies. Finally, these norms should help improve the interpretation of verbal fluency tasks and allow for greater diagnostic accuracy.
Background: Implicit racial bias denotes socio-cognitive attitudes towards other-race groups that are exempt from conscious awareness. In parallel, other-race faces are more difficult to differentiate relative to own-race faces - the ''Other- Race Effect.'' To examine the relationship between these two biases, we trained Caucasian subjects to better individuate other-race faces and measured implicit racial bias for those faces both before and after training. Methodology/Principal Findings: Two groups of Caucasian subjects were exposed equally to the same African American faces in a training protocol run over 5 sessions. In the individuation condition, subjects learned to discriminate between African American faces. In the categorization condition, subjects learned to categorize faces as African American or not. For both conditions, both pre- and post-training we measured the Other-Race Effect using old-new recognition and implicit racial biases using a novel implicit social measure - )
The article discusses a study which investigated whether subtitles, which provide lexical information, support perceptual learning about foreign speech. As explained, subtitles indicate which words are being spoken, which could improve lexically-guided learning about foreign speech sounds. As part of the study, Dutch participants were asked to watch videos containing unfamiliar regionally-accented English, with or without subtitles. Both English and Dutch subtitles were used which made it possible to compare the effects of subtitles in the language spoken in the videos with the effects of subtitles in the observers' native language.
Background: One of the most debated issues in the cognitive neuroscience of language is whether distinct semantic domains are differentially represented in the brain. Clinical studies described several anomic dissociations with no clear neuroanatomical correlate. Neuroimaging studies have shown that memory retrieval is more demanding for proper than common nouns in that the former are purely arbitrary referential expressions. In this study a semantic relatedness paradigm was devised to investigate neural processing of proper and common nouns. Methodology/Principal Findings: 780 words (arranged in pairs of Italian nouns/adjectives and the first/last names of well known persons) were presented. Half pairs were semantically related (''Woody Allen'' or ''social security''), while the others were not (''Sigmund Parodi'' or ''judicial cream''). All items were balanced for length, frequency, familiarity and semantic relatedness. Participants were to decide about the semantic relatedness of t)
The hands and mouth do not always slip together in British Sign Language: Dissociating articulatory channels in the lexicon David P. Vinson (d.vinson@ucl.ac.uk) Robin L. Thompson (robin.thompson@ucl.ac.uk) Robert Skinner (robert.skinner@ucl.ac.uk) Neil Fox (neil.fox@ucl.ac.uk) Gabriella Vigliocco (g.vigliocco@ucl.ac.uk) Deafness, Cognition and Language Research Centre, Department of Cognitive, Perceptual and Brain Sciences University College London, 26 Bedford Way, London, WC1H 0AP, UK LUNCH 1 are distinguished only by English-derived mouthings), but they are also commonplace in nonambiguous signs, occurring very frequently in spontaneous conversation, and are often considered to be part of the signs themselves (see Boyes Braem & Sutton- Spence, 2001, for further discussion) 2. However, there is little evidence concerning the precise nature of the link between mouthings and manual elements of lexical signs in language production, and the nature of the systems underlying their retrieval and production. It is certainly the case that the two must diverge at some point, because they rely upon different articulatory systems (hands vs. mouth). Our primary question concerns the extent to which these representations are linked before this divergence takes place. On one hand, mouthings might reflect the activation of representations based on a spoken language, which are accessed relatively independently from the sign language representations driving the manual component of signs. As such they would be incidental to the retrieval of the manual form, rather than being integrated before phonological and especially phonetic encoding.. On the other hand, although mouthings historically originated as a borrowed form from the surrounding spoken language, they may have become fully embedded within the sign language production system and thus completely integrated with the manual component of signs. In order to test these two alternatives, we employed a lexical retrieval task targeting the semantic level of representation: cyclic semantic blocking (Kroll & Stewart, 1994). In this task, participants repeatedly name objects presented in contexts of other objects that are either semantically related or unrelated to each other. In spoken languages, speakers are slower to name pictures when they are presented in the context of semantically related items, an Abstract We investigate the extent of integration between the hands and mouthing for lexical signs in British Sign Language, using picture naming and translation tasks that are sensitive to semantic similarity effects in lexical retrieval. Semantic errors in sign forms due to semantically related contexts were more common in translation from English than in picture naming, while semantic errors in mouth patterns were sensitive to semantic context only in picture naming, and not in translation from English. These results are consistent with an account whereby mouthing is accessed through a largely separable channel from manual components of the sign lexicon, rather than being bundled with manual components and incorporated into the sign language lexicon despite its original relationship to English. Effects did not differ between Deaf and hearing native signers, suggesting that stronger links between orthography and phonology in the hearing group do not play a role. Keywords: lexical retrieval, production, sign language, mouthing, semantic competition Introduction Signed language production involves the simultaneous use of multiple articulators; not only the two hands themselves, but also other articulators such as the body, face, and the mouth have both lexical and grammatical functions (e.g., to convey adjectival or adverbial information, or to mark negation, yes-no questions or relative clauses). In addition to mouth patterns that can be used to express adjectival or adverbial information, many lexical signs are associated with specific mouth patterns which are integral to a specific sign and are time-locked to production of the sign's manual component (i.e., the movement of the hands, Boyes Braem & Sutton-Spence, 2001). These mouth patterns are of two types: those originating within the sign language system, and those derived from a spoken language. The former, sometimes termed mouth gestures, use abstract vocal properties (e.g., inhalation/exhalation, mouth shape, or articulation) to reflect properties of the manual signs themselves (Woll & Sieratzki, 1998). The latter, instead (often termed mouthings ), are derived from the pronunciation of words in a spoken language. Sometimes mouthings are used to distinguish between ambiguous sign forms (for example, the British Sign Language (BSL) signs BREAKFAST and Signs in BSL are customarily represented as English glosses in capital letters. For example, in a set of 300 lexical signs produced by Deaf BSL signers for use in a lexical norming study (Vinson, et al. 2009), more than 90% included mouthing, although the sign models were given only general instructions to produce the signs as naturally as possible, and no mention was made of mouthing. This is likely an overestimate of the rate at which mouthing occurs in discourse (these signs were produced in isolation) but gives an impression of the importance of mouthing.
Author’s Introduction African languages have played an important role in the development of linguistic theory but their role in the fields of historical linguistics and linguistic typology has been less prominent. Africa’s linguistic diversity has been long underestimated given the dominance of the four‐family model proposed by Joseph Greenberg. Criticism of this model has long held among specialists in some of Africa’s smaller and lesser‐known language families, but has only recently become more widely acknowledged among linguists. Archaeologists, geneticists, and others continue to model African prehistory based on African linguistic classifications, which are outdated and which have failed to withstand scrutiny. This teaching and learning guide suggests a program to train scholars in recognizing and evaluating the standards by which various African language classifications have been made. Africa’s linguistic diversity will be shown to be far greater than what is suggested by the four‐family model. Author Recommends 1. Childs, G. Tucker. 2003. The classification of African languages. An Introduction to African Languages, 19–53. Amsterdam: John Benjamins. A highly readable introduction to the subject, ‘The classification of African languages’ provides an overview of the four‐phylum model, along with the discussion of some of the major critiques that have been made of it. 2. Dimmendaal, Gerrit J. and F.K. Erhard Voeltz. 2007. Africa. Encyclopedia of the World’s Endangered Languages, 579–634. ed. by Christopher Moseley. London: Routledge. ‘Africa’ provides an overview of the major changes to the picture of African language classification that have been accepted by leading historical linguists and specialists. This reading also includes a list of endangered African languages, and an up‐to‐date discussion of Africa’s linguistic areas. 3. Campbell, Lyle and William J. Poser. 2008. Africa. Language Classification: History and Method, 203–245. Cambridge: Cambridge University Press. The ‘Africa’ section of Campbell and Poser’s book surveys the history of African language classification. Examples are given showing problems with the types of evidence used to uphold Greenberg and pre‐Greenberg classifications. 4. Dimmendaal, Gerrit J. 2008. Language ecology and linguistic diversity on the African continent. Language and Linguistics Compass 2(5).840–858. DOI: DOI: 10.1111/j.1749-818X.2008.00085.x The article ‘Language ecology and linguistic diversity on the African continent’ discusses recent revisions to African linguistic classification and pays special attention to the development of spread zones and accretion zones on the continent. It also details several cases of contact‐induced linguistic changes. 5. Batibo, Herman M. 2005. The endangered languages of Africa. Language Decline and Death in Africa: Causes, Consequences and Challenges, 62–86. Clevedon: Multilingual Matters. The book ‘Language Decline and Death in Africa’ is an important survey of the factors that have been leading to a loss of linguistic diversity on the African continent. Chapter 5 ‘The endangered languages of Africa’ provides a concise summary of these factors, as well as a country‐by‐country outline of endangered languages. 6. Mous, Maarten. 2003. Loss of linguistic diversity in Africa. Language Death and Language Maintenance: Theoretical, Practical and Descriptive Approaches, 157–170. ed. by Mark Janse and Sijmen Tol. Amsterdam: John Benjamins. The article ‘Loss of linguistic diversity in Africa’ surveys the linguistic genetic diversity of the continent’s languages and provides an overview of the factors that have led to a language shift in Africa. Online Materials 1. The World Atlas of Language Structures Online (WALS) http://wals.info/index The World Atlas of Language Structures Online (WALS) is a database of over 58,000 datapoints of phonological, grammatical, and lexical features of a subset of the world’s languages. The database allows scholars see maps of the distribution of these features, and to read the 142 chapters examining these structural features across the typological database. The Online database is a project of Max Planck Institute for Evolutionary Anthropology and the Max Planck Digital Library, edited by Martin Haspelmath, Matthew S. Dryer, David Gil and Bernard Comrie (Munich: Max Planck Digital Library, 2008). WALS was also published as a book (with CD‐ROM) by Oxford University Press in 2005. 2. SIL Electronic Survey Reports http://www.sil.org/silesr/indexes/languages.asp The SIL Electronic Survey Reports website provides links to a large number of language surveys (in a downloadable pdf format) produced by Summer Institute of Linguistics (SIL) researchers. These reports include a great deal of linguistic information about a large number of African languages. Of particular interest to the comparative linguist is the large number of word lists included in the reports. 3. Jouni Maho’s Web Resources for African Languages http://www.africanlanguages.org/ Jouni Maho’s ‘Web Resources for African Languages’ pages provide a continuously updated list of online databases and downloadable documents on African languages. Links to both published and unpublished sources are included. The pages also provide a select print bibliography for each language family or isolate. 4. Comparative Bantu Online Dictionary (CBOLD) http://www.linguistics.berkeley.edu/CBOLD/ The Comparative Bantu Online Dictionary (CBOLD) at the University of California Berkeley as started in 1994 by Larry Hyman and John Lowe but includes contributions from scholars from many other institutions. It includes over 20 searchable Bantu language dictionaries, the Tervuren database of Bantu Lexical Reconstructions (BLR2) (with 9800 entries), and the Tanzanian Language Survey lexical database. 5. Ethnologue: Languages of the World http://www.ethnologue.com/ Ethnologue is an ‘encyclopedic reference work cataloging all of the world’s 6,912 known living languages’ produced by the Summer Institute of Linguistics (SIL). This website provides a quick and handy means of getting information about such topics as: the languages spoken in each
TempEval is a framework for evaluating systems that automatically annotate texts with temporal relations. It was created in the context of the SemEval 2007 workshop and uses the TimeML annotation language. The evaluation consists of three subtasks of temporal annotation: anchoring an event to a time expression in the same sentence, anchoring an event to the document creation time, and ordering main events in consecutive sentences. In this paper we describe the TempEval task and the systems that participated in the evaluation. In addition, we describe how further task decomposition can bring even more structure to the evaluation of temporal relations.
Since the inception of the Senseval series there has been a great deal of debate in the word sense disambiguation (WSD) community on what the right sense distinctions are for evaluation, with the consensus of opinion being that the distinctions should be relevant to the intended application. A solution to the above issue is lexical substitution, i.e. the replacement of a target word in context with a suitable alternative substitute. In this paper, we describe the English lexical substitution task and report an exhaustive evaluation of the systems participating in the task organized at SemEval-2007. The aim of this task is to provide an evaluation where the sense inventory is not predefined and where performance on the task would bode well for applications. The task not only reflects WSD capabilities, but also can be used to compare lexical resources, whether man-made or automatically created, and has the potential to benefit several natural-language applications.
The NLP community has shown a renewed interest in deeper semantic analyses, among them automatic recognition of semantic relations in text. We present the development and evaluation of a semantic analysis task: automatic recognition of relations between pairs of nominals in a sentence. The task was part of SemEval-2007, the fourth edition of the semantic evaluation event previously known as SensEval. Apart from the observations we have made, the long-lasting effect of this task may be a framework for comparing approaches to the task. We introduce the problem of recognizing relations between nominals, and in particular the process of drafting and refining the definitions of the semantic relations. We show how we created the training and test data, list and briefly describe the 15 participating systems, discuss the results, and conclude with the lessons learned in the course of this exercise.
Virtual reality technology is argued to be suitable to the simulation study of mass evacuation behavior, because of the practical and ethical constraints in researching this field. This article describes three studies in which a new virtual reality paradigm was used, in which participants had to escape from a burning underground rail station. Study 1 was carried out in an immersion laboratory and demonstrated that collective identification in the crowd was enhanced by the (shared) threat embodied in emergency itself. In Study 2, high-identification participants were more helpful and pushed less than did low-identification participants. In Study 3, identification and group size were experimentally manipulated, and similar results were obtained. These results support a hypothesis according to which (emergent) collective identity motivates solidarity with strangers. It is concluded that the virtual reality technology developed here represents a promising start, although more can be done to embed it in a traditional psychology laboratory setting.
Recently, Kuppens, Van Mechelen, and Rijmen (2008) developed a method that allows researchers to examine and disentangle the contributions of the different possible sources of variability in sequential processes that underlie psychological outcomes or behaviors. Although this method may prove valuable for many research domains in the social sciences, its use may be limited by its statistical complexity and the effort and programming skills required. We present an R package, called Desequens, intended to make this method easily accessible to social science researchers. The tool does not require any knowledge of R, so that R laymen can easily apply the method to their data as well. We demonstrate the use of Desequens by means of a didactic example.
creativeness / a pleasing field / of bloom Word associations are an important element of linguistic creativity. Traditional lexical knowledge bases such as WordNet formalize a limited set of systematic relations among words, such as synonymy, polysemy and hypernymy. Such relations maintain their systematicity when composed into lexical chains. We claim that such relations cannot explain the type of lexical associations common in poetic text. We explore in this paper the usage of Word Association Norms (WANs) as an alternative lexical knowledge source to analyze linguistic computational creativity. We specifically investigate the Haiku poetic genre, which is characterized by heavy reliance on lexical associations. We first compare the density of WAN-based word associations in a corpus of English Haiku poems to that of WordNet-based associations as well as in other non-poetic genres. These experiments confirm our hypothesis that the non-systematic lexical associations captured in WANs play an important role in poetic text. We then present Gaiku, a system to automatically generate Haikus from a seed word and using WAN-associations. Human evaluation indicate that generated Haikus are of lesser quality than human Haikus, but a high proportion of generated Haikus can confuse human readers, and a few of them trigger intriguing reactions.
We employ a single-trial correlational MEG analysis technique to investigate early processing in the visual recognition of morphologically complex words. Three classes of affixed words were presented in a lexical decision task: free stems (e.g., taxable), bound roots (e.g., tolerable), and unique root words (e.g., vulnerable, the root of which does not appear elsewhere). Analysis was focused on brain responses within 100-200 msec poststimulus onset in the previously identified letter string and visual word-form areas. MEG data were analyzed using cortically constrained minimum-norm estimation. Correlations were computed between activity at functionally defined ROIs and continuous measures of the words' morphological properties. ROIs were identified across subjects on a reference brain and then morphed back onto each individual subject's brain (n = 9). We find evidence of decomposition for both free stems and bound roots at the M170 stage in processing. The M170 response is shown to be sensitive to morphological properties such as affix frequency and the conditional probability of encountering each word given its stem. These morphological properties are contrasted with orthographic form features (letter string frequency, transition probability from one string to the next), which exert effects on earlier stages in processing ( approximately 130 msec). We find that effects of decomposition at the M170 can, in fact, be attributed to morphological properties of complex words, rather than to purely orthographic and form-related properties. Our data support a model of word recognition in which decomposition is attempted, and possibly utilized, for complex words containing bound roots as well as free word-stems.
Voice acoustic analysis is typically a labor-intensive, time-consuming process that requires the application of idiosyncratic parameters tailored to individual aspects of the speech signal. Such processes limit the efficiency and utility of voice analysis in clinical practice as well as in applied research and development. In the present study, we analyzed 1,120 voice files, using standard techniques (case-by-case hand analysis), taking roughly 10 work weeks of personnel time to complete. The results were compared with the analytic output of several automated analysis scripts that made use of preset pitch-range parameters. After pitch windows were selected to appropriately account for sex differences, the automated analysis scripts reduced processing time of the 1,120 speech samples to less than 2.5 h and produced results comparable to those obtained with hand analysis. However, caution should be exercised when applying the suggested preset values to pathological voice populations.
English is spoken worldwide by both native (L1) and nonnative (L2) speakers. It is therefore imperative to establish how easily L1 and L2 speakers understand each other. We know that L1 listeners adapt to foreign-accented speech very rapidly (Clarke & Garrett, 2004), and L2 listeners find L2 speakers (from matched and mismatched L1 backgrounds) as intelligible as native speakers (Bent & Bradlow, 2003). But foreign-accented speech can deviate widely from L1 pronunciation norms, for example when adult L2 learners experience difficulties in producing L2 phonemes that are not part of their native repertoire (Strange, 1995). For instance, Italian L2 learners of English often lengthen the lax English vowel /I/, making it sound more like the tense vowel /i/ (Flege et al., 1999). This blurs the distinction between words such as bin and bean. Unless listeners are able to adapt to this kind of pronunciation variance, it would hinder word recognition by both L1 and L2 listeners (e.g., /bin/ could mean either bin or bean). In this study we investigate whether Italian-accented English interferes with on-line word recognition for native English listeners and for nonnative English listeners, both those where the L1 matches the speaker accent (i.e., Italian listeners) and those with an L1 mismatch (i.e., Dutch listeners). Second, we test whether there is perceptual adaptation to the Italian-accented speech during the experiment in each of the three listener groups. Participants in all groups took part in the same cross-modal priming experiment. They heard spoken primes and made lexical decisions to printed targets, presented at the acoustic offset of the prime. The primes, spoken by a native Italian, consisted of 80 English words, half with /I/ in their standard pronunciation but mispronounced with an /i/ (e.g., trick spoken as treek), and half with /i/ in their standard pronunciation and pronounced correctly (e.g., treat). These words also appeared as targets, following either a related prime (which was either identical, e.g., treat-treat, or mispronounced, e.g., treek-trick) or an unrelated prime. All three listener groups showed identity priming (i.e., faster decisions to treat after hearing treat than after an unrelated prime), both overall and in each of the two halves of the experiment. In addition, the Italian listeners showed mispronunciation priming (i.e., faster decisions to trick after hearing treek than after an unrelated prime) in both halves of the experiment, while the English and Dutch listeners showed mispronunciation priming only in the second half of the experiment. These results suggest that Italian listeners, prior to the experiment, have learned to deal with Italian-accented English, and that English and Dutch listeners, during the experiment, can rapidly adapt to Italian-accented English. For listeners already familiar with a particular accent (e.g., through their own pronunciation), it appears that they have already learned how to interpret words with mispronounced vowels. Listeners who are less familiar with a foreign accent can quickly adapt to the way a particular speaker with that accent talks, even if that speaker is not talking in the listeners’ native language.
Se référant à divers ouvrages et auteurs de tous temps, styles ou nationalités, cette étude sur la représentation du substrat dialectal et étranger dans la littérature française et anglo-américaine, et sur sa traduction, cherche principalement à comprendre la démarche des auteurs recourant à la retranscription phonétique, pour ensuite mieux appréhender celle des traducteurs. Analysant en premier lieu les diverses motivations qui poussent ces écrivains à bouleverser les normes grammaticales et orthographiques pour transfigurer dans l’écrit, l’oral et la parole, puis s’interrogeant quant à la validité d’une méthode à appliquer à ces créations linguistiques, cet exposé tente de répondre, notamment par un examen énumératif des procédés matérialisant l’accent dialectal ou étranger, aux questions d’ordre lexical, grammatical ou morphosyntaxique qu’infère cette intrusion de la langue parlée dans le texte. Elucidant enfant les outils traductologiques mis en place dans ces littératures, ce travail propose un traducteur futur de pleinement s’en inspirer, et ainsi faire de l’achoppement, un argument à la créativité.
An important task of ontology learning is to enrich the vocabulary for domain ontologies using different sources of information. WordNet, an online lexical database covering many domains, has been widely used as a source from which to mine new vocabulary for ontology enrichment. However, since each word submitted to WordNet may have several different meanings (senses), existing approaches still face the problem of semantic disambiguation in order to select the correct sense for the new vocabulary to be added. In this paper, we present a similarity computation method that allows us to efficiently select the correct WordNet sense for a concept-word in a given ontology. Once the correct sense is identified, we can then enrich the concept's vocabularly using nearby words in WordNet. Experimental results using an amphibian ontology show that the similarity computation method reach a good average accuracy and our approach is able to enrich the vocabulary of each concept with words mined from WordNet synonyms and hypernyms.
The thesis studies the translation process for the laws of Finland as they are translated from Finnish into Swedish. The focus is on revision practices, norms and workplace procedures. The translation process studied covers three institutions and four revisions. In three separate studies the translation process is analyzed from the perspective of the translations, the institutions and the actors. The general theoretical framework is Descriptive Translation Studies. For the analysis of revisions made in versions of the Swedish translation of Finnish laws, a model is developed covering five grammatical categories (textual revisions, syntactic revisions, lexical revisions, morphological revisions and content revisions) and four norms (legal adequacy, correct translation, correct language and readability). A separate questionnaire-based study was carried out with translators and revisers at the three institutions. \n\nThe results show that the number of revisions does not decrease during the translation process, and no division of labour can be seen at the different stages. This is somewhat surprising if the revision process is regarded as one of quality control. Instead, all revisers make revisions on every level of the text. Further, the revisions do not necessarily imply errors in the translations but are often the result of revisers following different norms for legal translation. \n\nThe informal structure of the institutions and its impact on communication, visibility and workplace practices was studied from the perspective of organization theory. The results show weaknesses in the communicative situation, which affect the co-operation both between institutions and individuals. Individual attitudes towards norms and their relative authority also vary, in the sense that revisers largely prioritize legal adequacy whereas translators give linguistic norms a higher value. Further, multi-professional teamwork in the institutions studied shows a kind of teamwork based on individuals and institutions doing specific tasks with only little contact with others. This shows that the established definitions of teamwork, with people co-working in close contact with each other, cannot directly be applied to the workplace procedures in the translation process studied. Three new concepts are introduced: flerstegsrevidering (multi-stage revision), revideringskedja (revision chain) and normsyn (norm attitude). \n\nThe study seeks to make a contribution to our knowledge of legal translation, translation processes, institutional translation, revision practices and translation norms for legal translation. \n\n Keywords: legal translation, translation of laws, institutional translation, revision, revision practices, norms, teamwork, organizational informal structure, translation process, translation sociology, multilingual.
grammatical resources from treebanks for English Here, we extend the LFG grammar acquisition approach to Arabic and the Penn Arabic Treebank (ATB) (Maamouri and Bies, 2004), adapting and extending the methodology of Arabic is challenging because of its morphological richness and syntactic complexity. Currently 98% of ATB trees (without FRAG and X) produce a covering and connected f-structure. We conduct a qualitative evaluation of our annotation against a gold standard and achieve an f-score of 95%.
All too often work in computational linguistics on the acquisition of conceptual descriptions takes place in isolation from work on concepts in psychology and neural science. We feel this is a mistake as evidence from these related disciplines can provide us with better ways of evaluating our results. In the talk I will present work in CIMEC on using cognitive evidence to evaluate the results of lexical acquisition work - specifically, using feature norms to evaluate the acquisition of features, and using EEG data to evaluate the results of categorization experiments.
This paper presents a tool for extracting multi-word expressions from corpora in Modern Greek, which is used together with a parallel concordancer to augment the lexicon of a rule-based machine-translation system. The tool is part of a larger extraction system that relies, in turn, on a multilingual parser developed over the past decade in our laboratory. The paper reviews the various NLP modules and resources which enable the retrieval of Greek multi-word expressions and their translations: the Greek parser, its lexical database, the extraction and concordancing system.
The Semantic Web was designed to unambiguously define and use ontologies to encode data and knowledge on the Web. Many people find it difficult, however, to write complex RDF statements and queries because doing so requires familiarity with the appropriate ontologies and the terms they define. We describe a system that automatically maps a set of ordinary English words to a set of appropriate ontology terms on the Semantic Web. We use the Swoogle Semantic Web search engine to provide ontology terms and ontology correlation statistics, the WordNet lexical database to resolve synonyms, and a practical three step approach to find the most suitable ontology context as well as appropriate ontology terms.
Chinese maximal-length phrases(maximal-length noun phrases and prepositional phrases) possess remarkable linguistic properties.Bidirectional labeling results of Chinese maximal-length phrases obtained using sequential classifiers reveal complementary properties in both directions.In this paper,both left-right and right-left sequential labeling were employed to identify the Chinese maximal-length noun phrases and prepositional phrases.Then a novel fork position based probabilistic algorithm was developed to fuse the bidirectional results.Experiments were carried out on the Penn Chinese Treebank,a segmented,part-of-speech tagged,and fully bracketed corpus.The results confirmed that the proposed algorithm is able to effectively exploit the complementary strengths of the two directions.
In the rapidly growing field of online psychotherapeutic interventions, an increasing number of clinicians are seeking to extend therapeutic interventions into cyberspace. However, because communication with clients in this medium is often devoid of auditory and visual feedback, these clinicians are not able to rely on their clinical observations. It then becomes incumbent to develop a psychometrically and theoretically sound means of assessing emotion and mood states that can be easily utilized in this forum. Utilizing cross-culturally and empirically supported models of emotion structure shown to be influential in the self-report data, the Positive Affect and Negative Affect factors, this study seeks to develop and validate a theoretically and psychometrically sound non-verbal measure of mood state that can easily be used in online interventions. Twenty five mood terms reflecting the range of positive and negative affect were selected from Watson and Tellegen’s (1982) two factor model and a corresponding set of fifty full colored emoticons were generated for the study (two emoticons per mood term) which participants rated on perceived level of positive affect and perceived level of negative affect. Intentional validity was tested by determining whether participants perceived the emoticons as expressing the emotion intended while convergent validity was determined by comparing self-ratings on the developed measure to self-ratings on the PANAS-X. Analysis and examination of participant’s positive and negative affect ratings for each emoticon suggest that participants showed a tendency to perceive emotions, through the emoticons, primarily on the continuum of positivity/negativity and less on the level of intensity or arousal. Further results exploring intentional validity showed that the majority of participants perceived the emoticons as expressing the emotions intended. Results also demonstrated that the emoticons had convergent validity with the PANASX. Implications and limitations of the study are discussed as well as directions for further research.
This dissertation presents research examining the role of contextual patterns, salience, and individual differences in the determination of how much is incidentally remembered from a cognitive task performed during the exploration of a naturalistic outdoor environment. Previous empirical findings suggest that the human mind often selects cues for the storage and retrieval of information based upon a rigid, predetermined hierarchy, frequently disregarding useful contextual cues in favor of features most directly relevant to the information itself. Drawing from environmental psychological principles, factors are outlined that contribute to the salience of contextual cues, the most important of which is the cognitive integration of the context with the observer and the integration of both with the task or mental operation at hand. Such integration is referred to as "contextual integration" and may represent an over-arching schema that serves as a cognitive or affective indicator of personal significance. The first of two reported experiments demonstrated superior memory for contexrually integrated stimuli over those given more rudimentary consideration. The second experiment found changes in memory resulting from an interaction between the type of task performed and the mediating role of a cognitive style known as field-independence. This interaction supports the notion that there are predictable patterns to the cognitive management of contextual information. The effects of these patterns are better accommodated by contextual integration than any single construct such as personal relevance or depth of processing. Furthermore, arousal states and affective ratings of the environment, in Experiments 1 and 2, respectively, showed differential changes in reaction to more or less integrated situations. Conducted almost entirely in a natural environment, the research presented attempts to more closely merge the empirical ideals of the environmental and cognitive areas in psychology.
Taking as its basis a survey of the 20th century Korean lexicon, this paper explores its lexical properties through examination of its lexical character, its extralinguistic background and various lexical aspects, and provides an overview of important achievements in lexical studies through the construction and arrangement of lexical data and through the examination of lexical studies. The major results are as follows: First, three main properties of the 20th century Korean lexicon were found: (1) Although it consists of native words, Chinese words, and foreign words from the West, native words are conspicuous for the motivation process of forming words, and highly developed in terms of symbolic words and sense words. (2) Changes in politics and social structures in the 20th century are reflected both directly and indirectly in the Korean lexicon. (3) Complex aspects have appeared due to the expansion of the lexicon, the mass production of new words including foreign words, differentiation between North Korea and South Korea and between old and young generations, and the appearance of an Internet vocabulary in the lexicon of young generations. Second, three significant achievements in 20th century studies on the Korean lexicon were identified: (1) Following on the construction of a lexical database, standard words were established, dictionaries were written, frequencies of words were examined, and basic words were chosen. (2) The government and various civil organizations have focused on lexical purification, with satisfactory results. (3) Books on lexicology and lexical history were published, research on lexical fields and lexical relations was activated and methodologies for lexical education were explored. Lastly, the 20th century Korean lexicon evolved complex and diverse features in response to the demands of the times. On the one hand, lexical studies during this time showed great development and produced significant results, both in quantity and quality. On the other hand, problems continued to exist in areas such as: (ⅰ) limitations in awareness of the importance of the lexicon; (ⅱ) objectives, targets and methodologies of lexical studies; and (ⅲ) lack of research scholars in this field. These problems remain to be addressed by the 21st Korean lexicon.
Prior COVID-19 infection may elevate activity of the behavioral immune system-the psychological mechanisms that foster avoidance of infection cues-to protect the individual from contracting the infection in the future. Such "adaptive behavioral immunity" may come with psychological costs, such as exacerbating the global pandemic's disruption of social and emotional processes (i.e., pandemic disruption). To investigate that idea, we tested a mediational pathway linking prior COVID infection and pandemic disruption through behavioral immunity markers, assessed with subjective emotional ratings. This was tested in a sample of 734 Mechanical Turk workers who completed study procedures online during the global pandemic (September 2021-January 2022). Behavioral immunity markers were estimated with an affective image rating paradigm. Here, participants reported experienced disgust/fear and appraisals of sickness/harm risk to images varying in emotional content. Participants self-reported on their previous COVID-19 diagnosis history and level of pandemic disruption. The findings support the proposed mediational pathway and suggest that a prior COVID-19 infection is associated with broadly elevated threat emotionality, even to neutral stimuli that do not typically elicit threat emotions. This elevated threat emotionality was in turn related to disrupted socioemotional functioning within the pandemic context. These findings inform the psychological mechanisms that might predispose COVID survivors to mental health difficulties.
This paper presents a novel application of incorporating Alternating Structure Optimization (ASO) to conduct the task of text chunking of Semantic Role Labeling (SRL) in Chinese texts. ASO is a competent linear algorithm based on the theory of multi-task learning. In this paper, by constructing several SRL tasks to constitute a multi-task, we are able to encode the inference obtained by ASO algorithm as additional feature to further boost the performance of the target task employing Conditional Random Fields (CRFs). To our knowledge, our method is the first that incorporates multi-task learning into a statistical model in SRL for Chinese texts. We evaluate our approach on Penn Treebank data sets and obtain encouraging result.
WordNet, a lexical database for English, is organized around semantic and lexical relationships between synsets, concepts represented by sets of synonymous word senses. Offering reasonably comprehensive coverage of the nouns, verbs, adjectives, and adverbs of general English, WordNet is a widely used resource for dealing with the ambiguity that arises from homonymy, polysemy, and synonymy. WordNet is used in many information-related tasks and applications (e.g., word sense disambiguation, semantic similarity, lexical chaining, alignment of parallel corpora, text segmentation, sentiment and subjectivity analysis, text classification, information retrieval, text summarization, question answering, information extraction, and machine translation).
Grammar rules for Clausal Coordinate Ellipsis (CCE) are based nearly exclusively on linguistic judgments (intuitions).For German, the extent to which grammar rules based on this type of empirical evidence generate all and only CCE structures that populate text corpora, has only been explored with the TIGER treebank of written newspaper text.How well these rules fit spoken German is unknown.In this paper, we study the applicability of judgment-based CCE rules to spontaneously spoken German by means of the TBa-D/S treebank, which is based on dialogues for appointment scheduling and travel planning from the VERBMOBIL project.The judgment-based CCE rules are shown to hold nearly equally well for spoken as for written text: The proportion of deviations from the rules are virtually identical-less than 3% of the utterances/sentences that include a clausal coordination (compared to about 1% in the TIGER treebank).Moreover, the relative frequencies in VERBMOBIL of four main CCE types distinguished in the literature reveal a pattern that resembles the pattern observed in CGN2.0, the Corpus of Spoken Dutch.
Annotated data have recently become more important, and thus more abundant, in computational linguistics. They are used as training material for machine learning systems for a wide variety of applications from Parsing to Machine Translation (Quirk et al., 2005). Dependency representation is preferred for many languages because linguistic and semantic information is easier to retrieve from the more direct dependency representation. Dependencies are relations that are defined on words or smaller units where the sentences are divided into its elements called heads and their arguments, e.g. verbs and objects. Dependency parsing aims to predict these dependency relations between lexical units to retrieve information, mostly in the form of semantic interpretation or syntactic structure. Parsing is usually considered as the first step of Natural Language Processing (NLP). To train statistical parsers, a sample of data annotated with necessary information is required. There are different views on how informative or functional representation of natural language sentences should be. There are different constraints on the design process such as: 1) how intuitive (natural) it is, 2) how easy to extract information from it is, and 3) how appropriately and unambiguously it represents the phenomena that occur in natural languages. In this article, a review of statistical dependency parsing for different languages will be made and current challenges of designing dependency treebanks and dependency parsing will be discussed.
The lexical development system OntoNet is introduced, which includes a browser and an editor for the WordNet 3.0 database. The aim of the OntoNet project is to provide a comfortable and up-to-date access to the lexical database for the modification of WordNet or the development of new wordnets.
The Postmodern culture today breaks down historically a solid barrier, produced in the modern age, between the language and its users, signs and realities and subject and his object. This phenomenon brings about some translation problems at the same time; interpretative diversity, linguistic derivation in the mass media, permanent reproduction of translated text and meaning`s continuity, definition of translator etc. I examine these problems through a theoretical approach to the mediative nature of the act of translation. I stress on two inevitable aspects of the postmodern linguistic tendencies: connotation generalized in our translation activities and its re-mediative culture. Focusing on a cycling perspective of our translating activities, I could reach the conclusion that the translated text have no relation with the denotative meanings which have been comprised closed and original with the realities from the modern age. A corrected model could be proposed. I call this cycle model of translating process which comprises ① Mediation or Creation ② Re-mediation or Translation ③ Re-re-mediation or Comprehension ④ Verification, Deduction or Correction. Each activity has not only its own translating process but its circulated role for the time in which everyone can participate equally as a reader, a sender, a translator and an individual who makes its contextual needs and desires. This model could show that the translating process is not for fixing a linguistic sign to some closed meanings but for expanding its pertinent meanings to diverse situations. Translating activity is not for making a linguistic norm by itself, but for making appropriate communication with as much of the population as possible.
Un assunto sempre più condiviso nell’ambito degli studi sull’acquisizione sia di L1 che di L2 è che l’evidenza empirica privilegiata debba essere rappresentata da corpora di produzioni scritte o orali degli apprendenti, estensivamente annotate a molteplici livelli di rappresentazione linguistica. Più in generale, corpora lemmatizzati e annotati a livello morfosintattico fanno ormai parte dello strumentario comune del linguista. Accanto ad essi, si fa però strada l’esigenza di disporre di risorse testuali più sofisticate dal punto di vista delle modalità di esplorazione linguistica, come ad esempio corpora annotati a livello sintattico (le cosiddette treebank). Questi consentono infatti di osservare i processi di convergenza degli apprendenti verso la lingua “obiettivo” anche a livello di specifici tratti grammaticali astratti o di macro-strutture linguistiche.
Generative lexicalized parsing models, which are the mainstay for probabilistic parsing of English, do not perform as well when applied to languages with different language-specific properties such as free(r) word order or rich morphology. For German and other non-English languages, linguistically motivated complex treebank transformations have been shown to improve performance within the framework of PCFG parsing, while generative lexicalized models do not seem to be as easily adaptable to these languages.
Program analysis tools used in software maintenance must be robust and ought to be accurate. Many data-driven parsing approaches developed for natural languages are robust and have quite high accuracy when applied to parsing of software. We show this for the programming languages Java, C/C++, and Python. Further studies indicate that post-processing can almost completely remove the remaining errors. Finally, the training data for instantiating the generic data-driven parser can be generated automatically for formal languages, as opposed to the manually development of treebanks for natural languages. Hence, our approach could improve the robustness of software maintenance tools, probably without showing a significant negative effect on their accuracy.
Computer-aided Acquisition of Semantic Knowledge (CASK) is aimed at describing a number of semantic fields of a few European languages using data mining techniques elaborated within the framework of the new paradigm of computation known as Knowledge Discovery in Databases (KDD). CASK's motivation is to dig deeper in order to find building blocks which could be used in various sophisticated ways. The project is interdisciplinary involving scientific cooperation of experts in linguistics with information engineers. The task of linguists consists in an interactive (computer-aided) discovery of ontology-based definitions of feature structures using the SEMANA (Semantic Analyser) software which was designed especially in order to build linguistic databases with semantic knowledge.