Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
With increasing numbers of Web users, there is a necessity to improve their Web site navigation experience over the Internet and a range of Web applications have emerged recently for this purpose. Many researchers have stressed the importance of identifying semantic relatedness of Web pages in such Web applications as Web site navigation, automatic tour generation and adaptive Web applications. One approach to identifying semantic relatedness between documents is to use lexical databases and lexical chains. For example, an approach using lexical chains has been proposed by Green for identifying paragraph similarity in a document [1]. However, due to the unacceptable length of time needed for lexical chaining and the difficulty of global representation of documents, Green used synset weight vectors to compare semantic relatedness between two documents. But his approach to identifying paragraph similarities can be extended to identify semantic similarities between documents. In this study, an approach to identifying semantic similarity between Web pages incorporating weighted lexical chains (SRWLC) and document properties based on reiteration, density, length and semantic distance is proposed. The two approaches (the proposed approach and Green’s approach - SR Green ) were empirically compared by determining the semantic relatedness of Web pages using human subjects. The research hypothesis of this research is that the proposed approach identifies significantly more semantically related pages with a higher precision than the approach that has been proposed by Green. The null hypothesis is that there is no significant different in identification of semantically related pages between the two approaches. In this context precision is defined as the proportion of retrieved pages that are relevant. Web pages belonging to the Department of Computer Science, Keele University are used for the empirical evaluation of the two methods. The semantic relatedness of all pages was identified using both approaches and a Web-based page categorisation exercise using human subjects was carried out for this empirical evaluation. The evaluation is Web based, and therefore can be carried out on the subjects’ preferred Web browser at his or her preferred time & place. Therefore the distortion effects are minimised and the results of the evaluation are realistic and also can be reliably generalised to some extent. An invitation e-mail was sent to twenty subjects during the first week of October 2004 giving a link and guidelines for them to start the experiment. Once the link on the email was clicked, subjects were shown the initial Web page, giving an introduction to the experiment and instructions on how to continue. Twelve out of twenty invited subjects completed the experiment. Two subjects attempted the experiments but couldn’t finish because of network problems; another three subjects couldn’t finish because of time restrictions, and the other three subjects did not respond at all. Therefore responses from only twelve subjects are used for experimental evaluation. The Wilcoxon signed ranks test returns a p value of 0.004, indicating that the null hypothesis can be rejected, and that there is evidence to suggest that the SR WLC approach identifies significantly more semantically-related pages with higher precision than the SR Green approach. The SRWLC approach should be evaluated further using different Web site contents and language styles (e.g. American/British English). It would be interesting to use more subjects from different backgrounds to do the evaluation. This would determine whether the results of the evaluation are influenced by the human subject’s background, such as their status (student, staff, or other) his familiarity with the pages of the test database, gender differences or the level of English knowledge. Most importantly the approach is believed to be valid for more general browsing environments than a computer science Website and a wider study is desirable.
L’adolescenza è fase di ricerca della propria identità. Con lo scopo di esplorare i propri “mondi possibili”, gli adolescenti producono delle rappresentazioni cognitive di sé nel futuro. Questi ‘sé possibili’ possono rappresentare speranze, timori e aspettative plausibili. Tuttavia, l’esplorazione dell’identità è sempre un processo attivo nel contesto. Per analizzare i processi di esplorazione dell’identità in contesti normativi e non-normativi e per valutare il legame tra processi di esplorazione dell’identità e comportamenti trasgressivi sono stati effettuati due studi. Al primo studio – volto ad indagare il ruolo del contesto e dell’età nella produzione dei sé possibili in contesti normativi e non-normativi - hanno partecipato 105 soggetti di età compresa tra 14 e 19 anni, frequentanti scuole medie superiori di contesti considerati ‘normativi’ (N = 61) e scuole medie superiori di contesti formalmente riconosciuti come ‘a rischio’ (N = 44); ha, inoltre, partecipato alla ricerca un più piccolo numero di soggetti ‘devianti’ (N=10) di età compresa tra 16 e 18 anni. Gli strumenti utilizzati: questionario Possible Selves Questionnaire open-ended di Oyserman e Markus (1990), questionario Self -Perception Profile for adolescent (Harter, 1985). Per i soggetti devianti sono state utilizzate produzioni grafiche e un’intervista autobiografica. E’ stata seguita una duplice procedura di analisi: quantitativa e qualitativa (categoriale e lessicale). I risultati mostrano un effetto dell’età e del contesto nella produzione dei sé possibili, e, in particolare, viene evidenziato il ruolo della dimensione del sé temuto nella formazione dell’identità dei soggetti ‘ a rischio’. Nel secondo studio, si ipotizza che la trasgressione (intesa come rottura della norma e dei vincoli) possa rappresentare una forma di esplorazione con differenze di tipologie e funzioni dei comportamenti trasgressivi in rapporto alla fase adolescenziale di riferimento. A 90 soggetti (adolescenti appartenenti a tre fasce di età corrispondenti a tre livelli di scolarità: primo e ultimo amnio di scuola superiore e secondo anno di università) è stata proposta una traccia narrativa volta ad indagare la funzione, la struttura ed il significato della trasgressione. Il corpus narrativo raccolto è stato sottoposto ad un duplice livello di analisi: contenutistico-categoriale, cui sono seguite analisi statistiche ad hoc e lessicale. I risultati emersi confermano la funzione esplorativa dei comportamenti soggettivamente percepiti come trasgressivi e ne evidenziano una funzione evolutiva, specificandone peculiarità e differenze in funzione dell’età dei soggetti / [ENGLISH] During adolescence the subject is in search of his/her identity. In order to explore identity, adolescents produce cognitive representations of themselves in the future. These future–oriented possible selves can represent future, expected selves or feared selves The exploration of possible selves is a process in context, and, we suppose, it differs in relation to the peculiar contexts within which the adolescent is integrated. To analyse the exploration identity processes in normative and non-normative contexts and to test the relation between identity exploration and trasgressive behaviours, we made two studies. In the first research - to explore the context and the age effects on possibile selves production in different contexts - a total of 105 youths between the ages of 14 and 18 were analysed. Youths were drawn from three sub-samples distinguished by their belonging contexts (normative at risk and deviant). Measures: Possible Selves Questionnaire open-ended (Oyserman e Markus, 1990), Self -Perception Profile for adolescent (Harter, 1985); autobiographic interview (for deviant subjects). The data collected has been assessed at two levels of analysis: quantitative and qualitative (categorical and lexical analysis of textual content). The results show that there is a context and age role on the identity exploration: in particular the exploration of the at risk adolescents is characterized by the feared self. In the second study, we assume that transgression (that is to say breaking norms and limits) can represent a form of identity exploration with differences of types and functions in the transgressive behaviours in relation to the developmental phase. We proposed a narrative task to 90 adolescents (of three different age groups and school levels: first and last year at the high school and second year at the university). The aim was to investigate the function, the structure and the meaning of transgression. The textual corpus collected has been assessed at two levels of analysis: categorical content and lexical content. The results showed the exploratory and developmental functions of the behaviours, which the adolescents perceived as transgression, and highlighted some peculiarities in relation to the age.
This paper focuses on the electronic literacy practices of two Korean-American heritage language learners who manage Korean weblogs.Online users deliberately alter standard forms of written language and play with symbols, characters, and words to economize typing effort, mimic oral language, or convey qualities of their linguistic identity such as gender, age, and emotional states.However, little is known about the impact of computer-mediated nonstandard language use on heritage learners' linguistic development.Through in-depth case studies of two siblings, the study examines the linguistic and pragmatic practices of these learners online and the perceived effects of non-standard forms of computer-mediated language on their heritage language development and maintenance.The data show that electronic literacy practices provide authentic opportunities to use the language and support the development of a social network of Korean speakers, which results in greater sociopsychological attachment to the Korean language and culture.The informants report that the deviant language forms found in e-texts enable them to engage in online interactions without the pressures of having to spell the words correctly.However, they express frustrations in not being able to distinguish between correct and non-standard forms of the language, which appear to be affecting their offline language use. THE KOREAN CONTEXTThe Republic of Korea has one of the fastest-growing cybercommunities in the world.According to the Korea Network Information Center, over 63% of the entire South Korean population are Internet users, and 95% of individuals in the 6-29 age bracket report using it on a daily basis.Internet sites that enable users to create "personal spaces" to share and document their changing lives and keep connected with people they know are immensely popular among Koreans.A case in point is "Cyworld," an upgraded blog that features chatting, commentaries, pictures, music, a guest book, avatars and links to other homepages prompting users to network with their friends, family, and colleagues.As of August 2005, there are over 11 million Cyworld registered users.Participation in online forums such as Cyworld engages its members in a social process of learning through shared practices, internally constructed membership, and the formation of personal and group identities (Holmes & Jin Sook Lee Electronic literacy and heritage language maintenance Language Learning & Technology 94Myerhoff, 1999).Members are involved in a community of practice, where a group of people who come together around a joint enterprise develop common beliefs, values, and ways of doing things, which all influence the ways in which members communicate with one another (Eckert, 2000;Wenger, 1998).New forms of expression are constantly being negotiated and shared among online users, making it difficult to keep current with the changing face of electronic text.Computer-mediated communication is unique in that, despite its similarities to oral speech, it invites substantial deregulation effects on communication, which can foster the use of creative, non-standard language play (Sproull & Kiesler, 1986).Studies have documented non-standard 1 uses of language in online interactions (a) to mark certain individual characteristics such as provincial dialects, social class, gender, age, and/or personality traits, (b) to economize typing efforts, and/or (c) to mimic spoken language (Barnes, 2003;Herring, 2001;Song, 2002;Sproull & Kiesler, 1986).For example, Su ( 2004) found an emergent mock Taiwanese accent among Internet users as a form of language play to jointly construct "a young, lively, congenial, and witty presence" (p.61).Androutsopoulos ( 2000) also revealed that non-standard orthography in online fan media texts was representative of spoken language and purely graphemic modifications, which are used to serve as contextualization cues and cues of subcultural positioning.Although all natural languages inevitably change over time, drastic deviances from standard language ranging from non-standard orthography and incorrect grammar to unfamiliar lexical items and symbols have brought forth great concern about the preservation of standard orthography, grammar, and pragmatic uses of the Korean language (Choi, 2003;Kim, 2005;Park, 1989).Educators across grade levels in Korea are reporting that students display electronic textual features in their school work: they have difficulty with spelling and with the proper word spacing used to delineate word boundaries due to non-standard ways of Internet language use, which flout conventional norms of literacy practices (Ahn, 2000;Choi, 2003;Kim, 2005;Noh, 2000).For young children and Korean as foreign/second language learners who have not fully acquired literacy in the language, exposure to electronic texts may have adverse effects on their language development.However, Meskill, Mossop, and Bates (1999) state that "children in the age of electronic text are developing unique skills and strategies for inventing novel forms of understanding these texts that are quite often independent of formal instructional ('school') literacy training" (p.4), thus, highlighting the positive ways in which the development of electronic texts can benefit students' cognitive flexibility and skills.
Corpora annotated with structural and linguistic characteristics play a major role in nearly every area of language processing. During recent years a number of corpora and large data sets became known and available to research even in specialized fields such as medicine, but still however, targeted predominantly for the English language. This paper provides a description of the collection, encoding and linguistic processing of an ever growing Swedish medical corpus, the MEDLEX Corpus. MEDLEX consists of a variety of text-documents related to various medical text genres. The MEDLEX Corpus has been structurally annotated using the Corpus Encoding Standard for XML (XCES), lemmatized and automatically annotated with part-of-speech and semantic information (extended named entities and the Medical Subject Headings, MeSH, terminology). The results from the processing stages (part-of-speech, entities and terminology) have been merged into a single representation format and syntactically analysed using a cascaded finite state parser. Finally, the parser’s results are converted into a tree structure that follows the TIGER-XML coding scheme, resulting a suitable for further exploration and fairly large Treebank of Swedish medical texts. 1.
This paper investigates how to extend coverage of a domain independent lexicon tailored for natural language understanding. We introduce two algorithms for adding lexical entries from VerbNet to the lexicon of the Trips spoken dialogue system. We report results on the efficiency of the method, discussing in particular precision versus coverage issues and implications for mapping to other lexical databases.
Objective:This study aimed to examine the effects of haloperidol and amphetamine on human startle response modulated by emotionally-toned film clips. Method: Sixty participants, in two groups (one receiving haloperidol and the other receiving amphetamine) were tested using electromyography (EMG) to measure eye-blink muscle (orbicular oculi) while different emotions were induced by six 2-minute film clips. Results: An affective rating shows the negative and positive effects of the two drugs on emotional reactivity, neither amphetamine nor haloperidol had any impact on the modulation of the startle response. Conclusion: The methodological and theoretical aspects of the study and findings will be discussed.
The databases record instances of deponency, which is the term we have adopted to describe mismatches between morphology and morphosyntax. The prototypical example are the deponent verbs of Latin, which involve a mismatch between passive form and active meaning. That is, a normal Latin verb had active forms such as amō 'I love' and amāvī 'I have loved', which contrasted with the passive forms amor 'I am loved' and amātus sum 'I have been loved' (in this case, with a masculine subject). A deponent verb, on the other hand, looks like the passive but functions like the active, as in mīror 'I admire', mīrātus sum 'I have admired'. In the the databases we construe deponency in an extended fashion, covering any mismatch between the apparent morphosyntactic value of a morphological form and its actual value in a given syntactic context. Two databases are housed on this site, accessible through the links above. The cross-linguistic database looks at the presence of morphological mismatches in a controlled sample of genetically and geographically diverse languages (based on the 100-language sample from the World Atlas of Language Structures). The typological database records the logical space of deponency: what features may be affected, and what are the characteristics of the resulting paradigm? Every logical combination of parameters is represented by one exemplar (or where none has been found, this is noted too). The typological database is supplemented by a set of formal analyses of examples which hold particular interest for morphological theory.
Group identifications and intergroup relations among Turkish Dutch respondents What determines group identification processes among ethnic minority groups and how are these processes related to in-group and out-group evaluations? This article focuses on Turkish and Dutch identification among Turkish Dutch respondents and their feelings towards different ethnic and religious groups. The results show that Turkish identification is strong and Dutch identification rather weak, and that both group identifications are not strongly associated. Perceived socio-structural characteristics of intergroup relations (stability, legitimacy, permeability, and discrimination) affected both Turkish and Dutch identification. Group identification was positively related to (ethnic and religious) in-group evaluation, but there were few relationships with out-group evaluations. The affective ratings of Moroccans, Antilleans, Jews and non-believers were quite negative.
In this paper we describe the current state of a new Japanese lexical resource: the Hinoki treebank. The treebank is built from dictionary definition sentences, and uses an HPSG based Japanese grammar to encode both syntactic and semantic information. It is combined with an ontology based on the definition sentences to give a detailed sense level description of the most familiar 28,000 words of Japanese.
This paper presents an approach to dependency parsing which can utilize any standard machine learning (classification) algorithm. A decision list learner was used in this work. The training data provided in the form of a treebank is converted to a format in which each instance represents information about one word pair, and the classification indicates the existence, direction, and type of the link between the words of the pair. Several distinct models are built to identify the links between word pairs at different distances. These models are applied sequentially to give the dependency parse of a sentence, favoring shorter links. An analysis of the errors, attribute selection, and comparison of different languages is presented.
Data-driven grammatical function tag assignment has been studied for English using the Penn-II Treebank data. In this paper we address the question of whether such methods can be applied successfully to other languages and treebank resources. In addition to tag assignment accuracy and f-scores we also present results of a task-based evaluation. We use three machine-learning methods to assign Cast3LB function tags to sentences parsed with Bikel's parser trained on the Cast3LB treebank. The best performing method, SVM, achieves an f-score of 86.87% on gold-standard trees and 66.67% on parser output - a statistically significant improvement of 6.74% over the baseline. In a task-based evaluation we generate LFG functional-structures from the function-tag-enriched trees. On this task we achive an f-score of 75.67%, a statistically significant 3.4% improvement over the baseline.
On the norm of current Chinese characters,象 像 of non-noun mainly the verb lexical meaning have not got the ideal social effect yet in the division and usage.And the present state is ambiguous.By reviewing the differences in non-noun lexical meaning and evidence of their division,it is thought that only divided their usage clearly can they be used normally.
The Hamburg implementation of the Weighted Constraint Dependency Grammar formalism (WCDG) includes an example grammar with comprehensive coverage for written German. This manual is the annotation guideline that was used to define the goals of the grammar and to create the Hamburg Dependency Treebank also published in the course of this project.
Recently proposed deterministic classifier-based parsers (Nivre and Scholz, 2004; Sagae and Lavie, 2005; Yamada and Mat-sumoto, 2003) offer attractive alternatives to generative statistical parsers. Deterministic parsers are fast, efficient, and simple to implement, but generally less accurate than optimal (or nearly optimal) statistical parsers. We present a statistical shift-reduce parser that bridges the gap between deterministic and probabilistic parsers. The parsing model is essentially the same as one previously used for deterministic parsing, but the parser performs a best-first search instead of a greedy search. Using the standard sections of the WSJ corpus of the Penn Treebank for training and testing, our parser has 88.1% precision and 87.8% recall (using automatically assigned part-of-speech tags). Perhaps more interestingly, the parsing model is significantly different from the generative models used by other well-known accurate parsers, allowing for a simple combination that produces precision and recall of 90.9% and 90.7%, respectively.
We study unsupervised methods for learning refinements of the nonterminals in a treebank. Following Matsuzaki et al. (2005) and Prescher (2005), we may for example split NP without supervision into NP[0] and NP[1], which behave differently. We first propose to learn a PCFG that adds such features to nonterminals in such a way that they respect patterns of linguistic feature passing: each node's nonterminal features are either identical to, or independent of, those of its parent. This linguistic constraint reduces runtime and the number of parameters to be learned. However, it did not yield improvements when training on the Penn Treebank. An orthogonal strategy was more successful: to improve the performance of the EM learner by treebank preprocessing and by annealing methods that split nonterminals selectively. Using these methods, we can maintain high parsing accuracy while dramatically reducing the model size.
Abstract Facial masculinity may be used as a cue in female mate choice, as it reflects the success of the male genotype in its developmental environment. Women may maximize reproductive success by using a conditional strategy favoring highly masculine facial features for short‐term relationships and feminized facial features in men for long‐term relationships. Three studies examine reactions to masculinized and feminized male facial composites. Properties of the original composite image affect ratings of critical attributes and the magnitude of the differences in ratings between versions undergoing identical processes of geometric manipulation (Study 1). Both men and women attribute personality, behavior, and mating strategies consistent with predictions derived from the good genes and mating trade‐off hypotheses (Study 2). Participants accurately grouped behavioral tendencies related to high mating effort/risky strategies and high parenting effort/risk adverse strategies and associated mating effort more so with masculinized faces and parenting effort more so with feminized faces (Study 3). These results indicate that male facial masculinity serves as a visual cue for inferring personality and reproductive strategy.
espanolEl principal objetivo de este articulo es mostrar que el teatro popular en Camerun desde los anos 80 ha venido forjandose una subversiva dentro de un contexto de crisis multisectorial en que la calle se ha convertido en el nuevo terreno y laboratorio de expresion. Dicta sus leyes y moldea nuevas normas, mas alla de los limites de lo convencional. Asi que el teatro participa en la creacion de un espacio critico a partir del cual las capas populares afrontan los resortes de la dominacion y le imponen nuevos margenes. De ahi el surgimiento de una lengua de gozo que atestigua una norma linguistica fluctuante, dentro de una palabra teatralizante en la que el frances se hunde en una mezcla heterogenea hecha de innovaciones o de desviaciones diversas. EnglishThe main concern of this paper is to show that popular theatre in Cameroon since 1980 shas formed a subversive in a context of multisector-based crisis where the has become an area and laboratory of expression. This street dictates and build up its laws, shapes new norms beyond conventional limits. Thus, theatre develops a critical space from which popular strata faces springs of domination and impose on its new margins. Hence, the appearance of a pleasure language that testifies to a fluctuating linguistic norm in a dramatiozing word where French is incorporated into an heterogeneous mixture made up of various innovations and diversions. francaisLe principal objectif de cet article est de montrer que le theâtre populaire au Cameroun depuis les annees 80 s'est forge une subversive dans un contexte de crise multisectorielle ou la rue est devenue le nouveau terrain et laboratoire d'expression. Celle-ci dicte ses lois et faconne de nouvelles normes, au-dela des limites du conventionnel. Le theâtre procede ainsi a l'amenagement d'un espace critique a partir duquel les couches populaires affrontent les r essorts de la domination et lui imposent de nouvelles marges. D'ou l'emergence d'une langue de jouissance qui temoigne d'une norme linguistique fluctuante, dans une parole theâtralisante ou le francais est englue dans un melange composite fait d'innovations ou de deviations diverses.
The PropBank primarily adds semantic role labels to the syntactic constituents in the parsed trees of the Treebank. The goal is for automatic semantic role labeling to be able to use the domain of locality of a predicate in order to find its arguments. In principle, this is exactly what is wanted, but in practice the PropBank annotators often make choices that do not actually conform to the Treebank parses. As a result, the syntactic features extracted by automatic semantic role labeling systems are often inconsistent and contradictory. This paper discusses in detail the types of mismatches between the syntactic bracketing and the semantic role labeling that can be found, and our plans for reconciling them.
Mediated moderation occurs when the interaction between two variables affects a mediator, which then affects a dependent variable. In this article, we describe the mediated moderation model and evaluate it with a statistical simulation using an adaptation of product-of-coefficients methods to assess mediation. We also demonstrate the use of this method with a substantive example from the adolescent tobacco literature. In the simulation, relative bias (RB) in point estimates and standard errors did not exceed problematic levels of ±10%, although systematic variability in RB was accounted for by parameter size, sample size, and nonzero direct effects. Power to detect mediated moderation effects appears to be severely compromised under one particular combination of conditions: when the component variables that make up the interaction terms are correlated and partial mediated moderation exists. Implications for the estimation of mediated moderation effects in experimental and nonexperimental research are discussed.
In this paper, we describe the SALTO tool. It was originally developed for the annotation of semantic roles in the frame semantics paradigm, but can be used for graphical annotation of treebanks with general relational information in a simple drag-and-drop fashion. The tool additionally supports corpus management and quality control. 1.
Detecting idioms in a sentence is important to sentence understanding. This paper discusses the linguistic knowledge for idiom detection. The challenges are that idioms can be ambiguous between literal and idiomatic meanings, and that they can be “transformed” when expressed in a sentence. However, there has been little research on Japanese idiom detection with its ambiguity and transformations taken into account. We propose a set of linguistic knowledge for idiom detection that is implemented in an idiom dictionary. We evaluated the linguistic knowledge by measuring the performance of an idiom detector that exploits the dictionary. As a result, more than 90% of the idioms are detected with 90% accuracy.
An HMM-based single character recovery (SCR) model is proposed in this paper to extract a large set of atomic abbreviations and their full forms from a text corpus. By an “atomic abbreviation,” it refers to an abbreviated word consisting of a single Chinese character. This task is important since Chinese abbreviations cannot be enumerated exhaustively but the abbreviation process for compound words seems to be compositional. One can often decode an abbreviated word character by character to its full form. With a large atomic abbreviation dictionary, one may be able to handle multiple character abbreviation problems more easily based on the compositional property of abbreviations.
Several recent Information Extraction (IE) systems have been restricted to the identification facts which are described within a single sentence. It is not clear what effect this has on the difficulty of the extraction task or how the performance of systems which consider only single sentences should be compared with those which consider multiple sentences. This paper compares three IE evaluation corpora, from the Message Understanding Conferences, and finds that a significant proportion of the facts mentioned therein are not described within a single sentence. Therefore systems which are evaluated only on facts described within single sentences are being tested against a limited portion of the relevant information in the text and it is difficult to compare their performance with other systems. Further analysis demonstrates that anaphora resolution and world knowledge are required to combine information described across multiple sentences. This result has implications for the development and evaluation of IE systems.
Asian language processing presents formidable challenges to achieving multilingualism and multiculturalism in our society. One of the first and most obvious challenges is the multitude and diversity of languages: more than 2,000 languages are listed as languages in Asia by Ethnologue (Gordon 2005), representing four major language families: Austronesian, Trans-New Guinea, Indo-European, and Sino-Tibetan. 1The challenge is made more formidable by the fact that as a whole, Asian languages range from the language with most speakers in the world (Mandarin Chinese, close to 900 million native speakers) to the more than 70 nearly extinct languages (e.g. Pazeh in Taiwan, one speaker). As a result, there are vast differences in the level of language processing capability and the number of sharable resources available for individual languages. Major Asian languages such as Mandarin Chinese, Hindi, Japanese, Korean, and Thai have benefited...
This paper describes methods and tools used for the post-annotation checking of Prague Dependency Treebank 2.0 data. The annotation process was complicated by many factors: for example, the corpus is divided into several layers that must reflect each other; the annotation rules changed and evolved during the annotation process; some parts of the data were annotated separately and in parallel and had to be merged with the data later. The conversion of the data from the old format to a new one was another source of possible problems besides omnipresent human inadvertence. The checking procedures are classified according to several aspects, e.g. their linguistic relevance and their role in the checking process, and prominent examples are given. In the last part of the paper, the methods are compared and scored.
WordNet, a lexical database for English that is extensively used by computational linguists, has not previously distinguished hyponyms that are classes from hyponyms that are instances. This note describes an attempt to draw that distinction and proposes a simple way to incorporate the results into future versions of WordNet.
The movements of newborns have been thoroughly studied in terms of reflexes, muscle synergies, leg coordination, and target-directed arm/hand movements. Since these approaches have concentrated mainly on separate accomplishments, there has remained a clear need for more integrated investigations. Here, we report an inquiry in which we explicitly concentrated on taking such a perspective and, additionally, were guided by the methodological concept of home base behavior, which Ilan Golani developed for studies of exploratory behavior in animals. Methods from nonlinear dynamics, such as symbolic dynamics and recurrence plot analyses of kinematic data received from audiovisual newborn recordings, yielded new insights into the spatial and temporal organization of limb movements. In the framework of home base behavior, our approach uncovered a novel reference system of spontaneous newborn movements.
Computational Modeling of Bilingualism Symposium Organizer: Ping Li (pli@richmond.edu) Department of Psychology, University of Richmond Richmond, VA 23173 USA connectionist developmental lexical model (Li et al., 2004). It considers learner variables (e.g., time of L2 learning and proficiency) and input variables (word types and bilingual distance) to assess determinants of bilingual lexical acquisition. It examines the time course of acquisition, the emergence of structured lexical representations in L1 and L2, and the effect of learning history on learning plasticity. The model attempts to account for important processes such as competition, the extent to which the two lexicons compete for resources in the lexical space over time; entrenchment, the extent to which lexical structures are consolidated in L1 affects the learning of L2, and vice versa; and plasticity, the extent to which structural consolidation of L1 impacts the learning of L2. In the final talk Michael Thomas will review the recent application of connectionist models of language processing to bilingualism. Connectionist models have frequently appealed to two different architectures, localist interactive activation models and distributed processing models. These architectures have been used to explore different phenomena within bilingual language processing, including localist models of visual and auditory word recognition and distributed models of lexical and syntactic acquisition. Recent approaches employing self-organization attempt to bridge the two types of model. The range of existing models of monolingual language processing suggest clear avenues for future bilingual research to pursue, in particular focusing on dynamic aspects of (1) bilingual acquisition, (2) gradual changes in language dominance, (3) real-time switching between languages, (4) bilingual aphasia and recovery, and (5) language decay. Finally, Thomas will conclude with his recent modeling work exploring critical periods and their implications for second language acquisition. Ping Li will give an introduction and overview of the symposium at the beginning, and Yasuhiro Shirai will provide an integrative discussion at the end. Introduction Computational modeling and bilingualism have had until recently only limited interactions (see reviews in French & Jacquet, 2004; Hernandez, Li, & MacWhinney, 2005; Li & Farkas, 2002; Thomas & van Heuven, 2005). Bilingualism has been the norm rather than the exception in our globalized world, but the acquisition of two languages entails significant complexity that challenges empirical methodology. Computational modeling, because of its flexibility in parameter variation and hypothesis testing, is ideally suited for identifying mechanisms underlying bilingual language acquisition and representation. In this symposium, we propose to integrate current computational studies of bilingualism and second language acquisition. Summary of Presentations Robert French will begin by reviewing the state of the art in the study of bilingual lexical memory, pointing out crucial issues in the field. He will then present the BSRN, a bilingual simple recurrent network model. The model learns both English and French NVN strings, intermixed at the sentence level for the two languages. The simulations show that BSRN can develop distinct representations not only for individual lexical categories in each language (as in Elman, 1990), but also for the two languages in general. Thus, the model can display distinct behaviors for the bilingual’s two lexicons without invoking separate mechanisms for each language, providing evidence to the idea of “single mechanism, variable representations” from bilingualism. In the second talk Curt Burgess will present the bilingual HAL model. Using language co-occurrences to model language or memory raises a number of controversial issues on the nature of lexical and semantic representations. In addition, using lexical co-occurrence to model bilingualism introduces crucial theoretical considerations since it is incumbent on the model to account for the transformation of the different lexical codes of multiple languages into similar semantic representations. On the surface this may seem straightforward. However, since co-occurrence models function (at some point in the encoding process) by counting the number of co-occurrences between specific lexical items, one has to provide an account of how the co- occurrence vectors for L1 can merge or co-exist with the vectors for L2. Burgess will discuss this memory consolidation process, a step that is not required in high- dimensional models that encodes only one language. In the third talk Ping Li will present a self-organizing neural network model that simulates developmental stages of the bilingual lexicon. The model is based on DevLex, a References French, R., & Jacquet, M. (2004). Understanding bilingual memory: models and data. Trends in Cognitive Sciences, Hernandez, A., Li, P., & MacWhinney, B. (2005). The emergence of competing modules in bilingualism. Trends in Cognitive Sciences, 9, 220-225. Li, P., & Farkas, I. (2002). A self-organizing connectionist model of bilingual processing. In R. Heredia & J. Altarriba (eds.), Bilingual sentence processing. Elsevier. Thomas, M., & van Heuven, W. (2005). Computational models of bilingual comprehension. In J. F. Kroll & A. de Groot (eds.) Handbook of Bilingualism: Psycholinguistic Approaches. Oxford University Press.
The aim of this research was to study the influence of both the emotional content and the physical characteristics of affective stimuli on the psychophysiological, behavioral and cognitive indexes of the emotional response. We selected 54 pictures from the IAPS, depicting unpleasant, neutral, and pleasant contents, and used two picture sizes as experimental conditions (120 x 90 cm and 52 x 42 cm). Sixty-one subjects were randomly assigned to each experimental condition. We recorded the startle blink reflex, skin conductance response, heart rate, free viewing time, and picture valence and arousal ratings. In line with previous research (e.g., Bradley, Codispoti, Cuthbert, and Lang, 2001), our data showed an effect of the affective content on all the measurements recorded. Importantly, effects of the size of the affective pictures on emotional responses were not found, indicating that the emotional content is more important than the formal properties of the stimuli in evoking the emotional response.
While the processing of verbal and psychophysiological indices of emotional arousal have been investigated extensively in relation to the left and right cerebral hemispheres, it remains poorly understood how both hemispheres normally function together to generate emotional responses to stimuli. Drawing on a unique sample of nine high-functioning subjects with complete agenesis of the corpus callosum (AgCC), we investigated this issue using standardized emotional visual stimuli. Compared to healthy controls, subjects with AgCC showed a larger variance in their cognitive ratings of valence and arousal, and an insensitivity to the emotion category of the stimuli, especially for negatively-valenced stimuli, and especially for their arousal. Despite their impaired cognitive ratings of arousal, some subjects with AgCC showed large skin-conductance responses, and in general skin-conductance responses discriminated emotion categories and correlated with stimulus arousal ratings. We suggest that largely intact right hemisphere mechanisms can support psychophysiological emotional responses, but that the lack of interhemispheric communication between the hemispheres, perhaps together with dysfunction of the anterior cingulate cortex, interferes with normal verbal ratings of arousal, a mechanism in line with some models of alexithymia.
WordNet, a lexical database for English that is extensively used by computational linguists, has not previously distinguished hyponyms that are classes from hyponyms that are instances. This work describes an attempt to draw this distinction and reports the way in which the results were incorporated in the last version (2.1) of WordNet.
While syntactically annotated corpora known as treebanks have been available for many years, along with a variety of customized tools for querying these annota-
Studies on attribution in the moral domain often involve the use of specific behavior examples. To make valid comparisons across trait dimensions (such as honesty and friendliness), it is important to equate the intensities of the specific behaviors used. Pretesting specific behaviors can be a costly effort, but it is often necessary for research in social psychology. Our study provides a rich source of such pretested behaviors. Positive and negative examples of behaviors in the categories of honesty, loyalty, friendliness, charitableness, and cooperativeness were solicited from participants and then rated on the relevant trait dimension by an independent group. The result is data representing rankings, raw scores, andz-scores in an index of 500 behaviors across 10 trait categories that can be used by researchers to study moral and immoral behaviors. The full index of behaviors is available at www .psychonomic.org/archive/.
The paper presents new lexicon of verb valencies for the Czech language named VerbaLex. VerbaLex is based on three valuable language resources for Czech, three independent electronic dictionaries of verb valency frames. The first resource, Czech WordNet valency frames dictionary, was created during the Balkanet project and contains semantic roles and links to the Czech WordNet semantic network. The other resource, VALLEX 1.0, is a lexicon based on the formalism of the Functional Generative Description (FGD) and was developed during the Prague Dependency Treebank (PDT) project. The third source of information for VerbaLex is the syntactic lexicon of verb valencies denoted as BRIEF, which originated at FI MU Brno in 1996. The resulting lexicon, VerbaLex, comprehends all the information found in these resources plus additional relevant information such as verb aspect, verb synonymity, types of use and semantic verb classes based on the VerbNet project.
Previously, we introduced a new computational tool for nonlinear curve fitting and data set exploration: the Naturalistic University of Alberta Nonlinear Correlation Explorer (NUANCE) (Hollis & Westbury, 2006). We demonstrated that NUANCE was capable of providing useful descriptions of data for two toy problems. Since then, we have extended the functionality of NUANCE in a new release (NUANCE 3.0) and fruitfully applied the tool to real psychological problems. Here, we discuss the results of two studies carried out with the aid of NUANCE 3.0. We demonstrate that NUANCE can be a useful tool to aid research in psychology in at least two ways: It can be harnessed to simplify complex models of human behavior, and it is capable of highlighting useful knowledge that might be overlooked by more traditional analytical and factorial approaches. NUANCE 3.0 can be downloaded from the Psychonomic Society Archive of Norms, Stimuli, and Data at www.psychonomic.org/archive.
Web searchers reformulate their queries, as they adapt to search engine behavior, learn more about a topic, or simply correct typing errors. Automatic query rewriting can help user web search, by augmenting a user’s query, or replacing the query with one likely to retrieve better results. One example of query-rewriting is spell-correction. We may also be interested in changing words to synonyms or other related terms. For Japanese, the opportunities for improving results are greater than for languages with a single character set, since documents may be written in multiple character sets, and a user may express the same meaning using different character sets. We give a description of the characteristics of Japanese search query logs and manual query reformulations carried out by Japanese web searchers. We use characteristics of Japanese query reformulations to extend previous work on automatic query rewriting in English, taking into account the Japanese writing system. We introduce several new features for building models resulting from this difference and discuss their impact on automatic query rewriting. We also examine enhancements in the form of rules which block conversion between some character sets, to address Japanese homophones. The precision/recall curves show significant improvement with the new feature set and blocking rules, and are often better than the English counterpart.
This special issue on Data Resources, Evaluation, and Dialogue Interaction is based on five thoroughly revised and extended papers from the sixth SIGdial Workshop held in Lisbon, Portugal, in September 2005. SIGdial is a special interest group on discourse and dialogue whose parent organisations are the Association for Computational Linguistics (ACL) and the International Speech Communication Association (ISCA). SIGdial workshops accommodate a broad range of topics related to discourse and dialogue. Among these topics are data resources, evaluation, and dialogue interaction. The papers selected for this special issue have in common that they all deal with aspects of these topics and each paper has its focus on at least one of them.
This paper describes a methodology aimed at grouping Catalan verbs according to their syntactic behavior. Our goal is to acquire a small number of basic classes with a high level of accuracy, using minimal resources. Information on syntactic class, expensive and slow to compile by hand, is useful for any NLP task requiring specific lexical information. We show that it is possible to acquire this kind of information using only a POS-tagged corpus. We perform two clustering experiments. The first one aims at classifying verbs into transitive, intransitive and verbs alternating with a se-construction. Our system achieves an average 0.84 F-score, for a task with a 0.33 baseline. The second experiment aims at further distinguishing among pure intransitives and verbs bearing a prepositional object. The baseline for the task is 0.51 and the upperbound 0.98. The system achieves an average 0.88 F-score.
We introduce Talbanken05, a Swedish treebank based on a syntactically annotated corpus from the 1970s, Talbanken76, converted to modern formats. The treebank is available in three different formats, besides the original one: two versions of phrase structure annotation and one dependency-based annotation, all of which are encoded in XML. In this paper, we describe the conversion process and exemplify the available formats. The treebank is freely available for research and educational purposes. 1.
Traditionally, orthographic variants have been modelled as different ways of spelling the same word - described at the level of the lexeme. But when inflection is taken into account, this runs into a problem: different citation forms have different inflectional paradigm - and orthographic variation does not merely affect the citation form, but the entire paradigm. The MorDebe database therefore models orthographic variation as a relation between distinct, yet still token-identical lexemes. This paper discusses the advantage of that approach, and the full set of practical problems that arose during the structural treatment of orthographic variation in the MorDebe database.
Treebank data have been utilized as data sources for a wide range of tasks in computational linguistics, including statistical parsing, anaphora resolution, induction of valence lexica, etc. More recently, researchers have experimented with extracting semantic information from syntactically annotated data. Here, treebank data
This paper describes a parser which generates parse trees with empty elements in which traces and fillers are co-indexed. The parser is an unlexicalized PCFG parser which is guaranteed to return the most probable parse. The grammar is extracted from a version of the PENN treebank which was automatically annotated with features in the style of The annotation includes GPSG-style slash features which link traces and fillers, and other features which improve the general parsing accuracy. In an evaluation on the PENN treebank Its results for the empty category prediction task and the trace-filler coindexation task exceed all previously reported results with 84.1% and 77.4% fscore, respectively.
Standard techniques used in multilingual terminology management fail to describe legal terminologies as they are bound to different legal systems and terms do not share a common meaning. In the LexALP project, we use a technique defined for general lexical databases to achieve cross language interoperability between languages of the Alpine Convention. In this paper we present the methodology and tools developed for the collection, description and harmonisation of the legal terminology of spatial planning and sustainable development in the four languages of the countries of the Alpine Space.