Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The aim of this paper is to illustrate the potential of a parallel corpus in the context of (computer-assisted) language learning. In order to do so, we propose to answer two main questions (1) what corpus (data) to use and (2) how to use the corpus (data). We provide an answer to the what-question by describing the importance and particularities of compiling and processing a corpus for pedagogical purposes. In order to answer the how-question, we first investigate the central concepts of the interactionist theory of second language acquisition: comprehensible input, input enhancement, comprehensible output and output enhancement. By means of two case studies, we illustrate how the abovementioned concepts can be realized in concrete corpus-based language learning activities. We propose a design for a receptive and productive language task and describe how a parallel corpus can be at the basis of powerful language learning activities. The Dutch Parallel Corpus, a ten-million word sentence aligned and annotated parallel corpus, is used to develop these language tasks.
Abstract: In given clause the attempt is done to study a problem of translation of synonyms depending on their kinds and classification. The various synonyms require the various approaches while translating is proved. And also depending on kinds of synonyms the interpreter chooses various translational, more precisely lexical transformations. Relations of languages define quantity of use of translational transformations. In given clause the term "synonym " is understood in wide meaning. As synonyms can be both word and word collocation, it is possible assume at a grammatical level can be and synonymous phrases, of time there are phrases and it is possible about «the synonymous offers ". The synonymous offers this same, that the offers with one by those by meanings, but used in various functional styles. Synonymous words and word collocations used in various functional styles in this case are considered. Synonymous words and word collocations used only in art style in this case are considered. The essence of use and translation of synonyms is opened.
Neuropsychological diagnostic tests of visual perception mostly assess high-level processes like object recognition. Object recognition, however, relies on distinct mid-level processes of perceptual organization that are only implicitly tested in classical tests. The Leuven Perceptual Organization Screening Test (L-POST) fills a gap with respect to clinically oriented tests of mid-level visual function. In 15 online subtests, a range of mid-level processes are covered, such as figure–ground segmentation, local and global processing, and shape perception. We also test the sensitivity to a wide variety of perceptual grouping cues, like common fate, collinearity, proximity, and closure. To reduce cognitive load, a matching-to-sample task is used for all subtests. Our online test can be administered in 20–45 min and is freely available at www.gestaltrevision.be/tests. The online implementation enables us to offer a separate interface for researchers and clinicians to have immediate access to the raw and summary results for each patient and to keep a record of their patient’s entire data. Also, each patient’s results can be flexibly compared with a range of age-matched norm samples. In conclusion, the L-POST is a valuable screening test for perceptual organization. The test allows clinicians to screen for deficits in visual perception and enables researchers to get a broader overview of mid-level visual processes that are preserved or disrupted in a given patient.
A recent dramatic increase in the number and scope of chronometric and norming lexical megastudies offers the ability to conduct virtual experiments-that is, to draw samples of items with properties that vary in critical linguistic dimensions. This paper introduces a bootstrapping approach, which enables testing of research hypotheses against a range of samples selected in a uniform, principled manner and evaluates how likely a theoretically motivated pattern is in a broad distribution of possible outcome patterns. We apply this approach to conflicting theoretical and empirical accounts of the relationship between the psychological valence (positivity) of a word and its speed of recognition. To this end, we conduct three sets of multiple virtual experiments with a factorial and a regression design, drawing data from two lexical decision megastudies. We discuss the influence that criteria for stimuli selection, statistical power, collinearity, and the choice of dataset have on the efficacy and outcomes of the bootstrapping procedure.
Word aversion has been defined as the visceral dislike of a word, independent of that words meaning. For instance, many people report a feeling of disgust upon saying or reading the word Moist, but very few report strong negative feelings towards synonyms of moist, such as damp or humid. We created a list of words with neutral meanings, for which many people report aversive feelings and had participants rate these words for emotional valence. Using information provided by the Affective Norming for English Words (ANEW) database, we created a scale for emotional valence and normed both the aversive words and their synonyms. We then gave participants a surprise free recall task following the rating to see whether certain words were more memorable than others, as well as to see whether the aversive words clustered together, indicating semantic association. Following norming, we conducted a series of experiments, aimed to determine what is driving the general dislike of aversive words. Our experiments utilize a software called Mouse Tracker, which allows us a nuanced look into whether context, morphology, or both play a role in our aversion to certain words. Experiment 1a was a priming study that manipulated the context in which words are presented. Aversive words and their synonyms were given positive contextual primes (e.g. cake - MOIST), negative contextual primes (e.g. skin - MOIST), or neutral primes (e.g. xxx - MOIST) and synonyms were primed with the same words in different scripts. Experiment 1b asked participants to rate aversive words along side positively and negatively valenced words from the ANEW and experiment 2 had participants rating unprimed nonwords, half of which had similar morphology to aversive words, and half of which did not.
Hierarchical data sets arise when the data for lower units (e.g., individuals such as students, clients, and citizens) are nested within higher units (e.g., groups such as classes, hospitals, and regions). In data collection for experimental research, estimating the required sample size beforehand is a fundamental question for obtaining sufficient statistical power and precision of the focused parameters. The present research extends previous research from Heo and Leon (2008) and Usami (2011b), by deriving closed-form formulas for determining the required sample size to test effects in experimental research with hierarchical data, and by focusing on both multisite-randomized trials (MRTs) and cluster-randomized trials (CRTs). These formulas consider both statistical power and the width of the confidence interval of a standardized effect size, on the basis of estimates from a random-intercept model for three-level data that considers both balanced and unbalanced designs. These formulas also address some important results, such as the lower bounds of the needed units at the highest levels.
Historically, rural schools have been known for their active engagement with parents and communities as well as for smaller class sizes, safer school environments, a more individualized approach to learning, flexible scheduling, creative approaches to acquiring expanded curriculum offerings, and a lower rate of students dropping out of school (Chalker, 2002; Johnson, 2006; Keith, Keith, Quirk, Cohen-Rosenthal, & Franzese, 1996). Although some rural communities have experienced population decline in the last few years, others have experienced growth, with 13% of the population growth in rural schools consisting of other than Whites from European ancestry (Dougherty, 2012). Demographic trends indicate that ethnically diverse populations will continue growing in rural areas (Johnson, 2006). Some rural communities are culturally established with rich tradition, religion heritage, and unique social norms based on isolated location of these areas: Alaskan villages, Native American reservations, and Amish farming communities, to name but a few (Nelson, 2010).Today, not all students enrolled in rural schools have the multi-generational involvement of their parents and grandparents experienced by students in past years (Bauch, 2000). Increasingly migrant families, parents with limited education, and singleparent homes have become more prevalent in rural communities (Grey, 1997; Schafft, Prins, & Movit, 2008). Parents who come from non-rural areas, other parts of the country, different cultural and religious traditions, isolated locations, or other countries may need guidance in learning how to navigate the cultural, social, and linguistic norms of rural schools and their communities. With ongoing demographic changes in rural schools, it becomes important for rural educators to be culturally sensitive to the needs of their changing communities. This study explored the perceptions of rural educators regarding their understanding of parental involvement and their reflection of how parental involvement worked in their schools.Theoretical FrameworkThrough their meta-analysis of school leadership research, Marzano, Waters, & McNulty (2005) affirmed parental involvement as one of several factors determining the capacity of schools to bring students to optimal levels of academic achievement. Students with involved parents, regardless of income or background, are more likely to earn higher grades, achieve higher test scores, enroll in higher-level academic programs, attend school regularly, graduate from high-school, and enroll in postsecondary education programs (Englund, Luckner, Whaley, & Egeland, 2004; Keith et al., 1996; Turney & Kao, 2009).Traditional Perspectives: Parental Involvement in Children's EducationIn the United States, educators have typically viewed parental involvement as something occurring within the school: participation in parent-teacher conferences, volunteer activities or committee work at school, and/or involvement with fund-raising activities (Berger, 1991; Weiss, Kreider, Lopez, & Chatman-Nelson, 2010; Young, 1995; Zarate, 2007). For decades, school systems have struggled with involving parents in their children's education. At the same time, the real challenges affecting parental involvement and strategies that effectively engage parents have not been clearly identified (Anfara & Mertens, 2008; Semke & Sheridan, 2011). Too often, when educators address parental involvement, they assume that parents alone are responsible for connecting with schools while also providing followthrough support for their children at home including homework, monitoring student performance at school, and the use of other meaningful learning activities (Epstein, 2001; Jeynes, 2010). In some cases, teachers may judge parents based on misinterpretations about the motivation, interest, and support of specific parents (Cooper, Crosnoe, Suizzo, & Pituch, 2010; Turney & Kao, 2009; Zarate, 2007). …
Laboratory phonology has been widely employed to understand the interactional relationship between the acoustic cues of English Lexical Stress (ELS)—duration, fundamental frequency, and intensity. However, research on ELS production in polysyllabic words is limited, and cross-linguistic research in this domain even more so. Hence, the impacts of second language (L2) experience and first language (L1) background on ELS acquisition have not been fully explored. This study of 100 adult Mandarin (Chinese), Arabic (Saudi Arabian), and English (Midwest American) speakers examines their ELS productions in tokens containing seven different stress-moving suffixes; i.e., Level 1 [ + cyclic] derivations according to lexical phonology. Speech samples were systematically analyzed using Praat and compared using statistical sampling. Native-speaker productions provided norm values for cross-reference to yield insights into the proposed Salience Hierarchy of the Acoustic Correlates of Stress (SHACS). The author recently reported the main findings which support the idea that SHACS exists in L1 sound schemes, and that native-like command of these systems can be acquired by L2 learners through increased L2 input. Other results are expected to reveal the role of tonic accent shift, the idiosyncrasies of individual suffixes, conflicts with standard dictionary pronunciations, and the effects of frequency perception scales on SHACS.
This article presents the results of research into the language culture of professional communication given the fluid nature of modern communication. The scope of interest is direct linguistic relations between professional senders and recipients of texts. As in the general culture, in professional communication the paradigms which have been used until now (in which the observance of specific stylistic and correctness indicators is recommended) are being dropped. As such, many senders make use of different styles and varieties of utterance at the same time, combining them freely within one text (e.g., formal and informal, careful and colloquial, etc.). Colloquial lexis (e.g., ogarnij się, ruchy, pierdoły) and linguistic forms which are incorrect according to the contemporary norm (e.g., na wiosce, mi się to podoba) are being added to customarily careful utterances—sermons, news reports, and lectures at school and university. Using the language in a temporary way, without any limitations and rigid confines, is becoming the rule. Senders use this way of speaking both to gain acceptance and to surprise recipients with unusual stylistic and lexical combinations. These new linguistic customs are especially visible among the most skillful language users, those who have a broad stylistic workshop and a large store of rhetorical devices. They appears not to fear shaking up the language system in connection to the observed changes. In the fluid modernity there is a self-restraining mechanism: one’s responsibility to make every utterance comprehensible to the recipient. Professional communication would otherwise be impossible.
This article is dedicated to “Diálogo de la lengua” (1535) by a Spanish grammarian of the Renaissance Juan de Valdés whose personality and works are not so widely known in Russia. “Diálogo de la lengua” represents the Renaissance dialogue, heir to the Greek-Latin traditions and is dedicated to the Castilian language, the interrelation of its norms and Language Usage, stylistic and lexical issues. The attention is focused on the genre features of “Diálogo” and the relationship of the author and the character to represent him in the dialogue. The analysis is conducted in terms of a combination of various stylistic and genres features of the language in the “Diálogo” making it possible to distinguish the traits of scientific style in the work (and to consider the “Diálogo de la lengua” as a grammar of the Castilian language and linguistic comment), as well as a work of art (scenic or philosophical Socratic dialogue).
SECOND LANGUAGE LEARNERS FACE A DUAL CHALLENGE IN VOCABULARY LEARNING: First, they must learn new names for the 100s of common objects that they encounter every day. Second, after some time, they discover that these names do not generalize according to the same rules used in their first language. Lexical categories frequently differ between languages (Malt et al., 1999), and successful language learning requires that bilinguals learn not just new words but new patterns for labeling objects. In the present study, Chinese learners of English with varying language histories and resident in two different language settings (Beijing, China and State College, PA, USA) named 67 photographs of common serving dishes (e.g., cups, plates, and bowls) in both Chinese and English. Participants' response patterns were quantified in terms of similarity to the responses of functionally monolingual native speakers of Chinese and English and showed the cross-language convergence previously observed in simultaneous bilinguals (Ameel et al., 2005). For English, bilinguals' names for each individual stimulus were also compared to the dominant name generated by the native speakers for the object. Using two statistical models, we disentangle the effects of several highly interactive variables from bilinguals' language histories and the naming norms of the native speaker community to predict inter-personal and inter-item variation in L2 (English) native-likeness. We find only a modest age of earliest exposure effect on L2 category native-likeness, but importantly, we find that classroom instruction in L2 negatively impacts L2 category native-likeness, even after significant immersion experience. We also identify a significant role of both L1 and L2 norms in bilinguals' L2 picture naming responses.
To investigate emotional expression in discourse expressed by speakers with and without right brain-damage (RBD). Analysis of lexical emotional expression in narrative and procedural discourse. General community. Males with RBD and a matched control group. Not applicable. The frequency and type of three appraisal resources:- affect (how people feel), judgment (whether people’s behavior conforms/transgresses social norms) and appreciation (reactions to/evaluation of things).The attitudes were also categorized by their grading (amplified or downplayed). Quantitatively, the individuals with RBD used fewer appraisal resources in narratives but performed similarly in procedures. They were able to express emotions to a greater extent in the personal rather than the sequence-picture samples. In personal narratives, they tended to evaluate things or phenomena more frequently than expressing their own feelings, thereby distancing themselves. Furthermore, they demonstrated greater impairment on the negative, rather than the positive, topic, providing support for the valence hypothesis (the right hemisphere is considered to be dominant for negative emotions). They also tended to intensify their emotions more and mitigate negative emotions less which is less socially appropriate and may contribute to their social deficits. These individuals are considered to be socially disconnected from the world; this may be accounted for by their restricted emotional expression. Communicating emotions is important in building solidarity and in belonging. Affective difficulties are among the most important factors influencing rehabilitation outcome and produce the greatest burden for family and rehabilitation staff. The assessment and treatment of evaluation should be an integral part of rehabilitation for this much-neglected group.
Les normes qui se diffusent dans les organisations ont été souvent conçues par des ingénieurs et pour le secteur industriel. Il est intéressant de se demander si les responsables qualité de formation non scientifique ou technique traduisent ces outils de contrôle organisationnel de façon similaire à leurs homologues scientifiques ou techniciens. Pour répondre à cette question, nous avons étudié le discours de différents responsables qualité concernant la norme ISO 9001 selon une méthode d’analyse du champ lexical en nous appuyant sur le fait que le vocabulaire choisi traduit l’image de la norme diffusée. Nous avons trouvé deux types de modèle de contrôle organisationnel dans les discours étudiés et l’on peut penser qu’ils sont corrélés à la nature de la formation de base des personnes. Dans l’objectif de faciliter la légitimation de la norme ISO 9001 au sein de l’entreprise, ceci pourrait éclairer le choix du recrutement des responsables qualité. Mots-clés: contrôle organisationnel, responsable qualité, analyse de discours
The article deals with the linguistic issues of composing a reference book of regional toponyms – a genre that requires special consideration in national lexicography. The assortment of these issues gave the possibility to carry out complex description of regional toponyms on the basis of semantic, functional, and orpthologuos criteria that let unify the names of Volgograd region settlements that are registered in various documents. The significance of the composed reference book is determined by several factors – the presence of local subsystems of geographical names in Russian toponymy; the inconsistency of current orthography norms on using capital letter in compound proprius names and fused-with-hyphen spelling of toponyms and off-toponym derivations; the lack of linguistically justified explanation of peculiarities of grammatical norms in the field of proper names use. The reference book of regional toponyms is based on the object description (toponymic vocabulary), principles of lexical units selection (description of spelling and grammatical properties of toponyms, encyclopedic information), the glossary (full list of toponyms of Volgograd region), typical article. The articles in the reference book are arranged in lexicographical zones with grammatical and semantic markers, lexicographical illustrations, other lexicographical labels, word etymology including. The reference book on Volgograd region toponymy is addressed to executive and administration authorities, journalists, regional ethnographers.
Due to globalization there is an increase in the appearances of languages in the multilingual linguistic landscape in urban spaces. Commentators have described this state of affairs as super-, mega- or complex diversity. Mainstream sociolinguists have argued that languages have no fixed boundaries and that they are "fluid" in fact. The output of speech production and language use is actually referred to as "languaging". The terms implies that languages are rather resources but not fixed tools for communication. In this paper, I will argue that this theory to which I will refer as the superdiversity/languaging theory cannot cover multilingual data in terms of resources only, if phenomena of multilingual linguistic landscape are studied more carefully. It turns out that constructions that look like "languaging" are from a linguistic point of view in fact well-known cases of code-switching (or -mixing) with separate languages involved, a dominant language and clearly targeted messages for the speakers of the "underlying" language. Hence, I will conclude that linguistic data in multilingual urban spaces are not necessarily arranged in terms of resources but rather in terms of Fishmanian diglossia, triglossia, and so on. This implies that even in these cases of languaging there is no reason to operate with concepts of language other than recognizable languages that are characterized by a prototypical grammatical and lexical basic core. Hence, languages in this sense and not code-switched variants, like "English as a Lingua Franca" feed into strategies of transnational communication, although the output of transnational communication can<br/>be a code-switched variant of English as well. However, I agree with the proponents of the<br/>superdiversity/languaging theory that it is highly relevant to study the proliferation of all sorts of multilingualism in the context of complex linguistic diversity. This reveals not only the structures and rules of language and languages that I will define as linguistic categories in accordance with Chomskyan grammar but also provides insight into the quickly changing semantic and world view concepts due to globalization. However, the code-switched variants appearing in multilingual complex spaces are not suitable for linguistic diversity management that includes institutions. Institutions are by definition the outcome of norm-based governance strategies and will implement norm-based entities, like languages that are recognizable and make possible contextualized, sophisticated language use. This rules out highly individual, spontaneous production of language, like languaging-phenomena.
On the basis of a large amount of corpus-based studies on translation works, the translation universals hypothesis is proposed. As it claims, translations enjoy some general features and Baker (1993) summarizes them into three universals, namely simplification, explicitation, and normalization, which are supported by many following researches. However, some of the later studies contradict with these rules in several ways, and the usages of passive voice and pronouns are the two most controversial issues. Previous researches suggest that according to the universal features of explicitation and normalization, translated texts tend to have a lower frequency of pronouns while over-representing the passive voice. To examine such claimings, 160 original English abstracts from two leading journals in the field of translation studies, The Translator and Translation Studies, and another 160 English abstracts from Chinese Translator Journal and Chinese Science & Technology Translators Journal, which are translated from Chinese abstracts, are collected. Two corpora are then constructed, namely the Original English Abstracts Corpus (OEAC) and Translated English Abstracts Corpus (TEAC). The CLAWS Part-of-speech Tagger is used to tag the lexical items and word processing tool AntConc 3.2.4 is used for retrieving the words. The comparison between the two corpora suggests that the translated English abstracts contain a lower level of frequency in the use of both passive voice and pronouns, which partially query the hypothesis of explicitation and normalization. A detailed analysis shows a higher frequency of past-tense passives in the OEAC and more passives in perfect tense in the TEAC. The OEAC also contains more relative pronouns while the other contains more indefinite pronouns. The norm theory is utilized to account for such phenomena. The detailed results of the study are expected to shed some lights on professional translating and academic writing.
We present a dependency treebank of Buddhist Chinese texts, containing more than 50K characters drawn from four sutras in the Chinese Buddhist Canon. With dates of composition that span almost five centuries, these sutras bear witness to the evolution of the Chinese language. The treebank has been annotated using the part-of-speech tagset of the Penn Chinese Treebank, and the Stanford Dependencies for Chinese with slight modifications. The article first discusses the texts and the annotation framework of this treebank, and reports on inter-annotator agreement. It then describes the search platform, to which the treebank has been imported, and applies the treebank to an open question in Chinese historical linguistics—the emergence of the Chinese copula.
Listeners show a reliable bias towards interpreting speech sounds in a way that conforms to linguistic restrictions (phonotactic constraints) on the permissible patterning of speech sounds in a language. This perceptual bias may enforce and strengthen the systematicity that is the hallmark of phonological representation. Using Granger causality analysis of magnetic resonance imaging (MRI)- constrained magnetoencephalography (MEG) and electroencephalography (EEG) data, we tested the differential predictions of rule-based, frequency–based, and top-down lexical influence-driven explanations of processes that produce phonotactic biases in phoneme categorization. Consistent with the top-down lexical influence account, brain regions associated with the representation of words had a stronger influence on acoustic-phonetic regions in trials that led to the identification of phonotactically legal (versus illegal) word-initial consonant clusters. Regions associated with the application of lingu)
The article considers the problem of formation of grammar skills in teaching foreign language communication and skills to understand adequately different types of discourse in relation to a definite situation of communication, and also the ability to reproduce statements acceptable in a definite communicative situation. Nowadays discourse competence is considered one of the most significant skills. Discourse competence implies the ability to code and decode information with the help of a foreign language in accordance with its lexical, grammar, syntactical norms and also taking into consideration stylistic, genre, sociocultural, psychological and emotional factors, using the sources of cohesion and coherence to achieve a communicative goal. Discourse competence is provided by the knowledge of the strategies which are typical for the target culture and taking into account grammatical structures in a definite communicative situation. According to the modern goals of foreign language education, in our case, in forming grammatical authenticity of a learner's speech, only discourse should be the basis of the learning process as the adequacy of learners' speech behavior is measured by the achievement of a communicative goal in a definite situation of foreign language communication and not only by the correctness or incorrectness of a produced statement. A discourse basis of the learning process underlines the dynamic pragmatist character of the language, taking into account extralinguistic factors of the situation of intercultural communication.
How does the presence of a categorically related word influence picture naming latencies? In order to test competitive and noncompetitive accounts of lexical selection in spoken word production, we employed the picture-word interference (PWI) paradigm to investigate how conceptual feature overlap influences naming latencies when distractors are category coordinates of the target picture. Mahon et al. (2007. Lexical selection is not by competition: A reinterpretation of semantic interference and facilitation effects in the picture-word interference paradigm. Journal of Experimental Psychology. Learning, Memory, and Cognition, 33(3), 503-535. doi:10.1037/0278-7393.33.3.503 ) reported that semantically close distractors (e.g., zebra) facilitated target picture naming latencies (e.g., HORSE) compared to far distractors (e.g., whale). We failed to replicate a facilitation effect for within-category close versus far target-distractor pairings using near-identical materials based on feature production norms, instead obtaining reliably larger interference effects (Experiments 1 and 2). The interference effect did not show a monotonic increase across multiple levels of within-category semantic distance, although there was evidence of a linear trend when unrelated distractors were included in analyses (Experiment 2). Our results show that semantic interference in PWI is greater for semantically close than for far category coordinate relations, reflecting the extent of conceptual feature overlap between target and distractor. These findings are consistent with the assumptions of prominent competitive lexical selection models of speech production.
Recently, we reported on our efforts to build the first prototype of KurdNet. In this proposal, we highlight the shortcomings of the current prototype and put forward a detailed plan to transform this prototype to a full-fledged lexical database for the Kurdish language.
We present the Uppsala Persian Dependency Treebank (UPDT) with a syntactic annotation scheme based on Stanford Typed Dependencies. The treebank consists of 6,000 sentences and 151,671 tokens with an average sentence length of 25 words. The data is from different genres, including newspaper articles and fiction, as well as technical descriptions and texts about culture and art, taken from the open source Uppsala Persian Corpus (UPC). The syntactic annotation scheme is extended for Persian to include all syntactic relations that could not be covered by the primary scheme developed for English. In addition, we present open source tools for automatic analysis of Persian containing a text normalizer, a sentence segmenter and tokenizer, a part-of-speech tagger, and a parser. The treebank and the parser have been developed simultaneously in a bootstrapping procedure. The result of a parsing experiment shows an overall labeled attachment score of 82.05% and an unlabeled attachment score of 85.29%. The treebank is freely available as an open source resource.
The first semantic roles corpus in Persian language, containing about 30,000 sentences from contemporary Persian language, is manually annotated. This corpus, based on the concept of thematic roles of Fillmore, adds a layer of predicate-argument information to the syntactic structures of Persian Dependency Treebank. In this corpus, the verbs, propositional nouns and adjectives are regarded as the predicates of the sentences and are annotated according to their argument structure. The data was prepared based on Conference on Natural Language Learning (CoNLL) dependency format. Semantic tags used as the semantic annotations include thematic roles and functional tags. Thematic roles labels present the argument structure of the predicates of the sentences, and functional tags modify the verb or the whole sentence. The number of thematic roles tags and functional tags are 27 and 15, respectively. The two tags of NEGATION and MODALS are used as the functional tags.
In this paper, we propose a Connectivedriven Dependency Tree (CDT) scheme to represent the discourse rhetorical structure in Chinese language, with elementary discourse units as leaf nodes and connectives as non-leaf nodes, largely motivated by the Penn Discourse Treebank and the Rhetorical Structure Theory. In particular, connectives are employed to directly represent the hierarchy of the tree structure and the rhetorical relation of a discourse, while the nuclei of discourse units are globally determined with reference to the dependency theory. Guided by the CDT scheme, we manually annotate a Chinese Discourse Treebank (CDTB) of 500 documents. Preliminary evaluation justifies the appropriateness of the CDT scheme to Chinese discourse analysis and the usefulness of our manually annotated CDTB corpus.
There are few explicit discourse connectives in Chinese texts,which bring in new challenge for the traditional connective-grounded coherence annotation scheme.The paper proposes a new idea to deal with the problem.We introduce topic chain(TC)as a main coherence representation and design several topic-comment relations to describe the complex event relations among TC-linked sentences.Therefore,a new coherence annotation scheme based on TCs and connectives are built accordingly.The tentative confirmatory experiments on the Tsinghua Chinese Treebank(TCT)data set show that more than 76%and 50% Chinese complex sentences have TCs and connectives respectively.They can co-occur in most Chinese sentences.The phenomena verify the feasibility and availability of this scheme.
This paper presents the ITU Turkish Dependency Validation Set firstly introduced in 2007 [36] in order to serve as the test set of the CoNLL-XI shared task (shared task of the Conference on Computational Natural Language Learning 2007 [28] ). The dataset is available from http://web.itu.edu.tr/gulsenc/treebanks.html and is used by several academic studies so far.
See http://clld.org
The present paper aims to categorize different types of synonymous words and also to highlight their synonymic pattern as well as grammatical categories found in Wordnet of Assamese language. Synonymy is an important component of vocabulary of the language. It establishes lexical relation between words. In fact, the term ‘synonymy’ is applied to the two or more words which share the same semantic features. WorldNet is a lexical database consisting of synsets. A synset is constructed by assembling a set of synonyms that together define a unique sense and synset is the basic foundation of Wordnet. Assamese language is rich in synonyms. In Assamese WorldNet, more than 20,000 synsets are entered under the categories of Noun, Verb, Adverb and Adjective. These synsets can of different types according to their semantic similarity, connotation, denotation, stylistic variations etc.
Emotion effects in event-related brain potentials (ERPs) have previously been reported for a range of visual stimuli, including emotional words, pictures, and facial expressions. Still, little is known about the actual comparability of emotion effects across these stimulus classes. The present study aimed to fill this gap by investigating emotion effects in response to words, pictures, and facial expressions using a blocked within-subject design. Furthermore, ratings of stimulus arousal and valence were collected from an independent sample of participants. Modulations of early posterior negativity (EPN) and late positive complex (LPC) were visible for all stimulus domains, but showed clear differences, particularly in valence processing. While emotion effects were limited to positive stimuli for words, they were predominant for negative stimuli in pictures and facial expressions. These findings corroborate the notion of a positivity offset for words and a negativity bias for pictures and facial expressions, which was assumed to be caused by generally lower arousal levels of written language. Interestingly, however, these assumed differences were not confirmed by arousal ratings. Instead, words were rated as overall more positive than pictures and facial expressions. Taken together, the present results point toward systematic differences in the processing of written words and pictorial stimuli of emotional content, not only in terms of a valence bias evident in ERPs, but also concerning their emotional evaluation captured by ratings of stimulus valence and arousal.
This paper describes the multistage process for building Arabic WordNet (ArWn) to be used in mobile device. The goal of this paper is how to create corpus, starting with selecting an annotation task, designing the data with the annotation process, and finally evaluating the results for a particular goal. Therefore, the paper presents designing and implementing bi-lingual lexicon to be used in machine translation and language processing. Consequently, the paper takes into consideration language characteristics in both directions Arabic and English. The proposed system is based on WordNet lexical database with a semantic and commonsense knowledge. The proposed dictionary will be implemented for mobile devices, therefore; the cloud computing will be used in this implementation. Consequently, SQL Azure will be used to solve scalability, and interoperability of mobile users and other methods have been used for both Arabic and English languages. So, the SQL Azure will be used as the cloud database to solve both the scalability in the data with scale terabytes of data to millions of mobile users and the interoperability challenges. The system dictionary is developed and tested in Android mobile platform. Experimental results show that the proposed system has two versionsat work; offline and online. The online approach uses the mobiles computing in the cloud system to reduce the storage complexity of the mobile. Real time test will be used in order to evaluate the system access and respond times to display results.
The present study compared the impact of symbolic equivalence and opposition relations on fear generalisation. In a procedure using nonsense words, some stimuli became symbolically equivalent to an aversively conditioned stimulus while others were symbolically opposite. The generalisation of fear to symbolically related stimuli was then measured using behavioural avoidance, retrospective unconditioned stimulus expectancy and stimulus valence ratings. Equivalence relations facilitated fear generalisation while opposition relations constrained generalisation. The potential clinical implications of symbolic generalisation are discussed.
The valency lexicon PDT-Vallex has been built in close connection with the annotation of the Prague Dependency Treebank project (PDT) and its successors (mainly the Prague Czech-English Dependency Treebank project, PCEDT). It contains over 11000 valency frames for more than 7000 verbs which occurred in the PDT or PCEDT. It is available in electronically processable format (XML) together with the aforementioned treebanks (to be viewed and edited by TrEd, the PDT/PCEDT main annotation tool), and also in more human readable form including corpus examples (see the WEBSITE link below). The main feature of the lexicon is its linking to the annotated corpora - each occurrence of each verb is linked to the appropriate valency frame with additional (generalized) information about its usage and surface morphosyntactic form alternatives.
Recent resting-state functional magnetic resonance imaging (fMRI) studies using graph theory metrics have revealed that the functional network of the human brain possesses small-world characteristics and comprises several functional hub regions. However, it is unclear how the affective functional network is organized in the brain during the processing of affective information. In this study, the fMRI data were collected from 25 healthy college students as they viewed a total of 81 positive, neutral, and negative pictures. The results indicated that affective functional networks exhibit weaker small-worldness properties with higher local efficiency, implying that local connections increase during viewing affective pictures. Moreover, positive and negative emotional processing exhibit dissociable functional hubs, emerging mainly in task-positive regions. These functional hubs, which are the centers of information processing, have nodal betweenness centrality values that are at least 1.5 times larger than the average betweenness centrality of the network. Positive affect scores correlated with the betweenness values of the right orbital frontal cortex (OFC) and the right putamen in the positive emotional network; negative affect scores correlated with the betweenness values of the left OFC and the left amygdala in the negative emotional network. The local efficiencies in the left superior and inferior parietal lobe correlated with subsequent arousal ratings of positive and negative pictures, respectively. These observations provide important evidence for the organizational principles of the human brain functional connectome during the processing of affective information.
Processing unpleasant affective cues induces elevated momentary symptom reports, especially in persons with high levels of symptom reporting in daily life. The present study aimed to examine whether applying an emotion regulation strategy, i.e. affect labeling, can inhibit these emotion influences on symptom reporting. Student participants (N = 61) with varying levels of habitual symptom reporting completed six picture viewing trials of homogeneous valence (three pleasant, three unpleasant) under three conditions: merely viewing, emotional labeling, or content (non-emotional) labeling. Affect ratings and symptom reports were collected after each trial. Participants completed a motor inhibition task and self-control questionnaires as indices of their inhibitory capacities. Heart rate variability was also measured. Labeling, either emotional or non-emotional, significantly reduced experienced affect, as well as the elevated symptoms reports observed after unpleasant picture viewing. These labeling effects became more pronounced with increasing levels of habitual symptom reporting, suggesting a moderating role of the latter variable, but did not correlate with any index of general inhibitory capacity. Our findings suggest that using an emotion regulation strategy, such as labeling emotional stimuli, can reverse the effects of unpleasant stimuli on symptom reporting and that such strategies can be especially beneficial for individuals suffering from medically unexplained physical symptoms.
Lexical Database the Japanese WordNet is a useful tool in natural language processing. However, it is officially announced that Japanese WordNet contains 5% errors. In this paper, we discuss error detection methods in the Japanese WordNet. キーワード Thesaurus, WordNet, Japanese WordNet, 1. はじめに 日本語 WordNet[1,2]は Princeton 大学が開発した WordNet[3]を用いた言語データベースである。日本語 WordNet は自然言語処理において有用であり、様々な 実験に使用されている [4]。フリーの Web シソーラス サービスにおいて、日本語 WordNet は一般的に使用さ れている。しかしながら、現行の日本語 WordNet は間 違いを 5%ほど含んでいると作成者らが認めており [2]、 それらの間違いが日本語 WordNet の使いやすさに影響 を及ぼしている可能性がある。 本論文では、われわれが検証した日本語 WordNet の 間違い探知手法において議論する。間違い探知は日本 語 WordNet の間違い修正の第一段階である。この手法 は、大規模言語データベースの作成に有用であると考 える。我々は特に日本語 WordNet の似たような間違い 1 Weblio, http://ejje.weblio.jp の発見を主眼にしており、この間違いのことを「類義 語の間違い」と呼んでいる。 英語でない WordNet や WordNet に似た言語データベ ースの作成という点において、複数のプロジェクトが 行 わ れ て い る 。 日 本 語 WordNet や Chinese Open WordNet[5]は、ブートストラップの段階で、Princeton WordNet のマッピング手法を用いて半自動生成されて いる。 また、 Universal WordNet[6]や Babel Net[7]、 Open Multilingual WordNet[8]といった、WordNet の拡張によ る統合、多言語概念字句データベースの生成の試みも なされている。概念と語句、または複数の概念間の関 係は、Wikipedia やタグ付けコーパスのような様々な資 源から自動的に抽出することが可能である。それらに よって得られた統合データベースの品質は、生成者自 身や、ネットワークコミュニティによって評価されて きた。 WordNet は、オントロジーのひとつとしてみなすこ とができる。多言語 WordNet を生成する場合には、言 語数に応じたオントロジー間のマッピングをする必要 がある。そのため、オントロジーの間違いの検出と修 正、オントロジー間のマッピングに関する研究がなさ れてきた。これらの研究において、オントロジー内で 分類が間違っているものや、冗長もしくは不適切であ る、または間違った関係性を生成されている箇所を修 正する試みがなされてきた。 間違い検出の手法として、日本語 WordNet のみを使 用した手法を動詞に適用した場合をベースラインとし て提示する [9]。また、コーパスを用いた単語をベクト ル化し、これらのコサイン類似度によって名詞の間違 い検出の手法として使用できないかを議論する。 本論文では、第 2 節で WordNet と日本語 WordNet の 説明、第 3 節で本論文のコンセプトと WordNet の構造 における「同義語の間違い」の一例を紹介する。第 4 節では間違いの抽出法に関するわれわれの手法の説明、 第 5 節では手法を用いた場合の結果の提示を行う。第 6 節では、本手法の Princeton WordNet における応用例 と、関心を持っている別手法に関しての説明、第 7 節 で word2vec を用いた単語のベクトル化とそれらを用 いた間違い検出の実験、第 8 節に今後の展望と課題を 述べる。 2. WordNet と日本語 WordNet 2.1. Princeton WordNet Princeton WordNet は英語の大規模言語データベース である。名詞、動詞、形容詞、副詞といった品詞ごと に、明確なコンセプトを持った「Synset」という認知同 義語のセットに纏められる。各 Synset は固有の ID に よって管理されており、Gloss と呼ばれる、Synset の簡 単な意味を説明するテキストがリンクされている。 Synset は概念 -意味関係もしくは字句トークン関係で 相互リンクを持っている。単語が持つ意味を Synset に よってグループ化することができるため、WordNet は シソーラスとして使用できる。多義である単語が存在 するため、単語は複数の Synset に属することがある。 2.2. 日本語 WordNet 日本語 WordNet は Princeton WordNet を基にした、日 本語の語彙データベースである。日本語 WordNet のプ 2 Wikipedia, http://ja.wikipedia.org ロジェクトの目的は、誰でも自由に使用可能な大規模 日本語データベースを提供することである。このデー タベースは 2006 年から開発されている。 日本語 WordNet の構造は、Princeton WordNet に準拠 している [1]。しかし、日本語と英語という言語の違い が存在するため、日本語 WordNet は Princeton WordNet に含まれていないオリジナルの Synset を含んでいる。 また、日本語 WordNet は、シソーラスとしての精度よ り多数の概念を包括することに主眼を置いている。 現行の日本語 WordNet の規模は以下のとおりである。 ・57,238 概念(Synset 数) ・93,834 語(日本語) ・158,058 語義(単語 -synset ペア数) 図 1 日本語 WordNet の Synset-同義語間リンク例 日本語とリンクを持つ Synset は日本語の gloss を持 っている。日本語 WordNet のカバー範囲の拡張のため に、SUMO や Wikipedia、GoiTaikei[10]といった他のリ ソースが使用されている。 2.3. 他言語の WordNet と WordNet の拡張 Princeton WordNet を基にした、様々な言語の言語デ ータベース作成プロジェクトが存在する。一部のプロ ジェクトでは WordNet、Wikipedia、Wiktionary及びそ の他の言語資源を用いて、多言語のごくデータベース を作成しようと試みている。 既存の言語資源と新しいデータベース間のマッピ ングの正確さは、それによって出力されるデータベー スの整合性の正しさを証明する指標になるので非常に 重要である。新しい言語の WordNet を作成することは、 他の言語からなる新しいオントロジーで表現されてい る、既存のオントロジーからマッピングで作成すると みなすことができる。 3 Wiktionary, http://ja.wiktionary.org 3. 日本語 WordNet の間違い 間違いの訂正は、新しく作成したオントロジーや、 オントロジー間のマッピングの整合性の確認において 重要である。日本語 WordNet の現行のバージョンでは、 約 5%の間違いが含まれている。また、Chinese Open WordNet も、それに匹敵するエラー率である。本節の 残りでは、同義語における間違いにおける、エラーの 種類に焦点を当てる。 3.1. 同義語の間違い WordNet の構造において、「同義語の間違い」を、語 wmiss が属している synset(S とする)の Gloss と合致 しない語であると定義する。 図 2 では、Synset 02651424-v について図示している。 この Synset は「泊める」、「収容」、「宿る」、「持ち込む」 という 4 つの同義語を持っている。
The chapter argues that language, which rests on the sharing of linguistic norms, honest information, and moral norms, evolved through a co-evolutionary process with a pivotal role for intersubjectivity. Mainstream evolutionary models, based only on individual-level and gene-level selection, are argued to be incapable to account for such sharing of care, values and information, thus implying the need to evoke multi-level selection, including (cultural) group selection. Four of the most influential current theories of the evolution of human-scale sociality, those of Dunbar, Deacon, Tomasello and Hrdy, are compared and evaluated on the basis of their answers to five questions: (1) Why we and not others? (2) How: by what mechanisms? (3) When? (4) In what kind of social settings? (5) What are the implications for ontogeny? The conclusions are that the theories are to a large degree complementary, and that they all assume, explicitly or not, a role for group selection. Hrdy’s theory, focusing on the evolution of alloparenting, is argued to provide the best explanation for the onset of the evolution of human intersubjectivity, and can furthermore offer a Darwinian framework for Tomasello’s theory of shared intentionality. Deacon’s theory deals rather with the evolution of morality and its co-evolution with “symbolic reference”, but these are necessarily antecedent to the primary evolution of human intersubjectivity. Dunbar’s theory on the transition from “musical” vocal-grooming to vocal “gossip” can be seen as providing a partial explanation for evolution of spoken language, most likely with Homo heidelbergensis 0.5 MYA, but presupposes the capacities accounted for by the other models.
We develop an instance (token) based extension of the state of the art word (type) based part-ofspeech induction system introduced in (Yatbaz et al., 2012). Each word instance is represented by a feature vector that combines information from the target word and probable substitutes sampled from an n-gram model representing its context. Modeling ambiguity using an instance based model does not lead to significant gains in overall accuracy in part-of-speech tagging because most words in running text are used in their most frequent class (e.g. 93.69% in the Penn Treebank). However it is important to model ambiguity because most frequent words are ambiguous and not modeling them correctly may negatively affect upstream tasks. Our main contribution is to show that an instance based model can achieve significantly higher accuracy on ambiguous words at the cost of a slight degradation on unambiguous ones, maintaining a comparable overall accuracy. On the Penn Treebank, the overall many-to-one accuracy of the system is within 1% of the state-of-the-art (80%), while on highly ambiguous words it is up to 70% better. On multilingual experiments our results are significantly better than or comparable to the best published word or instance based systems on 15 out of 19 corpora in 15 languages. The vector representations for words used in our system are available for download for further experiments.
Recent years have seen an increased interest in and availability of parallel corpora. Large corpora from international organizations (e.g. European Union, United Nations, European Patent Office), or from multilingual Internet sites (e.g. OpenSubtitles) are now easily available and are used for statistical machine translation but also for online search by different user groups. This paper gives an overview of different usages and different types of search systems. In the past, parallel corpus search systems were based on sentence-aligned corpora. We argue that automatic word alignment allows for major innovations in searching parallel corpora. Some online query systems already employ word alignment for sorting translation variants, but none supports the full query functionality that has been developed for parallel treebanks. We propose to develop such a system for efficiently searching large parallel corpora with a powerful query language.
The feminist movement purports to improve conditions for women, and yet only a minority of women in modern societies self-identify as feminists. This is known as the feminist paradox. It has been suggested that feminists exhibit both physiological and psychological characteristics associated with heightened masculinization, which may predispose women for heightened competitiveness, sex-atypical behaviors, and belief in the interchangeability of sex roles. If feminist activists, i.e., those that manufacture the public image of feminism, are indeed masculinized relative to women in general, this might explain why the views and preferences of these two groups are at variance with each other. We measured the 2D:4D digit ratios (collected from both hands) and a personality trait known as dominance (measured with the Directiveness scale) in a sample of women attending a feminist conference. The sample exhibited significantly more masculine 2D:4D and higher dominance ratings than comparison samples representative of women in general, and these variables were furthermore positively correlated for both hands. The feminist paradox might thus to some extent be explained by biological differences between women in general and the activist women who formulate the feminist agenda.
We present an algorithm and implementation for extracting recurring fragments from treebanks. Using a tree-kernel method the largest common fragments are extracted from each pair of trees. The algorithm presented achieves a thirty-fold speedup over the previously available method on the Wall Street Journal dataset. It is also more general, in that it supports trees with discontinuous constituents. The resulting fragments can be used as a tree-substitution grammar or in classification problems such as authorship attribution and other stylometry tasks.
EngVallex is the English counterpart of the PDT-Vallex valency lexicon, using the same view of valency, valency frames and the description of a surface form of verbal arguments. EngVallex contains links also to PropBank and Verbnet, two existing English predicate-argument lexicons used, i.a., for the PropBank project. The EngVallex lexicon is fully linked to the English side of the PCEDT parallel treebank, which is in fact the PTB re-annotated using the Prague Dependency Treebank style of annotation. The EngVallex is available in an XML format in our repository, and also in a searchable form with examples from the PCEDT.