Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Background: During sentence processing we decode the sequential combination of words, phrases or sentences according to previously learned rules. The computational mechanisms and neural correlates of these rules are still much debated. Other key issue is whether sentence processing solely relies on language-specific mechanisms or is it also governed by domain-general principles. Methodology/Principal Findings: In the present study, we investigated the relationship between sentence processing and implicit sequence learning in a dual-task paradigm in which the primary task was a non-linguistic task (Alternating Serial Reaction Time Task for measuring probabilistic implicit sequence learning), while the secondary task were a sentence comprehension task relying on syntactic processing. We used two control conditions: a non-linguistic one (math condition) and a linguistic task (word processing task). Here we show that the sentence processing interfered with the probabilistic implicit seque)
Background: The well-established left hemisphere specialisation for language processing has long been claimed to be based on a low-level auditory specialization for specific acoustic features in speech, particularly regarding 'rapid temporal processing'. Methodology: A novel analysis/synthesis technique was used to construct a variety of sounds based on simple sentences which could be manipulated in spectro-temporal complexity, and whether they were intelligible or not. All sounds consisted of two noise-excited spectral prominences (based on the lower two formants in the original speech) which could be static or varying in frequency and/or amplitude independently. Dynamically varying both acoustic features based on the same sentence led to intelligible speech but when either or both acoustic features were static, the stimuli were not intelligible. Using the frequency dynamics from one sentence with the amplitude dynamics of another led to unintelligible sounds of comparable spectro-te)
This article describes the RMediation package,which offers various methods for building confidence intervals (CIs) for mediated effects. The mediated effect is the product of two regression coefficients. The distribution-of-the-product method has the best statistical performance of existing methods for building CIs for the mediated effect. RMediation produces CIs using methods based on the distribution of product, Monte Carlo simulations, and an asymptotic normal distribution. Furthermore, RMediation generates percentiles, quantiles, and the plot of the distribution and CI for the mediated effect. An existing program, called PRODCLIN, published in Behavior Research Methods, has been widely cited and used by researchers to build accurate CIs. PRODCLIN has several limitations: The program is somewhat cumbersome to access and yields no result for several cases. RMediation described herein is based on the widely available R software, includes several capabilities not available in PRODCLIN, and provides accurate results that PRODCLIN could not.
Domestic dogs are skillful at using the human pointing gesture. In this study we investigated whether dogs take contextual information into account when following pointing gestures, specifically, whether they follow human pointing gestures more readily in the context in which food has been found previously. Also varied was the human's tone of voice as either imperative or informative. Dogs were more sustained in their searching behavior in the 'context' condition as opposed to the 'no context' condition, suggesting that they do not simply follow a pointing gesture blindly but use previously acquired contextual information to inform their interpretation of that pointing gesture. Dogs also showed more sustained searching behavior when there was pointing than when there was not, suggesting that they expect to find a referent when they see a human point. Finally, dogs searched more in high-pitched informative trials as opposed to the low-pitched imperative trials, whereas in the latter do)
Personal naming practices exist in all human groups and are far from random. Rather, they continue to reflect social norms and ethno-cultural customs that have developed over generations. As a consequence, contemporary name frequency distributions retain distinct geographic, social and ethno-cultural patterning that can be exploited to understand population structure in human biology, public health and social science. Previous attempts to detect and delineate such structure in large populations have entailed extensive empirical analysis of naming conventions in different parts of the world without seeking any general or automated methods of population classification by ethno-cultural origin. Here we show how 'naming networks', constructed from forename-surname pairs of a large sample of the contemporary human population in 17 countries, provide a valuable representation of cultural, ethnic and linguistic population structure around the world. This innovative approach enriches and adds)
We analyze and extend a recently proposed model of linguistic diffusion in social networks, to analytically derive time to convergence, and to account for the innovation phase of lexical dynamics in networks. Our new model, the degree-biased voter model with innovation, shows that the probability of existence of a norm is inversely related to innovation probability. When the innovation rate in the population is low, variants that become norms are due to a peripheral member with high probability. As the innovation rate increases, the fraction of time that the norm is a peripheral-introduced variant and the total time for which a norm exists at all in the population decrease. These results align with historical observations of rapid increase and generalization of slang words, technical terms, and new common expressions at times of cultural change in some languages.
Bilingualism provides a unique opportunity for understanding the relative roles of proficiency and order of acquisition in determining how the brain represents language. In a previous study, we combined magnetoencephalography (MEG) and magnetic resonance imaging (MRI) to examine the spatiotemporal dynamics of word processing in a group of Spanish- English bilinguals who were more proficient in their native language. We found that from the earliest stages of lexical processing, words in the second language evoke greater activity in bilateral posterior visual regions, while activity to the native language is largely confined to classical left hemisphere fronto-temporal areas. In the present study, we sought to examine whether these effects relate to language proficiency or order of language acquisition by testing Spanish-English bilingual subjects who had become dominant in their second language. Additionally, we wanted to determine whether activity in bilateral visual regions was relat)
Background: The language faculty is probably the most distinctive feature of our species, and endows us with a unique ability to exchange highly structured information. In written language, information is encoded by the concatenation of basic symbols under grammatical and semantic constraints. As is also the case in other natural information carriers, the resulting symbolic sequences show a delicate balance between order and disorder. That balance is determined by the interplay between the diversity of symbols and by their specific ordering in the sequences. Here we used entropy to quantify the contribution of different organizational levels to the overall statistical structure of language. Methodology/Principal Findings: We computed a relative entropy measure to quantify the degree of ordering in word sequences from languages belonging to several linguistic families. While a direct estimation of the overall entropy of language yielded values that varied for the different families con)
The paper addresses the core problems of foreign students' communicative competence and justifies the «language culture» course as mean to build the required skills; the paper also proposes methodology of teaching those skills in the lexical norms study.
Understanding the role that social cues have on interpersonal choice, and their susceptibility to contextual effects, is of core importance to models of social decision-making. Language, on the other hand, is one of the main means of communication during social interactions in our culture. The present experiments tested whether positive and negative linguistic descriptions of alleged partners in a modified Ultimatum Game biased decisions made to the same set of offers, and whether the contextual uncertainty of the game modulated this biasing effect. The results showed that in an uncertain context, the same offers were accepted with higher probability when they were preceded by positive rather than by negative valenced trait-words. Participants also accepted fair offers with higher probability than unfair offers, but this effect did not interact with the valence of the social descriptive words. In addition, the speed of the decision was affected by valence: acceptance choices were fast)
This article focuses on tracheotomy, which is transformation of language, norms and speech. The article examines the philosophical problem of the transformation of language info speech, the latter is single unbroken wholeness. Norms are previously divided into d ialect and literal. Each of them is in turn d ivided into phonetic, lexical, orthographical, spelling, morphological, syntactic variety. For the disclosure of the issue there are different in terms of foreign and domestic linguistic differentiation with respect to language the language and speech rate. Particular attention is paid to the analysis of the norm in-depth by the well-known Uzbek linguists.
Tant dans le domaine de la psychologie que dans celui du traitement automatique des langues, les normes portant sur des proprietes semantiques, comme le caractere concret ou abstrait, la polarite ou le caractere emotionnel, constituent des ressources importantes. La construction manuelle de ces normes, par l’intermediaire d’evaluateurs, est couteuse, d’ou l’interet de developper des methodes de construction ou d’extension automatique. Plusieurs methodes ont ete proposees, mais elles portent sur une seule dimension: la polarite. Nous proposons de voir dans quelle mesure l’une d’entre elles peut etre etendue a six autres normes, et ce pour le francais et l’espagnol. Les experimentations confirment l’efficacite de la technique non seulement pour etendre une norme, mais egalement pour mettre en evidence des mots pour lesquels les valeurs attribuees par les evaluateurs sont sujettes a caution.
Hedges,as an important strategy and a pervasive feature in academic writing,can present claims with greater precision,limit the professional damage and give deference to the reader.However,it has often been mistaken as poor writing style and thus neglected for a long time.This paper presents a contrastive interlanguage analysis(CIA) of lexical hedges in two self-compiled corpora,i.e.Native Abstract Corpus(NAC) and Chinese Abstract Corpus(CAC).With the use of lexical hedges in English abstracts of linguistic research articles written by native speakers as the norm,this study analyzes the deviation in the use of lexical hedges by Chinese researchers—EFL learners at a higher linguistic level through comparison with a view to providing some suggestions for the teaching and learning of hedges in EAP classrooms.
Until recently, most systems performing temporal extraction and reasoning from text have focused on recognizing and normalizing temporal expressions alone, for which the TIDES annotation scheme has been adopted. Temporal awareness of a text, however, involves not only identifying the temporal expressions, but the events which these expressions anchor, as well as other events which must be ordered relative to them. Because of these broader concerns, TimeML has been developed as an annotation specification that encompasses not only temporal expressions, but all temporally relevant aspects of a text. The annotation schemes, however, are not interchangeable, resulting in incompatible corpora and accompanying extraction algorithms for each standard. In this paper, we describe an automatic migration process from the TIMEX2 tags of TIDES to the TIMEX3 tags of TimeML. This transformation procedure has been implemented and evaluated with two different corpora, obtaining 93.3 and 89.2% overall F-Measure respectively.
Background The importance and protective nature of children's early communication capacities, from birth to preschool years, in relation to later academic and social functioning is well established in the literature. Studies have shown links between early language competence (from birth to preschool) and later language, literacy, behavioural and social outcomes 1-3 as well as language, literacy and numeracy being shown to serve as key protective factors for positive life outcomes.4 The term ‘communication’ includes speech (the physical production of sounds), language (understanding and expression of spoken and written language, from sounds to words to sentences, to discourse), pragmatics (the social use of language in interactions), fluency (the smooth rhythm and pattern of talking) and voice (the production of sound through the vocal cords). 5, 6 Prelinguistic and early language development are the areas of communication that are of primary interest in this systematic review because they are the predominant aspects of communication that studies measure, when investigating the impact of parent responsiveness on children's communication development. Prelinguistic communication skills are the foundation skills that facilitate infants' communication competence.7-10 The prelinguistic period is typically from 0-12 months and skills include early vocal behaviours such as cooing and babbling,5 symbolic and functional play,(cited in 7) attention,11, 12 gestures such as facial expression,13 eye contact, turn taking, copying9 and phonetic (speech) perception.10 Language development encompasses the sub-components of sound and sound patterns (phonological development), words (lexical development), sentences and grammar (syntactic and morphological development) and the development of communicative competence, incorporating pragmatic skills (language use in a social context).5 Communication development starts from the prelinguistic period and is influenced by environmental factors including parental (particularly maternal) responsiveness and directiveness. Responsiveness refers to adults' ‘prompt, contingent, and appropriate’ (cited in 14 pp64) responses to a child's behaviours. This definition underpins the various aspects or descriptions of responsiveness that have been researched in relation to children's communicative development, for example: maternal encouragement,15 supportive parenting, 16 interpersonal timing, 17 and maternal behavioural and verbal responsiveness.14 Masur and colleagues14 also discuss the importance of considering directiveness (as well as responsiveness) when investigating the impact of parent speech and behaviour on children's communication development. Directiveness is described as being ‘characterised by attempts to command and control children's behaviour or attention’ 14(pp64) and may be supportive or intrusive in nature. Research has shown predictive relationships between parental responsiveness and directiveness and children's language development.14,16,18,19 Enhancing parent responsiveness to children's early communication can have positive effects on child development, social development, self esteem, the attachment relationship between parent and child literacy outcomes. 20-23 Hence, parents, as the primary caregivers, are in a powerful position to influence their child's communication development, and subsequent academic and social success, through the way they respond to their children from birth. The importance of this systematic review This systematic review aims to support and direct evidenced based practice and health promotion in speech pathology and related fields that work with parents and infants or children. To the reviewer's awareness, no systematic reviews on the relationship between parent responsiveness and children's communication development have been developed to date. Systematic reviews can play a role in the education of health professionals and lay people.24 A systematic review on the relationship between parental responsiveness and children's communication development could support health professionals and policy makers in easily accessing synthesised information on this topic. Prevalence data on speech-language difficulties in 2 to 4 ½ year olds has been reported as 5-8%. (cited in 25 & 26) Whilst this percentage is not categorized into causal factors, it is plausible, based on the research on parental responsiveness, that a proportion of these children have speech-language difficulties due to a reduced level of parental responsiveness in their early learning environment. Despite the established evidence regarding the importance of maternal responsiveness on children's early communication development, which in turn, influences later life outcomes, it is the reviewer's opinion that this information is not widely promoted in the general community to serve as a preventative measure. Using Gordon's operational classification of disease prevention, a universal or selective preventative measure would include public education as ‘an essential aspect of the strategy for optimal public health practice’. 27 (pp108) According to Gordon a universal preventative measure is desirable for everybody in the general population (i.e. all parents of infants or expectant parents), while a selective preventative measure is aimed at subgroups of the population who are considered to have characteristics that place them ‘at risk’ (i.e. parents who are at risk of being less responsive to their infants). This systematic review could encourage the focus of universal or selective health promotion and policy development on educating society about the benefits and importance of parent responsiveness in relation to child development outcomes. Health promotion and early parent education on the parent's role in children's early communication development could potentially reduce the number of preschool and school-age children with speech-language difficulties, hence reducing the economic, social and individual costs of this issue. This comprehensive systematic review will incorporate both quantitative and textual components. A preliminary search of the literature has found that studies on parental responsiveness and children's communication development are quantitative by nature. The textual component of this review will set the context of current thinking and action in society in relation to the quantitative component. The textual component is important because it will investigate whether the research is being put into action, or at least has a profile in society. The textual component may help to clarify the direction, if any that government needs to take regarding public education on this topic. A qualitative component of this systematic review is not included because a preliminary search did not identify any qualitative papers, and a qualitative approach is not required to answer the research question presented. Review question/objective The quantitative objective of this review is to determine the best available evidence on the relationship between parents' responsiveness to children's prelinguistic and early communication and their subsequent communication development. More specifically, the questions are: What are the attributes of parental responsiveness? That is: To delimit the attributes of parents' verbal and behavioural responsiveness and directiveness that influences children's preverbal and early communicative development. Do some attributes of parent responsiveness have more consequence to children's early communication development than others? Is the amount or frequency of parent responsiveness important? That is: Do varying levels of parent responsiveness impact differently on children's communication development? Are there parental factors (e.g. education level) within the well population that predict or influence responsiveness quality and quantity? If so, what are they? The textual objective is to identify the current social context within Australia, regarding the topic of parental responsiveness and children's communication development. More specifically, the questions are: Does current government policy on child development reflect the research evidence identified in the quantitative component of this systematic review? What is society's current awareness and standing (perception) on this topic, as identified through policy, expert and public opinion? Are there preventative universal or selective health promotion measures in place relating to the review question? If the answer to question 3 is ‘yes’, then what are they? Inclusion criteria Types of participants The quantitative component of this review will consider studies that include parents as the primary independent variable and children as the secondary, dependent variable. More specifically, the review will include studies with: 1. Parents Because a main goal of this review is to support public education for universal or selective health promotion, parents who are identified as falling within the well or at risk, but not clinically significant population will be included. Well parents refer to the general public who are not affected by current suffering. 27At risk parents may include parents whose social circumstances place them at risk of being less responsive to their children. For example, parents of low education or intellectual capacity, or of certain age. At risk parents will be included in this review because they have characteristics that place them in a position for selective preventive health promotion 27 and could provide insight into the outcomes of varying levels of parent responsiveness to children. The term clinically significant refers to parents who have clinical diagnoses that impact on their capacity to respond to their children. For example, hearing impairment, and mental illnesses such as psychoses, schizophrenia, clinical or post-natal depression. These parents are excluded from the review because they present compounding factors that are beyond the scope of this review. 2. Children Children who's language level is preverbal (i.e. prelinguistic period) up to production of early phrases (E.g.: two-word utterances) will be included in this systematic review. Based on child development norms, these ages would typically include 0 - 3 year olds, however, studies that have older cohorts will also be included, providing the earlier years are also represented within the study. Children who are typically developing, or defined as a ‘late talker’ or as having a specific speech/language issue will be included in this systematic review because these studies may reveal important information about parent responsiveness as a causal or influencing factor. Studies may or may not have control groups. Studies will not be considered for this review when children are identified as having any primary co-morbid condition such as syndromes, global developmental delays or disorders, Autism Spectrum Disorder, or hearing issue including hearing impairment and cochlear implant because this introduces too many confounding factors. Studies on bilingual children will not be included for the same reason. Consideration will be made as whether to include children who spend care time with a carer other than their primary parent/caregiver, for example, childcare. The amount of time spent in the care of persons/institutions other than their parents is important to consider because the review is examining the relationship between the parent's impact on the child through their responsiveness attributes and levels. It is beyond the scope of the review to consider the language development of children independent of their parent's responsiveness. The reviewer will examine the literature to determine the cut off point for time spent in childcare. Where this is not clear, the reviewer will contact the authors of the studies for this specific information. The textual component of this review will consider discourse and opinion reported or published by government agencies, experts, the public and media, about the systematic review question, that is of direct relevance or interest to Australia. Types of intervention(s) The quantitative component of the review will consider any studies that evaluate parent verbal and behavioural responsiveness and directiveness to their children's preverbal and or early linguistic communication. Studies may investigate parent responsiveness and or directiveness in the context of a home or clinical/education environment. The reviewer will take the environment (e.g. home, laboratory, community settings) in which studies gather their data into consideration throughout the review process. The textual component of this review will consider published and unpublished papers that describe society's and government's current attitudes and opinions regarding the topic of parental responsiveness to infant and early communication. Types of outcomes The quantitative component of this review will consider studies that include outcome measures of child prelinguistic and early language development. This includes, but is not limited to, measures of language milestones such as comprehension of first words, speech sound perception, babbling, first word production, first 50 words and first 2-word utterance. The process of the systematic review may reveal other important prelinguistic or early language outcomes, which may be considered for inclusion depending on the validity, reliability and standardisation of the tools used to obtain the data. The preferred type of assessment tools used to retrieve data about child language outcomes will be standardised language/communication assessments. However, parent reports and non-standardised assessments will also be considered for inclusion. The textual component of this review will consider discourse and opinion about the topic of parental responsiveness and children's communication development, as reported in textual or policy papers. The outcome will be the main themes and concepts identified through expert and society opinions, and government policy, in relation to the review question. Types of studies The quantitative component of the review will consider analytical epidemiological study designs including prospective and retrospective cohort studies, case control studies and analytical cross sectional studies for inclusion. Randomised control trials of parent responsiveness are not ethically possible, therefore will not be included in this review. Case series studies have not been identified in preliminary search of the topic, therefore will not be included in this review. The textual component will consider expert opinion, discussion papers, position papers, government policies and reports, conference papers, theses and dissertations, and other text relating to child development/health promotion/early education within the context and parameters of the review question. Discourse must be written in English and be of western culture. Discourse from Australia is of primary interest. Discourse from other countries that constitute western society (i.e.: the Americas, New Zealand and Western Europe)28 will only be included where it has been shown to be of interest to Australia. For example, an Australian expert has commented on a paper from another Western country. Search strategy The search strategy aims to find both published and unpublished studies. A three-step search strategy will be utilised for each component of this review. An initial limited search of PubMed and CINAHL will be undertaken followed by analysis of the text words contained in the title and abstract, and of the index terms used to describe article. A second search using all identified keywords and index terms will then be undertaken across all included databases. Where necessary, terms and indexing language will be adjusted to search the other databases listed. This process will be done in close consultation with the Research Librarian for Mental Health, Psychiatry, Psychology, University of Adelaide. Thirdly, the reference list of all identified reports and articles will be searched for additional studies. Studies published in English will be considered for inclusion in this review. As there are no other identified systematic reviews on this topic, any quantitative studies within an unlimited timeframe will be considered for inclusion in this review, in order to increase the breadth of the results and so not to miss any pertinent earlier studies. To keep textual information of current opinion and policy relevant and up to date, the timeframe will be the past 10 years (2002 - 2012). The databases to be searched include: PubMed PsycINFO CINAHL Embase Scopus Web of Science Mednar Proquest Dissertations and Theses Index to Theses Australian Digital Theses Program The Networked Digital Library of Theses and Dissertations (NDLDT) Keywords and concepts to be used for the initial search of PubMed and CINAHL will include:Table: No Caption available.The search for textual information will also include relevant websites in the English language, related to child development, literacy, parent-infant attachment, government policy on early childhood development and education, and media releases relating to the review question. An initial search to identify a comprehensive list of relevant websites for grey literature will be done through the Google search engine using initial key words seen above and additional keywords including: Government policy Early childhood Parent education Parent training Infant mental health Expert opinion(s) Individual countries (eg Australia, New Zealand, Canada, America, United Kingdom) Examples of potential grey literature sites include: Australian Government Department of Health and Ageing Australian Government Department of Education, Employment and Workplace Relations Council of Australian Governments The Hanen Centre. Speech and Language Development for Children Assessment of methodological quality Quantitative papers selected for retrieval will be assessed by two independent reviewers for methodological validity prior to inclusion in the review using standardised critical appraisal instruments from the Joanna Briggs Institute Meta Analysis of Statistics Assessment and Review Instrument (JBI-MAStARI) (Appendix I). Textual papers selected for retrieval will be assessed by two independent reviewers for authenticity prior to inclusion in the review using standardised critical appraisal instruments from the Joanna Briggs Institute Narrative, Opinion and Text Assessment and Review Instrument (JBI-NOTARI) (Appendix I). Any disagreements that arise between the reviewers will be resolved through discussion, or with a third reviewer. Data collection Quantitative data will be extracted from papers included in the review using the standardised data extraction tool from JBI-MAStARI (Appendix II). Textual data will be extracted from papers included in the review using the standardised data extraction tool from JBI-NOTARI (Appendix II). The data extracted will include specific details about the interventions, populations, study methods and outcomes of significance to the review question and specific objectives. Data synthesis Quantitative papers will, where possible, be pooled in statistical meta-analysis using JBI-MAStARI. All results will be subject to double data entry. Effect sizes expressed as relative risk for cohort studies and odds ratio for case control studies (for categorical data) and weighted mean differences (for continuous data) and their 95% confidence intervals will be calculated for analysis. A Random effects model will be used and heterogeneity will be assessed statistically using the standard Chi-square. Where statistical pooling is not possible the findings will be presented in narrative form including tables and figures to aid in data presentation where appropriate. Textual papers will, where possible be pooled using JBI-NOTARI. This will involve the aggregation or synthesis of conclusions to generate a set of statements that represent that aggregation, through assembling and categorising these conclusions on the basis of similarity in meaning. These categories are then subjected to a meta-synthesis in order to produce a single comprehensive set of synthesised findings that can be used as a basis for evidence-based practice. Where textual pooling is not possible the conclusions will be presented in narrative form. Conflicts of interest The primary reviewer is not aware of any conflicts of interest at the time of submitting the systematic review protocol. Acknowledgements The primary reviewer would like to acknowledge the support of the secondary reviewer, Matthew Kowald; her principal supervisor, Dr Aye Aye Gyi from the Joanna Briggs Institute, University of Adelaide; her associate supervisor, Dr Debbie James from Research and Evaluation Unit, Children, Youth and Women's Health Service, and Maureen Bell, Research Librarian for Mental Health, Psychiatry, Psychology, University of Adelaide. As this systematic review forms partial submission for the award of Masters of Clinical Sciences degree, a secondary reviewer will be used for critical appraisal only.
Many languages in Africa are written using Latin-based scripts, but often with extra diacritics (e.g. dots below in Igbo: \({\d i}, {\d o}, {\d u}\)) or modifications to the letters themselves (e.g. open vowels “e” and “o” in Lingala: ɛ, ɔ). While it is possible to render these characters accurately in Unicode, oftentimes keyboard input methods are not easily accessible or are cumbersome to use, and so the vast majority of electronic texts in many African languages are written in plain ASCII. We call the process of converting an ASCII text to its proper Unicode form unicodification. This paper describes an open-source package which performs automatic unicodification, implementing a variant of an algorithm described in previous work of De Pauw, Wagacha, and de Schryver. We have trained models for more than 100 languages using web data, and have evaluated each language using a range of feature sets.
The Internet has facilitated both the dissemination of anonymous texts as well as easy “borrowing” of ideas and words of others. This has raised a number of important questions regarding authorship. Can we identify the anonymous author of a text by comparing the text with the writings of known authors? Can we determine if a text, or parts of it, has been plagiarized? Such questions are clearly of both academic and commercial importance.
Reaction time tasks are used widely in basic and applied psychology. There is a need for an easy-to-use, freely available programme that can run simple and choice reaction time tasks with no special software. We report the development of, and make available, the Deary-Liewald reaction time task. It is initially tested here on 150 participants, aged from 18 to 80, alongside another widely used reaction time device and tests of fluid and crystallised intelligence and processing speed. The new task’s parameters perform as expected with respect to age and intelligence differences. The new task’s parameters are reliable, and have very high correlations with the existing task. We also provide instructions for downloading and using the new reaction time programme, and we encourage other researchers to use it.
The reconstruction of standardized texts in the Prague Dependency Treebank of Spoken Czech enables the comparison of authentic spoken utterances and standardized texts The authors concentrate on the questions: What does the syntactic identity of the Czech spoken and written texts consist of? What syntactic constructions are „natural“ in the spoken and in the written text? What is the difference in the density of the cohesive links, in the explicit and implicite relations between units?
Humans reached present-day Island Southeast Asia (ISEA) in one of the first major human migrations out of Africa. Population movements in the millennia following this initial settlement are thought to have greatly influenced the genetic makeup of current inhabitants, yet the extent attributed to different events is not clear. Recent studies suggest that south-to-north gene flow largely influenced present-day patterns of genetic variation in Southeast Asian populations and that late Pleistocene and early Holocene migrations from Southeast Asia are responsible for a substantial proportion of ISEA ancestry. Archaeological and linguistic evidence suggests that the ancestors of present-day inhabitants came mainly from north-to-south migrations from Taiwan and throughout ISEA approximately 4,000 years ago. We report a large-scale genetic analysis of human variation in the Iban population from the Malaysian state of Sarawak in northwestern Borneo, located in the center of ISEA. Genome-wide s)
Background: Diversity patterns of livestock species are informative to the history of agriculture and indicate uniqueness of breeds as relevant for conservation. So far, most studies on cattle have focused on mitochondrial and autosomal DNA variation. Previous studies of Y-chromosomal variation, with limited breed panels, identified two Bos taurus (taurine) haplogroups (Y1 and Y2; both composed of several haplotypes) and one Bos indicus (indicine/zebu) haplogroup (Y3), as well as a strong phylogeographic structuring of paternal lineages. Methodology and Principal Findings: Haplogroup data were collected for 2087 animals from 138 breeds. For 111 breeds, these were resolved further by genotyping microsatellites INRA189 (10 alleles) and BM861 (2 alleles). European cattle carry exclusively taurine haplotypes, with the zebu Y-chromosomes having appreciable frequencies in Southwest Asian populations. Y1 is predominant in northern and north-western Europe, but is also observed in several Ibe)
It has been shown that the human genome contains extensive copy number variations (CNVs). Investigating the medical and evolutionary impacts of CNVs requires the knowledge of locations, sizes and frequency distribution of them within and between populations. However, CNV study of Chinese minorities, which harbor the majority of genetic diversity of Chinese populations, has been underrepresented considering the same efforts in other populations. Here we constructed, to our knowledge, a first CNV map in seven Chinese populations representing the major linguistic groups in China with 1,440 CNV regions identified using Affymetrix SNP 6.0 Array. Considerable differences in distributions of CNV regions between populations and substantial population structures were observed. We showed that ∼35% of CNV regions identified in minority ethnic groups are not shared by Han Chinese population, indicating that the contribution of the minorities to genetic architecture of Chinese population could not)
Abstract Prior research on relative clauses (RCs) in Mandarin Chinese has led to conflicting results regarding ease of processing subject-extracted RCs (SRCs) versus object-extracted RCs (ORCs) and has often used animacy configurations that are rare in corpora. Building on animacy patterns observed in a corpus, we used self-paced reading to explore how animacy influences real-time processing of Chinese RCs. Experiment 1 tested SRCs, and found marginal facilitation effects with animate heads (subjects) and inanimate objects. Experiment 2 tested ORCs and found significant facilitation effects with inanimate head (objects). Experiment 3 showed that when the subject is animate and the object inanimate, ORCs are as easy to process as SRCs, but when the subject is inanimate and the object is animate, SRCs are processed faster. Thus, the animacy of the head and the embedded noun must be taken into account when evaluating processing ease. Keywords: AnimacyRelative clause (RC)Mandarin ChineseProcessing Acknowledgments We would like to thank audiences at the 14th Annual Conference on Architectures and Mechanisms for Language Processing (AMLaP), the 2008 Western Conference on Linguistics (WECOL) and the 83rd Annual Meeting of the Linguistics Society of America (LSA), where earlier versions of some of this research were presented. Preliminary analyses of some of the data reported here appeared in Wu, Kaiser, and Andersen (2010). The stimuli used in this research are a revised version of the stimuli used in Wu (2009). We thank Yanan Sheng for assistance in the stimulus revision and running of participants, Xiaomei Qiao, and Tangfeng Yang for assistance in carrying out norming studies, Mei Li for providing facilities in running Experiment 3 at Tongji University, and Rudolf Troike for help with finalising the translations of our Chinese stimuli. This research was partially supported by a project sponsored by the Scientific Research Foundation for Returned Overseas Chinese Scholars, State Education Ministry, and by a grant from the Shanghai Municipal Philosophy and Social Sciences Foundation (2010BYY003) to the first author. Notes 1As a reviewer pointed out, the relation between head animacy and RC-type is clear with object RCs (which tend to occur with inanimate heads), but less so with subject RCs. Indeed, in Mak et al.'s (2002, pp. 54–55) German corpus, the 144 subject-extracted RCs have inanimate heads almost as frequently as animate heads: 57% animate heads and 43% inanimate heads. Also, in Roland et al.'s (2007, p. 357) analysis of the English-language Brown corpus, 47% of 100 randomly-selected subject-extracted RCs have inanimate heads. However, existing corpus data from Chinese suggest that subject RCs' head animacy patterns (at least in Chinese) may vary depending on the grammatical role of the RC's head noun. For Chinese, Pu (2007, p. 45) and Wu (2009) found that (1) when SRCs modify sentential subjects, animate heads significantly outnumber inanimate heads, but (2) when SRCs modify sentential objects, there is no particular bias toward animate or inanimate heads. 2The term "experiencer" refers to a change of psychological state on a human participant caused by someone or something in the context of certain intransitive verbs (e.g., win, die); experience-theme verbs (e.g., love, discover, like); or causer-experiencer verbs (e.g., please, amuse, amaze, and annoy). 3Lin and Garnsey (Citation2010) manipulated animacy in their stimuli, but they also topicalised their RCs to a sentence-initial position and used null head nouns. Headless RCs and topicalisation in Mandarin normally occur only when supportive discourse contexts are given, but their stimuli were presented in isolation. Thus their stimuli had a marked structure, which may have complicated their results. 4One reviewer pointed out that the percentage of RCs where both nouns have the same animacy is 28% in Mak et al.'s (2002) Dutch corpus and 40% in their German corpus. However, viewed from another perspective, this means that the percentage of RCs with contrastive animacy configuration is 72% in Dutch and 60% in German, a pattern similar to Wu's (2009) corpus analysis. Furthermore, at least in Wu's (2009) analyses of Chinese Treebank Corpus, RCs with matched animacy (double-animates or double-inanimates) occurred significantly less frequently than RCs with nonmatched animacy (p'<.05). 5In Experiment 1, the log frequencies for the different verbs and for the embedded nouns were matched. The mean log frequencies for the verbs from the SUBTLEX-CH are as follows: 3.14 for Oi-Sa and Oa-Sa, 3.32 for Oa-Si and Oi-Si. The frequencies do not differ significantly, F(3, 76) = 0.1069, p=.9558. The mean log frequencies for the verbs from the 2008 frequency dictionary are as follows: 9.46 for Oi-Sa and Oa-Sa, 9.04 for Oa-Si and Oi-Si. These frequencies also do not differ significantly, F(3, 78) = 0.6015, p=.616. The mean log frequencies for the embedded nouns from the SUBTLEX-CH are as follows: 3.07 for Oi-Sa and Oi-Si, 3.28 for Oa-Sa and Oa-Si. The frequencies do not differ significantly, F(3, 78) = 0.1726, p=.9146. The log frequencies for the embedded nouns from the 2008 frequency dictionary are as follows: 9.38 for Oi-Sa and Oi-Si, 9.36 for Oa-Sa and Oa-Si. The frequencies also do not differ significantly, F(3, 82) = 0.0087, p=.9989. Log frequencies for the head nouns were matched for the 2008 dictionary (means: 8.5 for Oi-Sa and Oa-Sa, 9.05 for Oa-Si and Oi-Si). According to this corpus, the frequencies of the different head nouns do not differ significantly, F(3, 72) = 1.7263, p=.1692. However, according to the SUBTLEX-CH corpus, the log frequencies for the head nouns are not matched [means: 4.32 for Oi-Sa and Oa-Sa, 2.88 for Oa-Si and Oi-Si; F(3, 74) = 4.8757, p=.0038]. As said, the frequency check reported above are based on an incomplete list of words that have their frequencies listed in either resource. 6At the sentence-initial RC-verb position (pos 1, e.g., raokai "bypass"), there was a marginal main effect of Head Animacy (t=1.84, p=.0663), and a marginal interaction between Head Animacy and Embedded-noun Animacy (t=−1.8, p=.073). However, given that this is the first word region, these weak effects are probably due to lexical differences. 7In Experiment 3, the log frequencies were matched for the verbs, but not for the embedded nouns and for the head nouns. The mean log frequencies of the verbs from SUBTLEX-CH are as follows: 3.83 for SRCs with animate heads and for ORCs with inanimate heads, 3.41 for SRCs with inanimate heads and for ORCs with animate heads. The frequencies do not differ significantly, F(3, 82) = 0.2254, p=.8785. The mean log frequencies of the verb from the 2008 frequency dictionary are as follows: 9.26 for SRCs with inanimate heads and for ORCs with animate heads, 9.19 for SRCs with inanimate heads and for ORCs with animate heads. These frequencies also do not differ significantly, F(3, 76) = 0.1274, p=.9436. The log frequencies for the embedded nouns from the SUBTLEX-CH are as follows: 2.86 for SRCs with animate heads and for ORCs with animate heads, 4.03 for SRCs with inanimate heads and for ORCs with inanimate heads. The differences in frequencies are marginally significant, F(3, 74) = 2.581, p=.062. The log mean frequencies of the embedded nouns from the 2008 frequency dictionary are as follows: 9.52 for SRCs with animate heads and for ORCs with animate heads, 9.11 for SRCs with inanimate heads and for ORCs with inanimate heads). These frequencies differ significantly, F(3, 76) = 2.833, p=.044. Reversely for the head nouns, their log frequencies for SRCs with inanimate heads and for ORCs with inanimate heads are more frequent than the log frequencies for SRCs with animate heads and for ORCs with animate heads. However, because we used a Latin-square design such that the different nouns rotated through the different conditions, we do not think this affects our results.
1 IntroductionThis paper deals with contrastive analysis of computer terms in Slovenian and Serbian language. Problematic aspects of computer jargon in two similar and related languages are considered in context of cultural and linguistic differences noticed during translation from Slovenian to Serbian. The analytic framework is grounded in concrete translator's challenges of some lexical issues in translation process. As George Steiner claims that ordinary language, literally at every moment, mutates in many forms: words enter as old words lapse. Grammatical conventions are changed under pressure of idiomatic use or by cultural ordinance. (Steiner 1998:19)When we are translating a text into another language we are confronted with cultural differences and, for translator, translating professional texts represents a big challenge. He has to be a mediator from professional idiom in one language to professional idiom in other languages. In era of rapidly changing computer technology, spreading new ways of global communication and developing of social networks, we are surrounded by computers and we use some of computer terms in our everyday lives. But, different cultures have different language politics and they relate differently to translating of terms taken from English as modern-day global language. As has been stated, all speech communities have experienced need to modernize and keep abreast of developments in science, technology, etc. (Winford 2007:37), and the spread of English loanwords into many languages across globe may fill gaps in lexicon (Winford 2007:38), but this massive lexical borrowing caused by innovation in field of computer science and technology may be regulated. Lexical borrowing is determined by norms, network structures and language ideology of community: Loyalty to one's native language and pride in its autonomy may encourage resistance to any foreign incursions (Winford 2007:41).Lexical borrowing into lexicon can have consequences on phonology, morphology or on lexicon itself. New lexical fields in recipient language may be created by using computeresse (Winford 2007:58).We attend to show correlation between language politic and accepting of English IT terms. There are two opposite attitudes about this kind of lexical borrowing in professional language: a) usage of Anglicisms is recommended because of globalisation; they provide understanding between experts from different countries; English terms in computer jargon have same purpose as Latin terms in field of medicine and natural science at all; b) translation, or at least adaptation, of these terms is necessary, they become a part of lexical system of recipient language and its standardisation is strongly recommended. We will show which of these attitudes is more common among users of computer jargon.2 Material and methodsThe research is carried out on basis of data excerpted during translation of graduate work Odnos studentov do zasebnosti na spletni skupnosti Facebook and term paper Etika in internet: oglasevanje na internetu, both published on Internet page of Faculty of Economy, University of Ljubljana, and supplemented by examples collected from internet dictionaries of computer terms (http://www.islovar.org, http://sverapoj.nedohodnik.net/gloss/, http://www.mikroknjiga.rs/pub/rmk/index.php). Islovar is project of Linguistic Section of Slovenian Society INFORMATIKA, established in year 2000, with purpose of increasing concern for professional language, equalizing IT terms and recruiting experts and common people for its modelling, through free interactive use and interactive development of dictionary. This internet dictionary has a large repertoire of computer terms and number of entries is constantly growing. …
Despite tremendous advances in artificial language synthesis, no machine has so far succeeded in deceiving a human. Most research focused on analyzing the behavior of "good" machine. We here choose an opposite strategy, by analyzing the behavior of "bad" humans, i.e., humans perceived as machine. The Loebner Prize in Artificial Intelligence features humans and artificial agents trying to convince judges on their humanness via computer-mediated communication. Using this setting as a model, we investigated here whether the linguistic behavior of human subjects perceived as non-human would enable us to identify some of the core parameters involved in the judgment of an agents' humanness. We analyzed descriptive and semantic aspects of dialogues in which subjects succeeded or failed to convince judges of their humanness. Using cognitive and emotional dimensions in a global behavioral characterization, we demonstrate important differences in the patterns of behavioral expressiveness of the)
Neuropsychological and imaging studies have shown that the left supramarginal gyrus (SMG) is specifically involved in processing spatial terms (e.g. above, left of), which locate places and objects in the world. The current fMRI study focused on the nature and specificity of representing spatial language in the left SMG by combining behavioral and neuronal activation data in blind and sighted individuals. Data from the blind provide an elegant way to test the supramodal representation hypothesis, i.e. abstract codes representing spatial relations yielding no activation differences between blind and sighted. Indeed, the left SMG was activated during spatial language processing in both blind and sighted individuals implying a supramodal representation of spatial and other dimensional relations which does not require visual experience to develop. However, in the absence of vision functional reorganization of the visual cortex is known to take place. An important consideration with respec)
In contrast to quantity processing, up to date, the nature of ordinality has received little attention from researchers despite the fact that both quantity and ordinality are embodied in numerical information. Here we ask if there are two separate core systems that lie at the foundations of numerical cognition: (1) the traditionally and well accepted numerical magnitude system but also (2) core system for representing ordinal information. We report two novel experiments of ordinal processing that explored the relation between ordinal and numerical information processing in typically developing adults and adults with developmental dyscalculia (DD). Participants made "ordered" or "non-ordered" judgments about 3 groups of dots (non-symbolic numerical stimuli; in Experiment 1) and 3 numbers (symbolic task: Experiment 2). In contrast to previous findings and arguments about quantity deficit in DD participants, when quantity and ordinality are dissociated (as in the current tasks), DD parti)
Although number words are common in everyday speech, learning their meanings is an arduous, drawn-out process for most children, and the source of this delay has long been the subject of inquiry. Children begin by identifying the few small numerosities that can be named without counting, and this has prompted further debate over whether there is a specific, capacity-limited system for representing these small sets, or whether smaller and larger sets are both represented by the same system. Here we present a formal, computational analysis of number learning that offers a possible solution to both puzzles. This analysis indicates that once the environment and the representational demands of the task of learning to identify sets are taken into consideration, a continuous system for learning, representing and discriminating set-sizes can give rise to effective discontinuities in processing. At the same time, our simulations illustrate how typical prenominal linguistic constructions (''the)
The human population history in Southeast Asia was shaped by numerous migrations and population expansions. Their reconstruction based on archaeological, linguistic or human genetic data is often hampered by the limited number of informative polymorphisms in classical human genetic markers, such as the hypervariable regions of the mitochondrial DNA. Here, we analyse housekeeping gene sequences of the human stomach bacterium Helicobacter pylori from various countries in Southeast Asia and we provide evidence that H. pylori accompanied at least three ancient human migrations into this area: i) a migration from India introducing hpEurope bacteria into Thailand, Cambodia and Malaysia; ii) a migration of the ancestors of Austro-Asiatic speaking people into Vietnam and Cambodia carrying hspEAsia bacteria; and iii) a migration of the ancestors of the Thai people from Southern China into Thailand carrying H. pylori of population hpAsia2. Moreover, the H. pylori sequences reflect iv) the migra)
We measured perceived depth from the optic flow (a) when showing a stationary physical or virtual object to observers who moved their head at a normal or slower speed, and (b) when simulating the same optic flow on a computer and presenting it to stationary observers. Our results show that perceived surface slant is systematically distorted, for both the active and the passive viewing of physical or virtual surfaces. These distortions are modulated by head translation speed, with perceived slant increasing directly with the local velocity gradient of the optic flow. This empirical result allows us to determine the relative merits of two alternative approaches aimed at explaining perceived surface slant in active vision: an ''inverse optics'' model that takes head motion information into account, and a probabilistic model that ignores extra-retinal signals. We compare these two approaches within the framework of the Bayesian theory. The ''inverse optics'' Bayesian model produces veridi)
In this paper, we present our attempts to design and implement a large-coverage computational grammar for the Persian language based on the Generalized Phrase Structured Grammar (GPSG) model. This grammatical model was developed for continuous speech recognition (CSR) applications, but is suitable for other applications that need the syntactic analysis of Persian. In this work, we investigate various syntactic structures relevant to the modern Persian language, and then describe these structures according to a phrase structure model. Noun (N), Verb (V), Adjective (ADJ), Adverb (ADV), and Preposition (P) are considered basic syntactic categories, and X-bar theory is used to define Noun phrases, Verb phrases, Adjective phrases, Adverbial phrases, and Prepositional phrases. However, we have to extend Noun phrase levels in X-bar theory to four levels due to certain complexities in the structure of Noun phrases in the Persian language. A set of 120 grammatical rules for describing different phrase structures of Persian is extracted, and a few instances of the rules are presented in this paper. These rules cover the major syntactic structures of the modern Persian language. For evaluation, the obtained grammatical model is utilized in a bottom-up chart parser for parsing 100 Persian sentences. Our grammatical model can take 89 sentences into account. Incorporating this grammar in a Persian CSR system leads to a 31% reduction in word error rate.
We tested whether the intervening time between multiple glances influences the independence of the resulting visual percepts. Observers estimated how many dots were present in brief displays that repeated one, two, three, four, or a random number of trials later. Estimates made farther apart in time were more independent, and thus carried more information about the stimulus when combined. In addition, estimates from different visual field locations were more independent than estimates from the same location. Our results reveal a retinotopic serial dependence in visual numerosity estimates, which may be a mechanism for maintaining the continuity of visual perception in a noisy environment. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for indivi)
Exogenous neurotrophin delivery to the deaf cochlea can prevent deafness-induced auditory neuron degeneration, however, we have previously reported that these survival effects are rapidly lost if the treatment stops. In addition, there are concerns that current experimental techniques are not safe enough to be used clinically. Therefore, for such treatments to be clinically transferable, methods of neurotrophin treatment that are safe, biocompatible and can support long-term auditory neuron survival are necessary. Cell transplantation and gene transfer, combined with encapsulation technologies, have the potential to address these issues. This study investigated the survival-promoting effects of encapsulated BDNF over-expressing Schwann cells on auditory neurons in the deaf guinea pig. In comparison to control (empty) capsules, there was significantly greater auditory neuron survival following the cell-based BDNF treatment. Concurrent use of a cochlear implant is expected to result in )
This paper deals with the ways in which minority students in the Danish public school system bring mono-lingually based norms into their poly-lingual peer group interaction. In sequential micro-analyses of interaction we show how the students use the voice of an authority in their reproduction and negotiation of linguistic norms. We base our analyses on the Bakhtinian concept of double-voicing, the Goffmanian concept of keying and Tholander's concept of subteaching. We discuss in detail the relation between the local practices of the students and the linguistic norms expressed in broader society, for instance in newspaper editorials and government papers. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
The language of a speech community can only act as an identity marker for all of its speakers if linguistic norms are widely shared and if a minimal number of language varieties are spoken. This article examines briefly how a linguistic norm came to serve the whole of Iceland and how a situation of relative linguistic homogeneity was maintained for centuries. Sociolinguistic theory tells us that the speech community that we can reconstruct for early Iceland should lead to the establishment and maintenance of local norms. However, Iceland, arguably monodialectal, was certainly characterized by long-term linguistic homogeneity and remained a society where nucleated settlements barely formed over a thousand-year period. Scholars have argued that a mixture of dialects leveled shortly after the settlement of Iceland in the ninth century (Settlement). Studies show that dialect leveling requires dialect mixing, the convergence of people on one place, and sustained linguistic contact between the speakers. The settlement pattern of Iceland is indicative of population divergence (not convergence) and there is limited evidence of sustained contact. It is therefore proposed that the dialect leveling might be linked instead with significant population movements and social upheaval in mainland Scandinavia in the immediate pre-Viking period. The variety of Norse that was taken westward across the Atlantic might itself already have been the result of several earlier stages of mixing and koineization. It is only by combining linguistic, historical, and archaeological knowledge that this problem of how one linguistic norm came to serve the whole of Iceland can be understood. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Reviews the book, A Better Pencil: Readers, Writers, and the Digital Revolution by Dennis Baron (2009). This book addresses the impact of digital technologies on reading and writing practices by contextualizing these developments within the communication technologies that have preceded them. Though computers are often blamed for a perceived deterioration of culture and language, Baron notes that scorn and mistrust of communication technologies is not a new phenomenon; in fact, writing, the printing press, and the telegraph were all greeted at first with skepticism. And despite the pervasive anxiety that 'internet speech' is resulting in the destruction of grammar and spelling, Baron observes that online communities enforce their own linguistic norms. Rather than a linguistic free-for-all, Baron argues, 'spelling counts' in virtual spaces just as it does offline. In sum, Baron accounts for both the benefits and pitfalls of the digital revolution and analyzes it as merely another set of technological innovations, usefully situating them within a history of writing technologies. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Every language is a multiple system which includes different forms of linguistic reality. According to one of Languages for Specific Purposes (LSP) definitions it implies a language aimed at satisfying some professional needs. LSP is the result of communication needs of professionals using the language. Languages for Specific Purposes are codes different from the standard language, which have their rules. Some of the features of LSP include specific vocabulary and certain grammatical and lexical means. Everyday language cannot satisfy all the needs so loan words as well as loan translations are widely accepted. Medical language does not differ form other LSP as its vocabulary is based on certain principles. The aim of this paper is to analyze the standardization of medical terminology based on principles which will serve as a guideline for doctors and linguists in their common attempts to standardize medical language in order to avoid misunderstandings. The corpus is based on scientific, professional and popular articles and the analysis will show to which extent medical language is affected by the norms.
In this article, I undertake a qualitative analysis of third-person direct-object anaphoric reference in a corpus of Brazilian TV evening news programmes. A comparison of my results to previous studies (Bagno 2005, Duarte 1989, Schwenter and Silva 2003) reveals surprising findings in that figures for anaphoric pronouns (both clitics and tonic pronouns), as well as null objects, are extremely low. Instead, lexical NPs and passive constructions are used to establish anaphoric reference. While the use of lexical NPs to establish anaphoric reference has been analysed before, the use of passive constructions for this purpose has not been previously observed. I argue that the described pattern can be explained when taking into account the audiovisual quality of television combined with an audience design (Bell 1984) that aims to underline the seriousness of quality reporting.
This work aims to improve an N-gram-based statistical machine translation system between the Catalan and Spanish languages, trained with an aligned Spanish–Catalan parallel corpus consisting of 1.7 million sentences taken from El Periódico newspaper. Starting from a linguistic error analysis above this baseline system, orthographic, morphological, lexical, semantic and syntactic problems are approached using a set of techniques. The proposed solutions include the development and application of additional statistical techniques, text pre- and post-processing tasks, and rules based on the use of grammatical categories, as well as lexical categorization. The performance of the improved system is clearly increased, as is shown in both human and automatic evaluations of the system, with a gain of about 1.1 points BLEU observed in the Spanish-to-Catalan direction of translation, and a gain of about 0.5 points in the reverse direction. The final system is freely available online as a linguistic resource.
Studies in interlanguage pragmatics have shown that L2 learners’ proficiency has an influence on the occurrences of L1 pragmatic transfer. However, questions remain whether the relationship between L1 pragmatic transfer and L2 proficiency is positive or negative. This paper is designed to study L1 pragmatic transfer in requests made by Chinese learners of English at low L2 proficiency level and at high L2 proficiency level and how L1 pragmatic transfer is related to their L2 proficiency. Ten low proficiency learners of English, ten high proficiency learners of English?ten native speakers of English and ten native speakers of Chinese participate in this study. Requests are collected by means of a discourse completion test questionnaire and are analysed in terms of requestive semantic formulas based on the taxonomy of request strategies, internal modifiers and external modifiers. The research results reveal that L1 pragmatic transfer decreases with the increase of L2 proficiency such as learners’ use of direct strategies, lexical and phrasal downgraders, imperatives and grounder and no clear relationship is found between L1 pragmatic transfer and L2 proficiency in terms of the other request strategies, internal modifiers and external modifiers. These results provide partial support to negative correlation hypothesis —high proficiency L2 learners are less likely to transfer their native language pragmatic norms since they have enough control over L2.
In this paper, we analyze the behaviour of Singular Value Decomposition in a number of word similarity extraction tasks, namely acquisition of translation equivalents from comparable corpora. Special attention is paid to two different aspects: computational efficiency and extraction quality. The main objective of the paper is to describe several experiments comparing methods based on Singular Value Decomposition (SVD) to other strategies. The results lead us to conclude that SVD makes the extraction less computationally efficient and much less precise than other more basic models for the task of extracting translation equivalents from comparable corpora.
Measuring talkativeness is of interest to several areas of research. However, there are few brief, validated measures available. We examined test-retest reliability, inter-relationships and convergent/divergent validity for five brief measures of verbal productivity. Nineteen men and 32 women participated in four sessions, completing five speech tasks that varied in demand, purpose of speech and sociability. Several potential metrics (word count, duration and rate) were examined. All tasks except a novel Unprompted Speech task demonstrated good word count test-retest reliability (interclass correlation coefficients from .71 to .85). Factor analysis revealed low-demand, non-functional tasks formed one factor (“Voluntary Talkativeness”), while higher demand tasks formed a second factor (“Speech Ability”). This finding and examination of relationships with IQ, personality and gender indicate “Voluntary Talkativeness” is not wholly accounted for by verbal ability, and is only weakly related to self-reported personality. Recommendations for the measurement of “Voluntary Talkativeness” are made.
Fake content is flourishing on the Internet, ranging from basic random word salads to web scraping. Most of this fake content is generated for the purpose of nourishing fake web sites aimed at biasing search engine indexes: at the scale of a search engine, using automatically generated texts render such sites harder to detect than using copies of existing pages. In this paper, we present three methods aimed at distinguishing natural texts from artificially generated ones: the first method uses basic lexicometric features, the second one uses standard language models and the third one is based on a relative entropy measure which captures short range dependencies between words. Our experiments show that lexicometric features and language models are efficient to detect most generated texts, but fail to detect texts that are generated with high order Markov models. By comparison our relative entropy scoring algorithm, especially when trained on a large corpus, allows us to detect these “hard” text generators with a high degree of accuracy.
In today’s digital multilingual world, language technology is crucial for providing access to information and opportunities for economic development. With approximately two thousand different languages, Africa is a multilingual continent par excellence, presenting acute challenges for those seeking to promote and use African languages in the areas of business development, education and relief aid. In recent times a number of researchers and institutions, both from Africa and elsewhere, have come forward to share the common goal of developing capabilities in language technology for African languages. In 2009 and 2010, the first two workshops on African Language Technology were organized (De Pauw et al. 2009, 2010a) as a forum to bring together a wide range of researchers working in this domain.
Correlational research investigating the relationship between scores on self-report imagery questionnaires and measures of social desirable responding has shown only a weak association. However, researchers have argued that this research may have underestimated the size of the relationship because it relied primarily on the Marlowe–Crowne scale (MC; Crowne & Marlowe, Journal of Consulting Psychology, 24, 349–354, 1960), which loads primarily on the least relevant form of social desirable responding for this particular context, the moralistic bias. Here we report the analysis of data correlating the Vividness of Visual Imagery Questionnaire (VVIQ; Marks, Journal of Mental Imagery, 19, 153–166, 1973) with the Balanced Inventory of Desirable Responding (BIDR; Paulhus, 2002) and the MC scale under anonymous testing conditions. The VVIQ correlated significantly with the Self-Deceptive Enhancement (SDE) and Agency Management (AM) BIDR subscales and with the MC. The largest correlation was with SDE. The ability of SDE to predict VVIQ scores was not significantly enhanced by adding either AM or MC. Correlations between the VVIQ and BIDR egoistic scales were larger when the BIDR was continuously rather than dichotomously scored. This analysis indicates that the relationship between self-reported imagery and social desirable responding is likely to be stronger than previously thought.
The study is focused on how to make use of the lexical database Pralex for a processing of nominal entries, how to describe a particular specific phenomenon or a partial issue in an entry form as well as which nominal data are included (mandatory items). To clarify this, the authors follow the order of individual parts of an entry form. Special attention is paid to the issue of homonymy, variation, explanation of a meaning and division of polysemantic entries.
The paper addresses the core problems at building the communication competencies of foreign student and justifies the language culture course as mean to build the required skills, the paper also proposes methodology of teaching those skills in the lexical norms study.