Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The Kashmiri population is an ethno-linguistic group that resides in the Kashmir Valley in northern India. A longstanding hypothesis is that this population derives ancestry from Jewish and/or Greek sources. There is historical and archaeological evidence of ancient Greek presence in India and Kashmir. Further, some historical accounts suggest ancient Hebrew ancestry as well. To date, it has not been determined whether signatures of Greek or Jewish admixture can be detected in the Kashmiri population. Using genome-wide genotyping and admixture detection methods, we determined there are no significant or substantial signs of Greek or Jewish admixture in modern-day Kashmiris. The ancestry of Kashmiri Tibetans was also determined, which showed signs of admixture with populations from northern India and west Eurasia. These results contribute to our understanding of the existing population structure in northern India and its surrounding geographical areas. [ABSTRACT FROM AUTHOR], Copyright)
The increasing need of automated analyzing web texts especially the short texts on Social Network Services (SNS) brings new demands of computerized text analysis instruments. The psychometric properties are the basis of the extensive use of these instruments such as the Linguistic Inquiry and Word Count (LIWC). For this study, Sina Weibo statuses were analyzed via rater coding and Simplified Chinese version of LIWC (SCLIWC), in order to evaluate the validity of SCLIWC in detecting psychological expressions in Weibo statuses (n = 60) and in identifying the psychological meaning of a single Weibo status (n = 11). Significant correlations between human ratings and SCLIWC scores and the high sensitivities of capturing single statuses with certain expressions identified by raters, proved the validity of SCLIWC in detecting psychological expressions. The results also suggested that, the efficiency of SCLIWC in detecting psychological expressions of SNS short texts could be higher if using s)
Multilingualism is common offline, but we have a more limited understanding of the ways multilingualism is displayed online and the roles that multilinguals play in the spread of content between speakers of different languages. We take a computational approach to studying multilingualism using one of the largest user-generated content platforms, Wikipedia. We study multilingualism by collecting and analyzing a large dataset of the content written by multilingual editors of the English, German, and Spanish editions of Wikipedia. This dataset contains over two million paragraphs edited by over 15,000 multilingual users from July 8 to August 9, 2013. We analyze these multilingual editors in terms of their engagement, interests, and language proficiency in their primary and non-primary (secondary) languages and find that the English edition of Wikipedia displays different dynamics from the Spanish and German editions. Users primarily editing the Spanish and German editions make more compl)
An alternative method for deriving typicality judgments, applicable in young children that are not familiar with numerical values yet, is introduced, allowing researchers to study gradedness at younger ages in concept development. Contrary to the long tradition of using rating-based procedures to derive typicality judgments, we propose a method that is based on typicality ranking rather than rating, in which items are gradually sorted according to their typicality, and that requires a minimum of linguistic knowledge. The validity of the method is investigated and the method is compared to the traditional typicality rating measurement in a large empirical study with eight different semantic concepts. The results show that the typicality ranking task can be used to assess children’s category knowledge and to evaluate how this knowledge evolves over time. Contrary to earlier held assumptions in studies on typicality in young children, our results also show that preference is not so much )
In the largest, longitudinal study of young, deaf children before and three years after cochlear implantation, we compared symbolic play and novel noun learning to age-matched hearing peers. Participants were 180 children from six cochlear implant centers and 96 hearing children. Symbolic play was measured during five minutes of videotaped, structured solitary play. Play was coded as "symbolic" if the child used substitution (e.g., a wooden block as a bed). Novel noun learning was measured in 10 trials using a novel object and a distractor. Cochlear implant vs. normal hearing children were delayed in their use of symbolic play, however, those implanted before vs. after age two performed significantly better. Children with cochlear implants were also delayed in novel noun learning (median delay 1.54 years), with minimal evidence of catch-up growth. Quality of parent-child interactions was positively related to performance on the novel noun learning, but not symbolic play task. Early im)
Using a large social media dataset and open-vocabulary methods from computational linguistics, we explored differences in language use across gender, affiliation, and assertiveness. In Study 1, we analyzed topics (groups of semantically similar words) across 10 million messages from over 52,000 Facebook users. Most language differed little across gender. However, topics most associated with self-identified female participants included friends, family, and social life, whereas topics most associated with self-identified male participants included swearing, anger, discussion of objects instead of people, and the use of argumentative language. In Study 2, we plotted male- and female-linked language topics along two interpersonal dimensions prevalent in gender research: affiliation and assertiveness. In a sample of over 15,000 Facebook users, we found substantial gender differences in the use of affiliative language and slight differences in assertive language. Language used more by self-)
The conveyor system plays a vital role in improving the performance of flexible manufacturing cells (FMCs). The conveyor selection problem involves the evaluation of a set of potential alternatives based on qualitative and quantitative criteria. This paper presents an integrated multi-criteria decision making (MCDM) model of a fuzzy AHP (analytic hierarchy process) and fuzzy ARAS (additive ratio assessment) for conveyor evaluation and selection. In this model, linguistic terms represented as triangular fuzzy numbers are used to quantify experts’ uncertain assessments of alternatives with respect to the criteria. The fuzzy set is then integrated into the AHP to determine the weights of the criteria. Finally, a fuzzy ARAS is used to calculate the weights of the alternatives. To demonstrate the effectiveness of the proposed model, a case study is performed of a practical example, and the results obtained demonstrate practical potential for the implementation of FMCs. [ABSTRACT FROM AUTHO)
Life satisfaction refers to a somewhat stable cognitive assessment of one’s own life. Life satisfaction is an important component of subjective well being, the scientific term for happiness. The other component is affect: the balance between the presence of positive and negative emotions in daily life. While affect has been studied using social media datasets (particularly from Twitter), life satisfaction has received little to no attention. Here, we examine trends in posts about life satisfaction from a two-year sample of Twitter data. We apply a surveillance methodology to extract expressions of both satisfaction and dissatisfaction with life. A noteworthy result is that consistent with their definitions trends in life satisfaction posts are immune to external events (political, seasonal etc.) unlike affect trends reported by previous researchers. Comparing users we find differences between satisfied and dissatisfied users in several linguistic, psychosocial and other features. For )
Children’s interpretations of sentences containing focus particles do not seem adult-like until school age. This study investigates how German 4-year-old children comprehend sentences with the focus particle ‘nur’ (only) by using different tasks and controlling for the impact of general cognitive abilities on performance measures. Two sentence types with ‘only’ in either pre-subject or pre-object position were presented. Eye gaze data and verbal responses were collected via the visual world paradigm combined with a sentence-picture verification task. While the eye tracking data revealed an adult-like pattern of focus particle processing, the sentence-picture verification replicated previous findings of poor comprehension, especially for ‘only’ in pre-subject position. A second study focused on the impact of general cognitive abilities on the outcomes of the verification task. Working memory was related to children’s performance in both sentence types whereas inhibitory control was sel)
Objective: This study examines reading aloud in patients with amyotrophic lateral sclerosis (ALS) and those with frontotemporal dementia (FTD) in order to determine whether differences in patterns of speaking and pausing exist between patients with primary motor vs. primary cognitive-linguistic deficits, and in contrast to healthy controls. Design: 136 participants were included in the study: 33 controls, 85 patients with ALS, and 18 patients with either the behavioural variant of FTD (FTD-BV) or progressive nonfluent aphasia (FTD-PNFA). Participants with ALS were further divided into 4 non-overlapping subgroups—mild, respiratory, bulbar (with oral-motor deficit) and bulbar-respiratory—based on the presence and severity of motor bulbar or respiratory signs. All participants read a passage aloud. Custom-made software was used to perform speech and pause analyses, and this provided measures of speaking and articulatory rates, duration of speech, and number and duration of pauses. Thes)
Quantitative analysis of organismal form is an important component for almost every branch of biology. Although generally considered an easily-measurable structure, the quantification of gastropod shell form is still a challenge because many shells lack homologous structures and have a spiral form that is difficult to capture with linear measurements. In view of this, we adopt the idea of theoretical modelling of shell form, in which the shell form is the product of aperture ontogeny profiles in terms of aperture growth trajectory that is quantified as curvature and torsion, and of aperture form that is represented by size and shape. We develop a workflow for the analysis of shell forms based on the aperture ontogeny profile, starting from the procedure of data preparation (retopologising the shell model), via data acquisition (calculation of aperture growth trajectory, aperture form and ontogeny axis), and data presentation (qualitative comparison between shell forms) and ending with)
The doublecortin domain-containing 2 (DCDC2) gene, which is located on chromosome 6p22.1, has been widely suggested to be a candidate gene for dyslexia, but its role in typical reading development over time remains to be clarified. In the present study, we explored the role of DCDC2 in contributing to the individual differences in reading development from ages 6 to 11 years by analysing data from 284 unrelated children who were participating in the Chinese Longitudinal Study of Reading Development (CLSRD). The associations of eight single nucleotide polymorphisms (SNPs) in DCDC2 with the latent intercept and slope of children’s reading scores were examined in the first step. There was significant support for an association of rs807724 with the intercept for the reading comprehension measure of reading fluency, and the minor “G” allele was associated with poor reading performance. Next, we further tested the rs807724 SNP in association with the reading ability at each tested time and r)
The Pirahã language has been at the center of recent debates in linguistics, in large part because it is claimed not to exhibit recursion, a purported universal of human language. Here, we present an analysis of a novel corpus of natural Pirahã speech that was originally collected by Dan Everett and Steve Sheldon. We make the corpus freely available for further research. In the corpus, Pirahã sentences have been shallowly parsed and given morpheme-aligned English translations. We use the corpus to investigate the formal complexity of Pirahã syntax by searching for evidence of syntactic embedding. In particular, we search for sentences which could be analyzed as containing center-embedding, sentential complements, adverbials, complementizers, embedded possessors, conjunction or disjunction. We do not find unambiguous evidence for recursive embedding of sentences or noun phrases in the corpus. We find that the corpus is plausibly consistent with an analysis of Pirahã as a regular langua)
Objectives: Previous studies have demonstrated that microRNA-132 plays a vital part in and is actively associated with several cancers, with its tumor-suppressive role in hepatocellular carcinoma confirmed. The current study employed multiple bioinformatics techniques to establish gene signatures for hepatocellular carcinoma, microRNA-132 predicted target genes and the corresponding overlaps. Methods: Various assays were performed to explore the role and cellular functions of miR-132 in HCC and a successive panel of tasks was completed, including NLP analysis, miR-132 target genes prediction, comprehensive analyses (gene ontology analysis, pathway analysis, network analysis and connectivity analysis), and analytical integration. Later, HCC-related and miR-132-related potential targets, pathways, networks and highlighted hub genes were revealed as well as those of the overlapped section. Results: MiR-132 was effective in both impeding cell growth and boosting apoptosis in HCC cell l)
The morphology and distribution of lateral line neuromasts vary between ecomorphological types of anuran tadpoles, but little is known about how this structural variability contributes to differences in lateral-line mediated behaviors. Previous research identified distinct differences in one such behavior, positive rheotaxis towards the source of a flow, in two tadpole species, the African clawed frog (Xenopus laevis; type 1) and the American bullfrog (Rana catesbeiana; type 4). Because these two species had been tested under different flow conditions, we re-evaluated these findings by quantifying flow-sensing behaviors of bullfrog tadpoles in the same flow field in which X. laevis tadpoles had been tested previously. Early larval bullfrog tadpoles were exposed to flow in the dark, in the presence of a discrete light cue, and after treatment with the ototoxin gentamicin. In response to flow, tadpoles moved downstream, closer to a side wall, and higher in the water column, but they did)
Running a concurrent task while speaking clearly interferes with speech planning, but whether verbal vs. non-verbal tasks interfere with the same processes is virtually unknown. We investigated the neural dynamics of dual-task interference on word production using event-related potentials (ERPs) with either tones or syllables as concurrent stimuli. Participants produced words from pictures in three conditions: without distractors, while passively listening to distractors and during a distractor detection task. Production latencies increased for tasks with higher attentional demand and were longer for syllables relative to tones. ERP analyses revealed common modulations by dual-task for verbal and non-verbal stimuli around 240 ms, likely corresponding to lexical selection. Modulations starting around 350 ms prior to vocal onset were only observed when verbal stimuli were involved. These later modulations, likely reflecting interference with phonological-phonetic encoding, were observed)
A considerable body of sensory research has addressed the rules governing simultaneity judgments (SJs) and temporal order judgments (TOJs). In principle, neural events that register stimulus-arrival-time differences at an early sensory stage could set the limit on SJs and TOJs alike. Alternatively, distinct limits on SJs and TOJs could arise from task-specific neural events occurring after the stimulus-driven stage. To distinguish between these possibilities, we developed a novel reaction-time (RT) measure and tested it in a perceptual-learning procedure. The stimuli comprised dual-stream Rapid Serial Visual Presentation (RSVP) displays. Participants judged either the simultaneity or temporal order of red-letter and black-number targets presented in opposite lateral hemifield streams of black-letter distractors. Despite identical visual stimulation across-tasks, the SJ and TOJ tasks generated distinct RT patterns. SJs exhibited significantly faster RTs to synchronized targets than to )
A fundamental problem in linguistics is how literary texts can be quantified mathematically. It is well known that the frequency of a (rare) word in a text is roughly inverse proportional to its rank (Zipf’s law). Here we address the complementary question, if also the rhythm of the text, characterized by the arrangement of the rare words in the text, can be quantified mathematically in a similar basic way. To this end, we consider representative classic single-authored texts from England/Ireland, France, Germany, China, and Japan. In each text, we classify each word by its rank. We focus on the rare words with ranks above some threshold Q and study the lengths of the (return) intervals between them. We find that for all texts considered, the probability SQ(r) that the length of an interval exceeds r, follows a perfect Weibull-function, SQ(r) = exp(−b(β)rβ), with β around 0.7. The return intervals themselves are arranged in a long-range correlated self-similar fashion, where the autoc)
Despite Saussure’s famous observation that sound-meaning relationships are in principle arbitrary, we now have a substantial body of evidence that sounds themselves can have meanings, patterns often referred to as “sound symbolism”. Previous studies have found that particular sounds can be associated with particular meanings, and also with particular static visual shapes. Less well studied is the association between sounds and dynamic movements. Using a free elicitation method, the current experiment shows that several sound symbolic associations between sounds and dynamic movements exist: (1) front vowels are more likely to be associated with small movements than with large movements; (2) front vowels are more likely to be associated with angular movements than with round movements; (3) obstruents are more likely to be associated with angular movements than with round movements; (4) voiced obstruents are more likely to be associated with large movements than with small movements. All)
In this study we propose a novel, unsupervised clustering methodology for analyzing large datasets. This new, efficient methodology converts the general clustering problem into the community detection problem in graph by using the Jensen-Shannon distance, a dissimilarity measure originating in Information Theory. Moreover, we use graph theoretic concepts for the generation and analysis of proximity graphs. Our methodology is based on a newly proposed memetic algorithm (iMA-Net) for discovering clusters of data elements by maximizing the modularity function in proximity graphs of literary works. To test the effectiveness of this general methodology, we apply it to a text corpus dataset, which contains frequencies of approximately 55,114 unique words across all 168 written in the Shakespearean era (16th and 17th centuries), to analyze and detect clusters of similar plays. Experimental results and comparison with state-of-the-art clustering methods demonstrate the remarkable performance )
The negative symptoms of schizophrenia (SZ) are associated with a pattern of reinforcement learning (RL) deficits likely related to degraded representations of reward values. However, the RL tasks used to date have required active responses to both reward and punishing stimuli. Pavlovian biases have been shown to affect performance on these tasks through invigoration of action to reward and inhibition of action to punishment, and may be partially responsible for the effects found in patients. Forty-five patients with schizophrenia and 30 demographically-matched controls completed a four-stimulus reinforcement learning task that crossed action (“Go” or “NoGo”) and the valence of the optimal outcome (reward or punishment-avoidance), such that all combinations of action and outcome valence were tested. Behaviour was modelled using a six-parameter RL model and EEG was simultaneously recorded. Patients demonstrated a reduction in Pavlovian performance bias that was evident in a reduced Go )
Prior research found reliable and considerably strong effects of semantic achievement primes on subsequent performance. In order to simulate a more natural priming condition to better understand the practical relevance of semantic achievement priming effects, running texts of schoolbook excerpts with and without achievement primes were used as priming stimuli. Additionally, we manipulated the achievement context; some subjects received no feedback about their achievement and others received feedback according to a social or individual reference norm. As expected, we found a reliable (albeit small) positive behavioral priming effect of semantic achievement primes on achievement in math (Experiment 1) and language tasks (Experiment 2). Feedback moderated the behavioral priming effect less consistently than we expected. The implication that achievement primes in schoolbooks can foster performance is discussed along with general theoretical implications. [ABSTRACT FROM AUTHOR], Copyright )
Huntington’s disease (HD) is genetically determined but with variability in symptom onset, leading to uncertainty as to when pharmacological intervention should be initiated. Here we take a computational approach based on neurocognitive phenotyping, computational modeling, and classification, in an effort to provide quantitative predictors of HD before symptom onset. A large sample of subjects—consisting of both pre-manifest individuals carrying the HD mutation (pre-HD), and early symptomatic—as well as healthy controls performed the antisaccade conflict task, which requires executive control and response inhibition. While symptomatic HD subjects differed substantially from controls in behavioral measures [reaction time (RT) and error rates], there was no such clear behavioral differences in pre-HD. RT distributions and error rates were fit with an accumulator-based model which summarizes the computational processes involved and which are related to identified mechanisms in more detai)
How does linguistic structure affect children’s acquisition of early number word meanings? Previous studies have tested this question by comparing how children learning languages with different grammatical representations of number learn the meanings of labels for small numbers, like 1, 2, and 3. For example, children who acquire a language with singular-plural marking, like English, are faster to learn the word for 1 than children learning a language that lacks the singular-plural distinction, perhaps because the word for 1 is always used in singular contexts, highlighting its meaning. These studies are problematic, however, because reported differences in number word learning may be due to unmeasured cross-cultural differences rather than specific linguistic differences. To address this problem, we investigated number word learning in four groups of children from a single culture who spoke different dialects of the same language that differed chiefly with respect to how they grammat)
We integrate recent findings from the linguistics literature with the organizational justice literature to examine how the language used to encode justice violations influences fairness perceptions. The study focused on the use of non-agentive syntax to encode mistakes in Spanish ('The vase was broken') versus using agentive syntax in English ('She broke the vase') influences event fairness perceptions. We hypothesized that when justice violations are encoded using Spanish, because the non-agentive syntax makes the responsible party less salient, the event would be perceived as less unfair. In Study 1 (n = 111), English-speaking participants rated the fairness of an event in which a mistake was made and an employee received a negative outcome. They rated it as more unfair (p < .01, η² = .06) when the scenario was presented in agentive syntax. Experiment 2 (n = 70) used native English- and Spanish-speakers who watched a video of manager making a mistake. We found that Spanish-speakers used less agentive syntax (p < .01, η² = .21), perceived the event as less unfair (p < .001, η² = .23), and were more willing to help the manager who made the mistake. In Experiment 3 (n = 101) we replicated this effect controlling for cross-cultural differences and native language; further, we found an interaction between entity fairness (event vs. entity) and native language (Spanish vs. English) on citizenship intentions (p < .01, η² = .08). These results extend our understanding of how language may influence relevant workplace attitudes. (PsycINFO Database Record (c) 2017 APA, all rights reserved)
This study examines how preadolescent African American students in Washington, D.C., used a linguistic practice called ‘joning,’ a style of verbal play similar to ritual insults, in peer interactions. Sociolinguists have focused on how children socialize each other into vernacular styles appropriate for peer group use but often assume that they disalign with social and linguistic norms for classroom behavior. Drawing from a nine-month ethnographic study that the author conducted in an after-school program, this article analyzes the structure and function of joning as a vernacular style of African American Vernacular English and its uses in constructing classroom identities. Joning often facilitated student learning, but it was perceived as a socially and physically risky linguistic practice because of its uses as conflict talk in the local community. Focusing on preadolescence as a key stage of language socialization, this article shows how minority students modify peer-learned linguistic practices to pursue academic success on their own terms. (PsycINFO Database Record (c) 2017 APA, all rights reserved)
This theoretical paper discusses different linguistic theories that have dealt (or in some cases: not dealt) with how situated utterances are built in natural language: What is the role of abstract systems of linguistic norms or impersonal brain mechanisms? Can individual speakers make decisions about their own utterances? In this paper some traditional, structuralist, interactionist and dialogist theories are mutually contrasted. Starting out from a dialogist framework, a notion of participatory agency will be developed, based on the fact that speakers' situated languaging occurs in various activity types in direct or indirect interaction with others. Recent theories of interbodily dynamics, or intercorporeality, are discussed. A version of extended dialogism is proposed. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This research explores the cultural and linguistic strategies of immigrant youth to negotiate inclusion/exclusion, including language discrimination in Vancouver, Canada. My theoretical framework draws upon the Arendtian notions of ‘public space’, and ‘action and speech’ as well as Bourdieu’s concepts of ‘symbolic violence’ and ‘habitus’. My methodology is a critical qualitative approach. Fourteen immigrant youth, aged 15–25, were involved in this research. The findings of this study indicate that unlike second-generation immigrants, first-generation immigrant youth face cultural and linguistic challenges. Non-recognition of youths’ distinct linguistic and social capitals, the imposition of official languages and the regulation of the education and language market according to the dominant linguistic norms include forms of discrimination against Turkish minority youth in Canada. Taken together, the findings suggest that immigrant youths’ cultural and linguistic experiences of inclusion and exclusion cannot be dissociated from the wider politics of the nation-state, popular hegemony and social inequalities in the host society. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
Inviting Citizen Designers to Design Learning Management System (LMS) Interfaces for Student Agency in a Digital Cross-Cultural Contact Zone assesses how FYC students from periphery cultural and linguistic backgrounds perceive Blackboard Learn and other learning management system (LMS) interfaces. The report of an empirical study shows that the current LMS design does not provide writing students in general and writing students from periphery cultural and linguistic backgrounds in particular an opportunity of a higher-level interactivity with the LMS. The current design neither includes periphery students' cultural and linguistic norms and values, nor does it allow them to affect the existing design through their design activities. These LMSs are currently constraining users from higher-level interactions. As a result, writing students have to act as the LMS ask them to do, and they remain passive in these platforms. Based on the web usability test responses, this study proposes to invite Citizen Designers, writing students from periphery cultural and linguistic backgrounds, to design LMS interfaces to enhance user activities and transform them into cross-cultural platforms. This study analyzes interface designs by Citizen Designers to see how designers acquire their agency in a cross-cultural digital contact zone. This study concludes that Citizen Designers' participation in interface design helps them create favorable electronic environments that help them acquire their agency and enhance their (digital) writings and researches. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
This study examines electrocortical activity associated with visual and auditory sensory perception and lexical-semantic processing in nonverbal (NV) or minimally-verbal (MV) children with Autism Spectrum Disorder (ASD). Currently, there is no agreement on whether these children comprehend incoming linguistic information and whether their perception is comparable to that of typically developing children. Event-related potentials (ERPs) of 10 NV/MV children with ASD and 10 neurotypical children were recorded during a picture-word matching paradigm. Atypical ERP responses were evident at all levels of processing in children with ASD. Basic perceptual processing was delayed in both visual and auditory domains but overall was similar in amplitude to typically-developing children. However, significant differences between groups were found at the lexical-semantic level, suggesting more atypical higher-order processes. The results suggest that although basic perception is relatively preserve)
Coreference resolution is one of the fundamental and challenging tasks in natural language processing. Resolving coreference successfully can have a significant positive effect on downstream natural language processing tasks, such as information extraction and question answering. The importance of coreference resolution for biomedical text analysis applications has increasingly been acknowledged. One of the difficulties in coreference resolution stems from the fact that distinct types of coreference (e.g., anaphora, appositive) are expressed with a variety of lexical and syntactic means (e.g., personal pronouns, definite noun phrases), and that resolution of each combination often requires a different approach. In the biomedical domain, it is common for coreference annotation and resolution efforts to focus on specific subcategories of coreference deemed important for the downstream task. In the current work, we aim to address some of these concerns regarding coreference resolution in)
Purpose: The purpose of the present study was to extend previous research by analyzing the ability of adults who stutter to use phonological working memory in conjunction with lexical access to perform a word jumble task. Method: Forty English words consisting of 3-, 4-, 5-, and 6-letters (n = 10 per letter length category) were randomly jumbled using a web-based application. During the experimental task, 26 participants were asked to silently manipulate the scrambled letters to form a real word. Each vocal response was coded for accuracy and speech reaction time (SRT). Results: Adults who stutter attempted to solve fewer word jumble stimuli than adults who do not stutter at the 4-letter, 5-letter, and 6-letter lengths. Additionally, adults who stutter were significantly less accurate solving word jumble tasks at the 4-letter, 5-letter, and 6-letter lengths compared to adults who do not stutter. At the longest word length (6-letter), SRT was significantly slower for the adults who )
In the masked priming technique, physical identity between prime and target enjoys an advantage over nominal identity in nonwords (GEDA-GEDA faster than geda-GEDA). However, nominal identity overrides physical identity in words (e.g., REAL-REAL similar to real-REAL). Here we tested whether the lack of an advantage of the physical identity condition for words was due to top-down feedback from phonological-lexical information. We examined this issue with deaf readers, as their phonological representations are not as fully developed as in hearing readers. Results revealed that physical identity enjoyed a processing advantage over nominal identity not only in nonwords but also in words (GEDA-GEDA faster than geda-GEDA; REAL-REAL faster than real-REAL). This suggests the existence of fundamental differences in the early stages of visual word recognition of hearing and deaf readers, possibly related to the amount of feedback from higher levels of information. [ABSTRACT FROM AUTHOR], Copyrig)
Language is not only the representation of thinking, but also shapes thinking. Studies on bilinguals suggest that a foreign language plays an important and unconscious role in thinking. In this study, a software—Linguistic Inquiry and Word Count 2007—was used to investigate whether the learning of English as a foreign language (EFL) can foster Chinese high school students’ English analytic thinking (EAT) through the analysis of their English writings with our self-built corpus. It was found that: (1) learning English can foster Chinese learners’ EAT. Chinese EFL learners’ ability of making distinctions, degree of cognitive complexity and degree of thinking activeness have all improved along with the increase of their English proficiency and their age; (2) there exist differences in Chinese EFL learners’ EAT and that of English native speakers, i. e. English native speakers are better in the ability of making distinctions and degree of thinking activeness. These findings suggest that t)
The rate of lexical replacement estimates the diachronic stability of word forms on the basis of how frequently a proto-language word is replaced or retained in its daughter languages. Lexical replacement rate has been shown to be highly related to word class and word frequency. In this paper, we argue that content words and function words behave differently with respect to lexical replacement rate, and we show that semantic factors predict the lexical replacement rate of content words. For the 167 content items in the Swadesh list, data was gathered on the features of lexical replacement rate, word class, frequency, age of acquisition, synonyms, arousal, imageability and average mutual information, either from published databases or gathered from corpora and lexica. A linear regression model shows that, in addition to frequency, synonyms, senses and imageability are significantly related to the lexical replacement rate of content words–in particular the number of synonyms that a word)
As visual media spread to all domains of public and scientific life, nonverbal behavior is taking its place as an important form of communication alongside the written and spoken word. An objective and reliable method of analysis for hand movement behavior and gesture is therefore currently required in various scientific disciplines, including psychology, medicine, linguistics, anthropology, sociology, and computer science. However, no adequate common methodological standards have been developed thus far. Many behavioral gesture-coding systems lack objectivity and reliability, and automated methods that register specific movement parameters often fail to show validity with regard to psychological and social functions. To address these deficits, we have combined two methods, an elaborated behavioral coding system and an annotation tool for video and audio data. The NEUROGES–ELAN system is an effective and user-friendly research tool for the analysis of hand movement behavior, including gesture, self-touch, shifts, and actions. Since its first publication in 2009 in Behavior Research Methods, the tool has been used in interdisciplinary research projects to analyze a total of 467 individuals from different cultures, including subjects with mental disease and brain damage. Partly on the basis of new insights from these studies, the system has been revised methodologically and conceptually. The article presents the revised version of the system, including a detailed study of reliability. The improved reproducibility of the revised version makes NEUROGES–ELAN a suitable system for basic empirical research into the relation between hand movement behavior and gesture and cognitive, emotional, and interactive processes and for the development of automated movement behavior recognition methods.
We investigated the linguistic patterns in the discourse of four generations of the collective leadership of the Communist Party of China (CPC) from 1921 to 2012. The texts of Mao Zedong, Deng Xiaoping, Jiang Zemin, and Hu Jintao were analyzed using computational linguistic techniques (a Chinese formality score) to explore the persuasive linguistic features of the leaders in the contexts of power phase, the nation’s education level, power duration, and age. The study was guided by the elaboration likelihood model of persuasion, which includes a central route (represented by formal discourse) versus a peripheral route (represented by informal discourse) to persuasion. The results revealed that these leaders adopted the formal, central route more when they were in power than before they came into power. The nation’s education level was a significant factor in the leaders’ adoption of the persuasion strategy. The leaders’ formality also decreased with their increasing age and in-power times. However, the predictability of these factors for formality had subtle differences among the different types of leaders. These results enhance our understanding of the Chinese collective leadership and the role of formality in politically persuasive messages.
Semantic textual similarity is a measure of the degree of semantic equivalence between two pieces of text. We describe the SemSim system and its performance in the *SEM 2013 and SemEval-2014 tasks on semantic textual similarity. At the core of our system lies a robust distributional word similarity component that combines latent semantic analysis and machine learning augmented with data from several linguistic resources. We used a simple term alignment algorithm to handle longer pieces of text. Additional wrappers and resources were used to handle task specific challenges that include processing Spanish text, comparing text sequences of different lengths, handling informal words and phrases, and matching words with sense definitions. In the *SEM 2013 task on Semantic Textual Similarity, our best performing system ranked first among the 89 submitted runs. In the SemEval-2014 task on Multilingual Semantic Textual Similarity, we ranked a close second in both the English and Spanish subtasks. In the SemEval-2014 task on Cross-Level Semantic Similarity, we ranked first in Sentence–Phrase, Phrase–Word, and Word–Sense subtasks and second in the Paragraph–Sentence subtask.
Semantic similarities are a cross-field research in Natural Language Processing and Ontologies with some possible fallout in Artificial Intelligence. Formerly, similarities were computed following a syntactical treatment to support case-based reasoning. Textual similarities are now guided by semantic machineries, offering various ways to compute relatedness measures. In this paper, we present both a logical and a visual framework aiming to reason with them. For that reason, we introduced FLH±, a fragment of description logic underpinning the well-known lexical database Wordnet. We illustrated this framework with the path length relatedness, one of the historical similarity measures occurring in a taxonomy. The core of our framework orchestrates the computation of similarity scores supported by REVERB, STANFORD CORENLP and WORDNET:SIMILARITY APIs and interfaces global similarities in graphical way by positioning them on segments. We also depicted some experimental results to confront our computational framework with some empirical data.
My research focuses on the study of grammatical change in the recent history of the English language; in particular, I am currently writing my PhD dissertation on from Early Modern English to Present-Day English. In this PhD project, I analyse and compare the different factors that appear to influence in Present-Day English with earlier stages of the language (Late Modern English).The concept of ellipsis refers to a syntactic strategy in which expected elements have been left unpronounced in certain constructions. This omission triggers a mismatch between meaning (the intended message) and sound (what is in fact uttered). In particular, my research focuses on those examples of (Miller 2011, Miller and Pullum 2013), i.e. ellipsis types that occur after the following licensors (that is, those elements that license ellipsis): modal verbs, auxiliaries be, have and do, infinitival marker to and negator not. The main aim is to carry out an empirical analysis of from Late Modern English to Present-Day English (1700-1914), both quantitatively and qualitatively, by means of data retrieved from the Penn Corpora of Historical English. This project pays attention to syntactic variation, genre distribution and discourse variables (type of anaphora, mismatches in polarity, aspect, voice, modality, tense; comparison of clause types; distance, linking, type of focus).Esta tesis doctoral versa sobre la variacion diacronica de la elipsis sintactica desde ingles moderno temprano hasta la actualidad. La elipsis representa un desajuste entre el significado de lo que se dice (la intencion de un mensaje) y lo que se pronuncia en realidad. En el ambito de la linguistica moderna, la elipsis se estudia en el campo la semantica, la sintaxis, la pragmatica, la psicolinguistica, la linguistica de corpus, etc. y constituye la novedad de esta investigacion el tratamiento empirico de este fenomeno linguistico tratando de juntar variables procedentes de las distintas teorias consultadas. Cabe destacar que no existe un estudio multidisciplinar sobre este concepto sintactico en ingles moderno. Existen solo unos pocos estudios relativamente recientes pero se centran unicamente en el estudio del ingles actual. Por esta razon, se decidio llevar a cabo la investigacion tanto de los aspectos formales como de los funcionales de la elipsis y su evolucion diacronica teniendo en cuenta distintas variables discursivas. Para eso, se creo una base de datos en la que apareceria el ejemplo de Post-Auxiliary Ellipsis (elipsis despues de un auxiliar) con su numero identificador; el genero al que pertenece (de los dieciocho diferentes que existen en el corpus utilizado, el Penn Treebank); el licensor (aquel elemento que posibilita la elipsis); el tipo de union entre el antecedente de la elipse y la clausula eliptica (coordinacion, subordinacion, parataxis, etc.); el tipo de conector entre el antecedente y la parte la distancia existente entre el antecedente y la clausula eliptica (numero de clausulas); contexto sintactico en el que aparece la elipsis (clausulas principales, subordinadas, question-tags, etc); tipo de anafora (anaforico, cataforico, exoforico); categoria del antecedente (sintagma verbal, nominal, adjetival o no constituyente); categoria del material elidido (sintagma verbal, nominal, adjetival o no constituyente); presencia o ausencia de cambio de referente; comparacion del aspecto, la voz, la modalidad y el tiempo del antecedente con respecto a la parte presencia o ausencia de question-tag; comparacion entre el tipo de clausula del antecedente (declarativa, interrogativa o imperativa) y el de la clausula elidida; y tipo de foco de la parte elidida (auxiliary-choice (eleccion de auxiliar), subject-choice (eleccion de sujeto) o ambos).Esta tese de doutoramento versa sobre a variacion diacronica da elipse sintactica dende ingles moderno temperan ata a actualidade. A elipse representa un desaxuste entre o significado do que se di (a intencion dunha mensaxe) e o que se pronuncia en realidade. No ambito da linguistica moderna, a elipse estudase no campo a semantica, a sintaxe, a pragmatica, a psicolinguistica, a linguistica de corpus, etc. e constitue a novidade desta investigacion o tratamento empirico deste fenomeno linguistico tratando de xuntar variables procedentes das distintas teorias consultadas. Cabe destacar que non existe un estudo multidisciplinar sobre este concepto sintactico en ingles moderno. Existen so uns poucos estudos relativamente recentes pero centranse unicamente no estudo do ingles actual. Por esta razon, decidiuse levar a cabo a investigacion tanto dos aspectos formais coma dos funcionais da elipse e a sua evolucion diacronica tendo en conta distintas variables discursivas. Para iso, creouse unha base de datos na que apareceria o exemplo de Post-Auxiliary Ellipsis (elipse despois dun auxiliar) co seu numero identificador; o xenero ao que pertence (dos dezaoito diferentes que existen no corpus utilizado, o Penn Treebank); o licensor (aquel elemento que posibilita a elipse); o tipo de union entre o antecedente da elipse e a clausula eliptica (coordinacion, subordinacion, parataxe, etc.); o tipo de conector entre o antecedente e a parte a distancia existente entre o antecedente e a clausula eliptica (numero de clausulas); contexto sintactico no que aparece a elipse (clausulas principais, subordinadas, question-tags, etc); tipo de anafora (anaforico, cataforico, exoforico); categoria do antecedente (sintagma verbal, nominal, adxectival ou non constituinte); categoria do material elidido (sintagma verbal, nominal, adxectival ou non constituinte); presenza ou ausencia de cambio de referente; comparacion do aspecto, a voz, a modalidade e o tempo do antecedente con respecto a parte presenza ou ausencia de question-tag; comparacion entre o tipo de clausula do antecedente (declarativa, interrogativa ou imperativa) e o da clausula elidida; e tipo de foco da parte elidida (auxiliary-choice (eleccion de auxiliar), subject-choice (eleccion de suxeito) ou ambos).
The rapid accumulation of data in social media (in million and billion scales) has imposed great challenges in information extraction, knowledge discovery, and data mining, and texts bearing sentiment and opinions are one of the major categories of user generated data in social media. Sentiment analysis is the main technology to quickly capture what people think from these text data, and is a research direction with immediate practical value in ‘big data’ era. Learning such techniques will allow data miners to perform advanced mining tasks considering real sentiment and opinions expressed by users in additional to the statistics calculated from the physical actions (such as viewing or purchasing records) user perform, which facilitates the development of real-world applications. However, the situation that most tools are limited to the English language might stop academic or industrial people from doing research or products which cover a wider scope of data, retrieving information from people who speak different languages, or developing applications for worldwide users. More specifically, sentiment analysis determines the polarities and strength of the sentiment-bearing expressions, and it has been an important and attractive research area. In the past decade, resources and tools have been developed for sentiment analysis in order to provide subsequent vital applications, such as product reviews, reputation management, call center robots, automatic public survey, etc. However, most of these resources are for the English language. Being the key to the understanding of business and government issues, sentiment analysis resources and tools are required for other major languages, e.g., Chinese. In this tutorial, audience can learn the skills for retrieving sentiment from texts in another major language, Chinese, to overcome this obstacle. The goal of this tutorial is to introduce the proposed sentiment analysis technologies and datasets in the literature, and give the audience the opportunities to use resources and tools to process Chinese texts from the very basic preprocessing, i.e., word segmentation and part of speech tagging, to sentiment analysis, i.e., applying sentiment dictionaries and obtaining sentiment scores, through step-by-step instructions and a hand-on practice. The basic processing tools are from CKIP Participants can download these resources, use them and solve the problems they encounter in this tutorial. This tutorial will begin from some background knowledge of sentiment analysis, such as how sentiment are categorized, where to find available corpora and which models are commonly applied, especially for the Chinese language. Then a set of basic Chinese text processing tools for word segmentation, tagging and parsing will be introduced for the preparation of mining sentiment and opinions. After bringing the idea of how to pre-process the Chinese language to the audience, I will describe our work on compositional Chinese sentiment analysis from words to sentences, and an application on social media text (Facebook) as an example. All our involved and recently developed related resources, including Chinese Morphological Dataset, Augmented NTU Sentiment Dictionary (aug-NTUSD), E-hownet with sentiment information, Chinese Opinion Treebank, and the CopeOpi Sentiment Scorer, will also be introduced and distributed in this tutorial. The tutorial will end by a hands-on session of how to use these materials and tools to process Chinese sentiment. Content Details, Materials, and Program please refer to the tutorial URL: http://www.lunweiku.com/
Representation of syntactic structure is a core area of research in Computational Linguistics, disambiguating distinctions in meaning that are crucial for correct interpretation of language. Development of algorithms and statistical models over the past three decades has led to systems that are accurate enough to be deployed in industry, playing a key role in products such as Google Search and Apple Siri. However, syntactic parsers today are usually constrained to tree representations of language, and performance is interpreted through a single metric that conveys no linguistic information regarding remaining errors.In this dissertation, we present new algorithms for error analysis and parsing. The heart of our approach to error analysis is the use of structural transformations to identify more meaningful classes of errors, and to enable comparisons across formalisms. For parsing, we combine a novel dynamic program with careful choices in syntactic representation to create an efficient parser that produces graph structured output. Together, these developments allowed us to evaluate the outstanding challenges in parsing and to address a key weakness in current work.First, we present a search algorithm that, given two structures, finds a sequence of modifications leading from one structure to the other. We applied this algorithm to syntactic error analysis, where one structure is the output of a parser, the other is the correct parse, and each modification corresponds to fixing one error. We constructed a tool based on the algorithm and analyzed variations in behavior between parsers, types of text, and languages. Our observations shine light on several assumptions about syntactic errors, showing some to be true and others to be false. For example, prepositional phrase attachment errors are indeed a major issue, while coordination scope errors do not hurt performance as much as expected.Next, we describe an algorithm that builds a parse in one syntactic representation to match a parse in another representation. Specifically, we build phrase structure parses from Combinatory Categorial Grammar derivations. Our approach follows the philosophy of CCG, defining specific phrase structures for each lexical category and generic rules for combinatory steps. The new parse is built by following the CCG derivation bottom-up, gradually building the corresponding phrase structure parse. This produced significantly more accurate parses than past work, and enabled us to compare performance of several parsers across formalisms.Finally, we address a weakness we observed in phrase structure parsers: the exclusion of syntactic trace structures for computational convenience. We present an efficient dynamic programming algorithm that constructs the graph structure that has the highest score under an edge-factored scoring function. We define a parse representation compatible with the algorithm, and show how certain linguistic distinctions dramatically impact coverage. We also show various ways to modify the algorithm to improve performance by exploiting properties of observed linguistic structure. This approach to syntactic parsing is the first to cover virtually all structure encoded in the Penn Treebank.
We address in this paper some theoretical and practical issues relating to generation, processing, and management of Parallel Translation Corpus (PTC) in Indian languages, which is under development in a consortium-mode project (ILCI-II) 1 under the aegis of DeitY, Govt. of India. These issues are discussed here for the first time keeping in mind the ready application of PTC in various domains of linguistics including computational linguistics, Natural Language Processing, applied linguistics, lexicography, translation, language description, etc. In a normative manner, we define what is a PTC; describe the process of its construction; identify its features; exemplify the processes of text alignment in PTC; discuss the methods of text analysis; propose for restructuring of translational units; define the process of extraction of translational equivalents; propose for generation of bilingual lexical database and Term Bank from a structured PTC; and finally identify the areas where a PTC and information extracted from it may be utilized. Since construction of PTC in Indian languages is full of hurdles, we try to construct a roadmap with a focus on techniques and methodologies that may be applied for achieving the task. The issues are brought under focus to justify the present work that is trying to construct PTC for some Indian languages for future reference and application.
Individuals differ in their ability to feel their own and others' internal states, with those that have more autistic and less empathic traits clustering at the clinical end of the spectrum. However, when we consider semantic competence, this group could compensate with a higher capacity to imagine the meaning of words referring to emotions. This is indeed what we found when we asked people with different levels of autistic and empathic traits to rate the degree of imageability of various kinds of words. But this was not the whole story. Individuals with marked autistic traits demonstrated outstanding ability to imagine theoretical concepts, i.e., concepts that are commonly grasped linguistically through their definitions. This distinctive characteristic was so pronounced that, using tree-based predictive models, it was possible to accurately predict participants' inclination to manifest autistic traits, as well as their adherence to autistic profiles - including whether they fell above or below the diagnostic threshold - from their imageability ratings. We speculate that this quasi-perceptual ability to imagine theoretical concepts represents a specific cognitive pattern that, while hindering social interaction, may favor problem solving in abstract, non-socially related tasks. This would allow people with marked autistic traits to make use of perceptual, possibly visuo-spatial, information for "higher" cognitive processing.
This paper address the problem of extracting learning objects or keywords from bunch of Documents,with the goal of use this objects for various purpose like Resume filtering,Email filtering,content classification etc. Keyword extraction, concept finding are in learning objects is very important subject in today’s eLearning environment. Keywords are subset of words that contains the useful information about the content of the document. Keyword extraction is a process that is used to get the important keywords from documents. In this proposed System I Calculate the TF-IDF of each word, then Decision tree algorithm is used for feature selection process using wordnet dictionary. WordNet is a lexical database of English which is used to find similarity from the candidate words. The words having highest similarity are taken as keywords.
The primary goal in this thesis is to identify better syntactic constraint or bias, that is language independent but also efficiently exploitable during sentence processing. We focus on a particular syntactic construction called center-embedding, which is well studied in psycholinguistics and noted to cause particular difficulty for comprehension. Since people use language as a tool for communication, one expects such complex constructions to be avoided for communication efficiency. From a computational perspective, center-embedding is closely relevant to a left-corner parsing algorithm, which can capture the degree of center-embedding of a parse tree being constructed. This connection suggests left-corner methods can be a tool to exploit the universal syntactic constraint that people avoid generating center-embedded structures. We explore such utilities of center-embedding as well as left-corner methods extensively through several theoretical and empirical examinations. Our primary task is unsupervised grammar induction. In this task, the input to the algorithm is a collection of sentences, from which the model tries to extract the salient patterns on them as a grammar. This is a particularly hard problem although we expect the universal constraint may help in improving the performance since it can effectively restrict the possible search space for the model. We build the model by extending the left-corner parsing algorithm for efficiently tabulating the search space except those involving center-embedding up to a specific degree. We examine the effectiveness of our approach on many treebanks, and demonstrate that often our constraint leads to better parsing performance. We thus conclude that left-corner methods are particularly useful for syntax-oriented systems, as it can exploit efficiently the inherent universal constraints in languages.
Chinese word segmentation and Part-of-speech (POS) tagging have been studied for decades. However, most of the previous works mainly focus on pipeline method which will lead to error propagation. In order to make word segmentation and POS tagging jointly in one model, in this paper, we propose an effective neural network model to improve the accuracy of the segmentation and tagging. Our model works based on the hierarchical Long Short-Term Memory (LSTM) and trained jointly in one objective function. What's more, to better utilizing the transition features between tags, we further introduce the transition matrix which can help to search the best tagging sequence. Experiment on Chinese Treebank shows that our model achieves competitive accuracy on word segmentation and POS tagging.
Treebanks are curial for natural language processing (NLP). In this paper, we present our work for annotating a Chinese treebank in scientific domain (SCTB), to address the problem of the lack of Chinese treebanks in this domain. Chinese analysis and machine translation experiments conducted using this treebank indicate that the annotated treebank can significantly improve the performance on both tasks. This treebank is released to promote Chinese NLP research in scientific domain.
One of the widely used approaches to Sentiment Analysis (SA) is lexicon-based approach that depends on sentiment-annotated lexical resources (such as SentiWordNet (SWN)). A broad variety of such resources are Synsetbased Lexical Databases (SLDs) (e.g. SWN is based on WordNet (WN)) and represent sentiment degrees of synonym groups of LDs, called "synsets." However, synsets themselves were open to criticism because although, in reality, not all the members of a synset represent its meaning with the same degree, in SLDs, they are, identically, considered as members of their synset. Therefore, the fuzzy version of synsets was proposed in a small number of previous studies. Fuzzy synsets can upgrade such lexicon-based SA by which the future SA systems can discriminate between word-senses of a same synset, how much each of them contains the sentiment load of that synset. But, to the best of our knowledge, none of the studies on fuzzy synsets has proposed any algorithm for providing fuzzy versions of "predefined synsets" of an SLD. In this study, we present the idea of an algorithm for constructing fuzzy version of any SLD of any language, given a corpus of that language and a word-sense-disambiguation system of that language/SLD.