Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
One of the widely used approaches to Sentiment Analysis (SA) is lexicon-based approach that depends on sentiment-annotated lexical resources (such as SentiWordNet (SWN)). A broad variety of such resources are Synsetbased Lexical Databases (SLDs) (e.g. SWN is based on WordNet (WN)) and represent sentiment degrees of synonym groups of LDs, called "synsets." However, synsets themselves were open to criticism because although, in reality, not all the members of a synset represent its meaning with the same degree, in SLDs, they are, identically, considered as members of their synset. Therefore, the fuzzy version of synsets was proposed in a small number of previous studies. Fuzzy synsets can upgrade such lexicon-based SA by which the future SA systems can discriminate between word-senses of a same synset, how much each of them contains the sentiment load of that synset. But, to the best of our knowledge, none of the studies on fuzzy synsets has proposed any algorithm for providing fuzzy versions of "predefined synsets" of an SLD. In this study, we present the idea of an algorithm for constructing fuzzy version of any SLD of any language, given a corpus of that language and a word-sense-disambiguation system of that language/SLD.
We describe the Corpus of Spoken Icelandic (ÍS-TAL) which is made up of 15 hours of spontaneous naturally occurring conversa-tions, 31 conversations in all. The corpus comprises 184,080 tokens, 14,297 types and 9,221 lemmas. It has been transcribed using standard orthography. We present a list of the 30 most common lemmas in the corpus and compare it to a list of the most frequent lemmas in the written language, concluding that the differences between the two lists are smaller than expected. We have tagged the corpus morphologically with a statistical tagger that had been trained on written texts. The results are much better than we expected, and the tagging accuracy is as least as high as for the written texts. The final part of the paper is a report on a work in progress. We have been experimenting with converting the morphological tagging into a shallow syntactic markup by applying a few simple hand-written rules. Even though the analysis we get by using this procedure is bound to be incomplete and contain several errors, we conclude that the results are promising and we can use this method to build a simple yet useful treebank with minimal effort. 1.
Research finds we make spontaneous trait inferences from facial appearance, even after brief exposures to a face (i.e., less than or equal to 100 ms). We examined spontaneous impressions of criminality from facial appearance, testing whether these impressions persist after repeated presentation (i.e., one to three exposures) and increased exposure duration (100, 500, or 1,000 ms) to the face. Judgement confidence and response times were recorded. Other participants viewed the faces for an unlimited period of time, rating trustworthiness, dominance and criminal appearance. We found evidence that participants spontaneously make criminal appearance attributions. These inferences persisted with repeated presentation and increased exposure duration, were related to trustworthiness and dominance ratings, and were made with high confidence. Implications are discussed.
Background Conventionally, it is believed that high-frequency auditory information is important for speech understanding. This is only partly true, as recent studies have demonstrated the importance of low-frequency information. This research was taken up to develop, standardize, and validate auditory low-frequency word lists in Hindi, an Indian language. Material and Methods The first phase of the study involved collection of bisyllabic words followed by verification by a native linguist. Words were then short-listed based on familiarity ratings given by 10 adult native speakers; those words were recorded and the best recorded words selected through subjective and objective analysis. Then, using Fast Fourier Transform and k-means clustering, words with more energy below 1.5 kHz were isolated. Finally, equally difficult 10 word lists were generated by obtaining psychometric function curves. Finally, lists were administered on 40 adult normal hearing particip Results Results showed a similar trend of increase in speech identification scores with increase in SL across all lists except list 4. During the final phase, developed lists were validated on 10 simulated low-frequency cochlear hearing loss participants. Hearing loss was simulated using Matlab and National Institute for Occupational Safety and Health (NIOSH) software. Results of validation revealed that auditory low-frequency word lists were sensitive enough to tap the speech understanding difficulty in the simulated condition. Conclusions The developed word lists can be used clinically to assess communication ability in individuals with rising hearing loss. The word lists also have the potential to assess the performance after amplification provided to individuals with rising hearing loss.
We address in this paper some theoretical and practical issues relating to generation, processing, and management of Parallel Translation Corpus (PTC) in Indian languages, which is under development in a consortium-mode project (ILCI-II) 1 under the aegis of DeitY, Govt. of India. These issues are discussed here for the first time keeping in mind the ready application of PTC in various domains of linguistics including computational linguistics, Natural Language Processing, applied linguistics, lexicography, translation, language description, etc. In a normative manner, we define what is a PTC; describe the process of its construction; identify its features; exemplify the processes of text alignment in PTC; discuss the methods of text analysis; propose for restructuring of translational units; define the process of extraction of translational equivalents; propose for generation of bilingual lexical database and Term Bank from a structured PTC; and finally identify the areas where a PTC and information extracted from it may be utilized. Since construction of PTC in Indian languages is full of hurdles, we try to construct a roadmap with a focus on techniques and methodologies that may be applied for achieving the task. The issues are brought under focus to justify the present work that is trying to construct PTC for some Indian languages for future reference and application.
This paper presents the conversion of Syn-TagRus dependency structures into Penn Treebank style phrase structures, whose resulting data will be used to train a statistical constituency parser for Russian and create a large-scale constituency-parsed corpus. The implemented conversion includes various innovative features in order to create phrase structure trees that are closest to Penn Treebank style while optimally preserving information of the original dependency structure annotations. We believe the newly converted phrase structure treebank will be not only an adequate training dataset for our ongoing project but also a valuable resource for traditional and computational linguistic research.
The authors try to transform dependency tree into phrase structure tree, and detect annotation errors automatically based on manual rules. The method is used in processing Peking University Multi-view Chinese Treebank(PMT). Although PMT has been manually checked twice before processed by this method, 1529 errors are detected among the 50275 sentences and the precision is 100%. The errors mainly belong to three types: word segmentation error, mismatching between POS and syntactic role, and syntactic role error. This method can further improve treebank quality, and be applied to other dependency treebanks.
Recent neural network sequence models with softmax classifiers have achieved their best language modeling performance only with very large hidden states and large vocabularies. Even then they struggle to predict rare or unseen words even if the context makes the prediction unambiguous. We introduce the pointer sentinel mixture architecture for neural sequence models which has the ability to either reproduce a word from the recent context or produce a word from a standard softmax classifier. Our pointer sentinel-LSTM model achieves state of the art language modeling performance on the Penn Treebank (70.9 perplexity) while using far fewer parameters than a standard softmax LSTM. In order to evaluate how well language models can exploit longer contexts and deal with more realistic vocabularies and larger corpora we also introduce the freely available WikiText corpus.
The internship is a critical part of graduate training and often the only opportunity to receive on-site clinical supervision during school psychology practice. Nonetheless, the process of pairing interns with field supervisors is not standardized and sometimes relies on factors such as logistics and supervisor credentials rather than a consideration of interpersonal variables that could optimize the internship experience. Related fields have found mixed evidence for a relationship between personality similarity within a supervisory dyad and outcomes such as a strong supervisory relationship, satisfaction with supervision, and supervisee effectiveness. This study examined the influence of personality similarity on ratings of supervisory working alliance, supervision satisfaction, and intern work readiness. This study also evaluated the predictive power of personality, supervisory working alliance, and systemic factors on intern work readiness and supervision satisfaction. Lastly, this study assessed the development of the supervisory working alliance and intern work readiness over time. Twenty-six dyads were recruited for participation in this study, including 24 practicing school psychologists serving as field supervisors and 26 school psychology interns. Data collection occurred at the midpoint and end of the internship year. Participants completed a demographic questionnaire, personality inventory, and measures of supervisory working alliance, supervision satisfaction, supervisee work readiness, and systemic factors. Results indicated that personality similarity among supervisors and interns is not related to supervisory working alliance, supervision satisfaction, or supervisee work readiness. However, supervisor ratings of supervisory working alliance were predictive of intern work readiness, and intern ratings of supervisory working alliance were predictive of supervision satisfaction. Systemic factors were not predictive of intern work readiness or supervision satisfaction. For supervisors, the supervisory working alliance significantly decreased over time, while intern ratings remained consistent from midyear to the end of the year. Intern development from midyear to the end of year could not be determined due to low scale reliability. Future studies should further examine factors that contribute to the supervisory working alliance and validate measures specific to the school context. More research is needed to establish the conditions and interpersonal characteristics that enable an optimal internship experience for both supervisors and supervisees in school psychology.
The pol-nkjp1m-pargram-dev structure bank was created using POLFIE: an LFG grammar of Polish. This structure bank contains sentences from the NKJP1M subcorpus of NKJP which were not included in Skladnica treebank. The pol-nkjp1m-pargram-dev structure bank can be accessed via INESS treebanking system in two ways: • use the direct link: http://clarino.uib.no/iness/lfg-sentences?&treebank=pol-nkjp1m-pargram-dev • go to http://iness.uib.no --> choose Treebank Selection in the menu on the left-hand side --> choose POLFIE in Treebank Collections --> choose pol-nkjp1m-pargram-dev
This paper deals with the declension of foreign words and loanwords in Croatian newspapers in the period from 1950 to 1999. The usage of foreign nouns, of a-type declension, with vowel final sounds in nominative in newspaper texts is compared with Croatian literary-linguistic norm. This comparison uncovers which forms are either inconsistent in terms of their usage or not in accordance with the linguistic norm. It can be seen that the grammatical norms of the second half of the 20th century were insufficient for news writing style incessantly inundated with new expressions - particularly foreign names which should be adjusted to Croatian declension system. That also points out the need for more detailed description in the present morphological norm.
Semantic analysis of sentences can only be carried out using Dependency Parsing. Dependency parser accepts words in a sentence and builds dependency relation among the words resulting in a unique tree for each sentence. An Indian Panini is the first to develop semantic analysis for Sanskrit using a dependency framework. Western researchers in the near past have also deliberated on dependency parsing so that automated dependency parser can be generated. Dependency Parser is useful in information extraction, question-answering, text summarization etc. For many indian languages namely Bengali, Kannada, Malayalam and Marathi a dependency based Treebank is in the development stage. Also for Hindi, Telugu and Tamil dependency Treebanks are already developed. The Treebank data can be used by a dependency parser generator like Maltparser to develop a Dependency parser. In this paper few dependency parsing algorithms are discussed
This article invetigates some of the most relevant issues concerning the relation between language and gender, understood as the social elaboration of the properties associated to the sex. It is in this perspective that we can interpret the notion of 'linguistic sexism', that is the whjole of the linguistic devices that the literature deals with as indicators of the different roles and the different power associated to the sexual differences. The analysis of the relation between language and gender naturally implies a more general reflection on the relation between language and thought and on the link between language and spciety, and between linguistic norm and linguistic use.
Krokodyl is an experimental hybrid deep depencency parser of Polish. Krokodyl has been developed at the Institute of Computer Science, Polish Academy of Sciences (IPI PAN) within the CLARIN-PL project. It was create to evaluate a hybrid approach to parsing: combining syntactic, lexical and semantic features for dependency parsing. It uses a number of tools as components of the feature generation chain, namely the Spejd Grammar, MALT parsing engine, MATE, the Polish Wordnet, the Skladnica treebank.
In NLP data drives research, as evidenced by the frequency with which seminal works of database engineering such as The Penn Treebank have been employed as a basis for experimentation. Traditionally large-scale expertly annotated corpora are expensive and time consuming to produce. This paradigm drove researchers to adopt automated methods for generating labelled data with available tools such as Freebase, DBpedia, and the "infoboxes" found on Wikipedia pages. These knowledge bases have been, or are in the process of being, subsumed by Wikidata, an initiative to concentrate such disparate data repositories in an organized machine readable format. This resource is an important research tool. In this paper, we review our experience using Wikidata in constructing a large annotated corpus under distant supervision, moreover we make the materials, the code used to generate our annotations, freely available to all interested parties.
Ubiquity Press is an open access publisher of peer-reviewed academic journals, books and data. We operate a highly cost-efficient model that makes quality open access publishing affordable for everyone. Ubiquity has joined forces with De Gruyter. Read the announcement.
This article proposes an ontology design pattern for leading knowledge providers to represent knowledge in more normalized, precise and interrelated ways, hence in ways that help the matching and exploitation of knowledge from different sources. This pattern is a knowledge sharing best practice that is domain and language independent. It can be used as a criteria for measuring the quality of an ontology. This pattern is: using binary relation types directly derived from concept types, especially role types or types of process. The article explains and illustrates this pattern, and relates it to other patterns and general ontology quality criteria. It also provides an ontology for automatically deriving relation types from concept types (e.g., those from lexical ontologies such as those derived from the WordNet lexical database). This derivation helps normalizing knowledge, reduces having to introduce new relation types and helps keeping all the types organized.
Term sense disambiguation is very essential for different approaches of NLP, including Internet search engines, information retrieval, Data mining, classification etc. However, the old methods using case frames and semantic primitives are not qualify for solving term ambiguities which needs a lot of information with sentences. This new approach introduces a building structure system of natural language knowledge. In this paper all surface case patterns is classified in advance with the consideration of the meaning of noun. Moreover, this paper introduces an efficient data structure using a trie which define the linkage among leaves and multi-attribute relations. By using this linkage multi-attribute relations, we can get a high frequent access among verbs and noun with an automatic generation of hierarchical relationships. In our experiment a large tagged corpus (Pan Treebank) is used to extract data. In our approach around 11,000 verbs and nouns is used for verifying the new method and made a hierarchy group of its noun. Moreover, the achievement of term disambiguating using our trie structure method and linking trie among leaves is 6% higher than old method.
Natural sleep provides a powerful model system for studying the neuronal correlates of awareness and state changes in the human brain. To quantitatively map the nature of sleep-induced modulations in sensory responses we presented participants with auditory stimuli possessing different levels of linguistic complexity. Ten participants were scanned using functional magnetic resonance imaging (fMRI) during the waking state and after falling asleep. Sleep staging was based on heart rate measures validated independently on 20 participants using concurrent EEG and heart rate measurements and the results were confirmed using permutation analysis. Participants were exposed to three types of auditory stimuli: scrambled sounds, meaningless word sentences and comprehensible sentences. During non-rapid eye movement (NREM) sleep, we found diminishing brain activation along the hierarchy of language processing, more pronounced in higher processing regions. Specifically, the auditory thalamus showe)
The present study examines the effect of language experience on vocal emotion perception in a second language. Native speakers of French with varying levels of self-reported English ability were asked to identify emotions from vocal expressions produced by American actors in a forced-choice task, and to rate their pleasantness, power, alertness and intensity on continuous scales. Stimuli included emotionally expressive English speech (emotional prosody) and non-linguistic vocalizations (affect bursts), and a baseline condition with Swiss-French pseudo-speech. Results revealed effects of English ability on the recognition of emotions in English speech but not in non-linguistic vocalizations. Specifically, higher English ability was associated with less accurate identification of positive emotions, but not with the interpretation of negative emotions. Moreover, higher English ability was associated with lower ratings of pleasantness and power, again only for emotional prosody. This sugg)
The claim that Eskimo languages have words for different types of snow is well-known among the public, but has been greatly exaggerated through popularization and is therefore viewed with skepticism by many scholars of language. Despite the prominence of this claim, to our knowledge the line of reasoning behind it has not been tested broadly across languages. Here, we note that this reasoning is a special case of the more general view that language is shaped by the need for efficient communication, and we empirically test a variant of it against multiple sources of data, including library reference works, Twitter, and large digital collections of linguistic and meteorological data. Consistent with the hypothesis of efficient communication, we find that languages that use the same linguistic form for snow and ice tend to be spoken in warmer climates, and that this association appears to be mediated by lower communicative need to talk about snow and ice. Our results confirm that variati)
We study rank-frequency relations for phonemes, the minimal units that still relate to linguistic meaning. We show that these relations can be described by the Dirichlet distribution, a direct analogue of the ideal-gas model in statistical mechanics. This description allows us to demonstrate that the rank-frequency relations for phonemes of a text do depend on its author. The author-dependency effect is not caused by the author’s vocabulary (common words used in different texts), and is confirmed by several alternative means. This suggests that it can be directly related to phonemes. These features contrast to rank-frequency relations for words, which are both author and text independent and are governed by the Zipf’s law. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, dow)
Music, like languages, is one of the key components of our culture, yet musical evolution is still poorly known. Numerous studies using computational methods derived from evolutionary biology have been successfully applied to varied subset of linguistic data. One of the major drawback regarding musical studies is the lack of suitable coded musical data that can be analysed using such evolutionary tools. Here we present for the first time an original set of musical data coded in a way that enables construction of trees classically used in evolutionary approaches. Using phylogenetic methods, we test two competing theories on musical evolution: vertical versus horizontal transmission. We show that, contrary to what is currently believed, vertical transmission plays a key role in shaping musical diversity. The signal of vertical transmission is particularly strong for intrinsic musical characters such as metrics, rhythm, and melody. Our findings reveal some of the evolutionary mechanisms )
Whether using two languages enhances executive functions is a matter of debate. Here, we take a novel perspective to examine the bilingual advantage hypothesis by comparing bi-dialect with mono-dialect speakers’ performance on a non-linguistic task that requires executive control. Two groups of native Chinese speakers, one speaking only the standard Chinese Mandarin and the other also speaking the Southern-Min dialect, which differs from the standard Chinese Mandarin primarily in phonology, performed a classic Flanker task. Behavioural results showed no difference between the two groups, but event-related potentials recorded simultaneously revealed a number of differences, including an earlier P2 effect in the bi-dialect as compared to the mono-dialect group, suggesting that the two groups engage different underlying neural processes. Despite differences in the early ERP component, no between-group differences in the magnitude of the Flanker effects, which is an index of conflict reso)
Compositional “language of thought” models have recently been proposed to account for a wide range of children’s conceptual and linguistic learning. The present work aims to evaluate one of the most basic assumptions of these models: children should have an ability to represent and compose functions. We show that 3.5–4.5 year olds are able to predictively compose two novel functions at significantly above chance levels, even without any explicit training or feedback on the composition itself. We take this as evidence that children at this age possess some capacity for compositionality, consistent with models that make this ability explicit, and providing an empirical challenge to those that do not. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles )
Research on cross-linguistic comparisons of the neural correlates of reading has consistently found that the left middle frontal gyrus (MFG) is more involved in Chinese than in English. However, there is a lack of consensus on the interpretation of the language difference. Because this region has been found to be involved in writing, we hypothesize that reading Chinese characters involves this writing region to a greater degree because Chinese speakers learn to read by repeatedly writing the characters. To test this hypothesis, we recruited English L1 learners of Chinese, who performed a reading task and a writing task in each language. The English L1 sample had learned some Chinese characters through character-writing and others through phonological learning, allowing a test of writing-on-reading effect. We found that the left MFG was more activated in Chinese than English regardless of task, and more activated in writing than in reading regardless of language. Furthermore, we found )
Recognizing speech in adverse listening conditions is a significant cognitive, perceptual, and linguistic challenge, especially for children. Prior studies have yielded mixed results on the impact of bilingualism on speech perception in noise. Methodological variations across studies make it difficult to converge on a conclusion regarding the effect of bilingualism on speech-in-noise performance. Moreover, there is a dearth of speech-in-noise evidence for bilingual children who learn two languages simultaneously. The aim of the present study was to examine the extent to which various adverse listening conditions modulate differences in speech-in-noise performance between monolingual and simultaneous bilingual children. To that end, sentence recognition was assessed in twenty-four school-aged children (12 monolinguals; 12 simultaneous bilinguals, age of English acquisition ≤ 3 yrs.). We implemented a comprehensive speech-in-noise battery to examine recognition of English sentences acro)
Treebank is one of important resources in the natural language processing. Compared with the rich and mature Chinese corpus, Vietnamese Syntactic Analysis is much more difficult. This paper presents a new approach which uses Chinese-Vietnamese bilingual word alignment corpus to build Vietnamese Dependency Treebank. Firstly, the aligned word processing was made by Chinese-Vietnamese sentence alignment; Secondly, the dependency parsing was done with Chinese sentences. Finally, Vietnamese Dependency Parsing Treebank was generated by Chinese-Vietnamese Languages align relationship and Chinese Dependency Tree, At the same time, The Vietnamese phrase tree converted into dependency Treebank can significantly improve the accuracy of dependency analysis. Experimental results show that this approach can simplify the process of manual collection and annotation of Vietnamese Treebank, and it can save manpower and time to build the Vietnamese Treebank. Experimental results show that the accuracy of this method compared to machine learning methods has improved significantly.
Background: Due to their increased vulnerability, immigrants are considered a priority group for communicable disease prevention and control in Europe. This study aims to compare influenza vaccination coverage (IVC) between regular immigrants and Italian citizens at risk for its complications and evaluate factors affecting differences. Methods: Based on data collected by the National Institute of Statistics during a population-based cross-sectional survey conducted in Italy in 2012–2013, we analysed information on 42,048 adult residents (≥ 18 years) at risk for influenza-related complications and with free access to vaccination (elderly residents ≥ 65 years and residents with specific chronic diseases). We compared IVC between 885 regular immigrants and 41,163 Italian citizens using log-binomial models and stratifying immigrants by area of origin and length of stay in Italy (recent: < 10 years; long-term: ≥ 10 years). Results: IVC among all immigrants was 16.9% compared to 40.2% am)
Language is not only the representation of thinking, but also shapes thinking. Studies on bilinguals suggest that a foreign language plays an important and unconscious role in thinking. In this study, a software—Linguistic Inquiry and Word Count 2007—was used to investigate whether the learning of English as a foreign language (EFL) can foster Chinese high school students’ English analytic thinking (EAT) through the analysis of their English writings with our self-built corpus. It was found that: (1) learning English can foster Chinese learners’ EAT. Chinese EFL learners’ ability of making distinctions, degree of cognitive complexity and degree of thinking activeness have all improved along with the increase of their English proficiency and their age; (2) there exist differences in Chinese EFL learners’ EAT and that of English native speakers, i. e. English native speakers are better in the ability of making distinctions and degree of thinking activeness. These findings suggest that t)
The ancestry of the Colombian population comprises a large number of well differentiated Native communities belonging to diverse linguistic groups. In the late fifteenth century, a process of admixture was initiated with the arrival of the Europeans, and several years later, Africans also became part of the Colombian population. Therefore, the genepool of the current Colombian population results from the admixture of Native Americans, Europeans and Africans. This admixture occurred differently in each region of the country, producing a clearly stratified population. Considering the importance of population substructure in both clinical and forensic genetics, we sought to investigate and compare patterns of genetic ancestry in Colombia by studying samples from Native and non-Native populations living in its 5 continental regions: the Andes, Caribe, Amazonia, Orinoquía, and Pacific regions. For this purpose, 46 AIM-Indels were genotyped in 761 non-related individuals from current popula)
Online Voting Advice Applications (VAAs) are survey-like instruments that help citizens to shape their political preferences and compare them with those of political parties. Especially in multi-party democracies, their increasing popularity indicates that VAAs play an important role in opinion formation for citizens, as well as in the public debate prior to elections. Hence, the objectivity and transparency of VAAs are crucial. In the design of VAAs, many choices have to be made. Extant research in survey methodology shows that the seemingly arbitrary choice to word questions positively (e.g., ‘The city council should allow cars into the city centre’) or negatively (‘The city council should ban cars from the city centre’) systematically affects the answers. This asymmetry in answers is in line with work on negativity bias in other areas of linguistics and psychology. Building on these findings, this study investigated whether question polarity also affects the answers to VAA statemen)
A defining trait of linguistic competence is the ability to combine elements into increasingly complex structures to denote, and to comprehend, a potentially infinite number of meanings. Recent magnetoencephalography (MEG) work has investigated these processes by comparing the response to nouns in combinatorial (blue car) and non-combinatorial (rnsh car) contexts. In the current study we extended this paradigm using electroencephalography (EEG) to dissociate the role of semantic content from phonological well-formedness (yerl car). We used event-related potential (ERP) recordings in order to better relate the observed neurophysiological correlates of basic combinatorial operations to prior ERP work on comprehension. We found that nouns in combinatorial contexts (blue car) elicited a greater centro-parietal negativity between 180-400ms, independent of the phonological well-formedness of the context word. We discuss the potential relationship between this ‘combinatorial’ effect and clas)
Our current understanding of pre-Columbian history in the Americas rests in part on several trends identified in recent genetic studies. The goal of this study is to reexamine these trends in light of the impact of post-Columbian admixture and the methods used to study admixture. The previously-published data consist of 645 autosomal microsatellite genotypes from 1046 individuals in 63 populations. We used STRUCTURE to estimate ancestry proportions and tested the sensitivity of these estimates to the choice of the number of clusters, K. We used partial correlation analyses to examine the relationship between gene diversity and geographic distance from Beringia, controlling for non-Native American ancestry (from Africa, Europe and East Asia), and taking into account alternative paths of migration. Principal component analysis and multidimensional scaling were used to investigate the relationships between Andean and non-Andean populations and to explore gene-language correspondence. We )
The Kashmiri population is an ethno-linguistic group that resides in the Kashmir Valley in northern India. A longstanding hypothesis is that this population derives ancestry from Jewish and/or Greek sources. There is historical and archaeological evidence of ancient Greek presence in India and Kashmir. Further, some historical accounts suggest ancient Hebrew ancestry as well. To date, it has not been determined whether signatures of Greek or Jewish admixture can be detected in the Kashmiri population. Using genome-wide genotyping and admixture detection methods, we determined there are no significant or substantial signs of Greek or Jewish admixture in modern-day Kashmiris. The ancestry of Kashmiri Tibetans was also determined, which showed signs of admixture with populations from northern India and west Eurasia. These results contribute to our understanding of the existing population structure in northern India and its surrounding geographical areas. [ABSTRACT FROM AUTHOR], Copyright)
The increasing need of automated analyzing web texts especially the short texts on Social Network Services (SNS) brings new demands of computerized text analysis instruments. The psychometric properties are the basis of the extensive use of these instruments such as the Linguistic Inquiry and Word Count (LIWC). For this study, Sina Weibo statuses were analyzed via rater coding and Simplified Chinese version of LIWC (SCLIWC), in order to evaluate the validity of SCLIWC in detecting psychological expressions in Weibo statuses (n = 60) and in identifying the psychological meaning of a single Weibo status (n = 11). Significant correlations between human ratings and SCLIWC scores and the high sensitivities of capturing single statuses with certain expressions identified by raters, proved the validity of SCLIWC in detecting psychological expressions. The results also suggested that, the efficiency of SCLIWC in detecting psychological expressions of SNS short texts could be higher if using s)
An alternative method for deriving typicality judgments, applicable in young children that are not familiar with numerical values yet, is introduced, allowing researchers to study gradedness at younger ages in concept development. Contrary to the long tradition of using rating-based procedures to derive typicality judgments, we propose a method that is based on typicality ranking rather than rating, in which items are gradually sorted according to their typicality, and that requires a minimum of linguistic knowledge. The validity of the method is investigated and the method is compared to the traditional typicality rating measurement in a large empirical study with eight different semantic concepts. The results show that the typicality ranking task can be used to assess children’s category knowledge and to evaluate how this knowledge evolves over time. Contrary to earlier held assumptions in studies on typicality in young children, our results also show that preference is not so much )
In the largest, longitudinal study of young, deaf children before and three years after cochlear implantation, we compared symbolic play and novel noun learning to age-matched hearing peers. Participants were 180 children from six cochlear implant centers and 96 hearing children. Symbolic play was measured during five minutes of videotaped, structured solitary play. Play was coded as "symbolic" if the child used substitution (e.g., a wooden block as a bed). Novel noun learning was measured in 10 trials using a novel object and a distractor. Cochlear implant vs. normal hearing children were delayed in their use of symbolic play, however, those implanted before vs. after age two performed significantly better. Children with cochlear implants were also delayed in novel noun learning (median delay 1.54 years), with minimal evidence of catch-up growth. Quality of parent-child interactions was positively related to performance on the novel noun learning, but not symbolic play task. Early im)
Using a large social media dataset and open-vocabulary methods from computational linguistics, we explored differences in language use across gender, affiliation, and assertiveness. In Study 1, we analyzed topics (groups of semantically similar words) across 10 million messages from over 52,000 Facebook users. Most language differed little across gender. However, topics most associated with self-identified female participants included friends, family, and social life, whereas topics most associated with self-identified male participants included swearing, anger, discussion of objects instead of people, and the use of argumentative language. In Study 2, we plotted male- and female-linked language topics along two interpersonal dimensions prevalent in gender research: affiliation and assertiveness. In a sample of over 15,000 Facebook users, we found substantial gender differences in the use of affiliative language and slight differences in assertive language. Language used more by self-)
While it has long been known that the pupil reacts to cognitive load, pupil size has received little attention in cognitive research because of its long latency and the difficulty of separating effects of cognitive load from the light reflex or effects due to eye movements. A novel measure, the Index of Cognitive Activity (ICA), relates cognitive effort to the frequency of small rapid dilations of the pupil. We report here on a total of seven experiments which test whether the ICA reliably indexes linguistically induced cognitive load: three experiments in reading (a manipulation of grammatical gender match / mismatch, an experiment of semantic fit, and an experiment comparing locally ambiguous subject versus object relative clauses, all in German), three dual-task experiments with simultaneous driving and spoken language comprehension (using the same manipulations as in the single-task reading experiments), and a visual world experiment comparing the processing of causal versus conce)
In this study we propose a novel, unsupervised clustering methodology for analyzing large datasets. This new, efficient methodology converts the general clustering problem into the community detection problem in graph by using the Jensen-Shannon distance, a dissimilarity measure originating in Information Theory. Moreover, we use graph theoretic concepts for the generation and analysis of proximity graphs. Our methodology is based on a newly proposed memetic algorithm (iMA-Net) for discovering clusters of data elements by maximizing the modularity function in proximity graphs of literary works. To test the effectiveness of this general methodology, we apply it to a text corpus dataset, which contains frequencies of approximately 55,114 unique words across all 168 written in the Shakespearean era (16th and 17th centuries), to analyze and detect clusters of similar plays. Experimental results and comparison with state-of-the-art clustering methods demonstrate the remarkable performance )
The negative symptoms of schizophrenia (SZ) are associated with a pattern of reinforcement learning (RL) deficits likely related to degraded representations of reward values. However, the RL tasks used to date have required active responses to both reward and punishing stimuli. Pavlovian biases have been shown to affect performance on these tasks through invigoration of action to reward and inhibition of action to punishment, and may be partially responsible for the effects found in patients. Forty-five patients with schizophrenia and 30 demographically-matched controls completed a four-stimulus reinforcement learning task that crossed action (“Go” or “NoGo”) and the valence of the optimal outcome (reward or punishment-avoidance), such that all combinations of action and outcome valence were tested. Behaviour was modelled using a six-parameter RL model and EEG was simultaneously recorded. Patients demonstrated a reduction in Pavlovian performance bias that was evident in a reduced Go )
Prior research found reliable and considerably strong effects of semantic achievement primes on subsequent performance. In order to simulate a more natural priming condition to better understand the practical relevance of semantic achievement priming effects, running texts of schoolbook excerpts with and without achievement primes were used as priming stimuli. Additionally, we manipulated the achievement context; some subjects received no feedback about their achievement and others received feedback according to a social or individual reference norm. As expected, we found a reliable (albeit small) positive behavioral priming effect of semantic achievement primes on achievement in math (Experiment 1) and language tasks (Experiment 2). Feedback moderated the behavioral priming effect less consistently than we expected. The implication that achievement primes in schoolbooks can foster performance is discussed along with general theoretical implications. [ABSTRACT FROM AUTHOR], Copyright )