Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This work improves a novel Service Selection Method for the development of Service-Oriented Applications in the context of the Service-Oriented Computing (SOC) paradigm. We have defined a Semantic-Structural Scheme to assess Web Services on Interface Compatibility exploring the available information from WSDL documents. The structural information involves data types from return, parameters and exceptions. The semantic information concerns identifiers from parameters and operation names. The lexical database WordNet is used as a semantic basis. Two appraisal values were defined: compatibility gap and adaptability gap. The former is centered on functional aspects. The latter explains the adaptation effort to a successful integration. We validated those appraisals values through different experiments with a data-set of 465 real-life Web Services and measured the results using three metrics from the Information Retrieval field.
Discourse relations bind smaller linguistic units into coherent texts. However, automatically identifying discourse relations is difficult, because it requires understanding the semantics of the linked arguments. A more subtle challenge is that it is not enough to represent the meaning of each argument of a discourse relation, because the relation may depend on links between lower-level components, such as entity mentions. Our solution computes distributional meaning representations by composition up the syntactic parse tree. A key difference from previous work on compositional distributional semantics is that we also compute representations for entity mentions, using a novel downward compositional pass. Discourse relations are predicted from the distributional representations of the arguments, and also of their coreferent entity mentions. The resulting system obtains substantial improvements over the previous state-of-the-art in predicting implicit discourse relations in the Penn Discourse Treebank.
BACKGROUND: The standard clinical acquisition for left ventricular functional parameter analysis with cardiovascular magnetic resonance (CMR) uses a multi-breathhold multi-slice segmented balanced SSFP sequence. Performing multiple long breathholds in quick succession for ventricular coverage in the short-axis orientation can lead to fatigue and is challenging in patients with severe cardiac or respiratory disorders. This study combines the encoding efficiency of a six-fold undersampled 3D stack of spirals balanced SSFP sequence with 3D through-time spiral GRAPPA parallel imaging reconstruction. This 3D spiral method requires only one breathhold to collect the dynamic data. METHODS: Ten healthy volunteers were recruited for imaging at 3 T. The 3D spiral technique was compared against 2D imaging in terms of systolic left ventricular functional parameter values (Bland-Altman plots), total scan time (Welch's t-test) and qualitative image rating scores (Wilcoxon signed-rank test). RESULTS: Systolic left ventricular functional values were not significantly different (i.e. 3D-2D) between the methods. The 95% confidence interval for ejection fraction was -0.1 ± 1.6% (mean ± 1.96*SD). The total scan time for the 3D spiral technique was 48 s, which included one breathhold with an average duration of 14 s for the dynamic scan, plus 34 s to collect the calibration data under free-breathing conditions. The 2D method required an average of 5 min 40s for the same coverage of the left ventricle. The difference between 3D and 2D image rating scores was significantly different from zero (Wilcoxon signed-rank test, p < 0.05); however, the scores were at least 3 (i.e. average) or higher for 3D spiral imaging. CONCLUSION: The 3D through-time spiral GRAPPA method demonstrated equivalent systolic left ventricular functional parameter values, required significantly less total scan time and yielded acceptable image quality with respect to the 2D segmented multi-breathhold standard in this study. Moreover, the 3D spiral technique used just one breathhold for dynamic imaging, which is anticipated to reduce patient fatigue as part of the complete cardiac examination in future studies that include patients.
This study examines the methodology of global foreign accent ratings in studies on L2 speech production. In three experiments, we test how variation in raters, range within speech samples, as well as instructions and procedures affects ratings of accent in predominantly monolingual speakers of German, non-native speakers of German, as well as long-term emigrants from Germany, that is, L1 attriters. The findings show that rater differences do not result in systematic changes in rating patterns. In contrast, range effects and effects of familiarity with accented speech lead to shifts in absolute and relative ratings. Including more strongly foreign-accented samples leads to lower judgements for the entire group of L2 speakers compared to natives. Similarly, lower familiarity with foreign accent results in more variable and more strongly foreign-accented judgements. We discuss the implications for research on L2 pronunciation as well as for the interpretation of nativeness in L2 studies and language testing more generally.
This paper proposes a simple yet effective framework of soft cross-lingual syntax projection to transfer syntactic structures from source language to target language using monolingual treebanks and large-scale bilingual parallel text. Here, soft means that we only project reliable dependencies to compose high-quality target structures. The projected instances are then used as additional training data to improve the performance of supervised parsers. The major issues for this idea are 1) errors from the source-language parser and unsupervised word aligner; 2) intrinsic syntactic non-isomorphism between languages; 3) incomplete parse trees after projection. To handle the first two issues, we propose to use a probabilistic dependency parser trained on the target-language treebank, and prune out unlikely projected dependencies that have low marginal probabilities. To make use of the incomplete projected syntactic structures, we adopt a new learning technique based on ambiguous labelings. For a word that has no head words after projection, we enrich the projected structure with all other words as its candidate heads as long as the newly-added dependency does not cross any projected dependencies. In this way, the syntactic structure of a sentence becomes a parse forest (ambiguous labels) instead of a single parse tree. During training, the objective is to maximize the mixed likelihood of manually labeled instances and projected instances with ambiguous labelings. Experimental results on benchmark data show that our method significantly outperforms a strong baseline supervised parser and previous syntax projection methods. 1
This is the first attempt at characterizing reading difficulty in Hindi using naturally occurring sentences. We created the Potsdam-Allahabad Hindi Eyetracking Corpus by recording eye-movement data from 30 participants at the University of Allahabad, India. The target stimuli were 153 sentences selected from the beta version of the Hindi-Urdu treebank. We find that word- or low-level predictors (syllable length, unigram and bigram frequency) affect first-pass reading times, regression path duration, total reading time, and outgoing saccade length. An increase in syllable length results in longer fixations, and an increase in word unigram and bigram frequency leads to shorter fixations. Longer syllable length and higher frequency lead to longer outgoing saccades. We also find that two predictors of sentence comprehension difficulty, integration and storage cost, have an effect on reading difficulty. Integration cost (Gibson, 2000) was approximated by calculating the distance (in words) between a dependent and head; and storage cost (Gibson, 2000), which measures difficulty of maintaining predictions, was estimated by counting the number of predicted heads at each point in the sentence. We find that integration cost mainly affects outgoing saccade length, and storage cost affects total reading times and outgoing saccade length. Thus, word-level predictors have an effect in both early and late measures of reading time, while predictors of sentence comprehension difficulty tend to affect later measures. This is, to our knowledge, the first demonstration using eye-tracking that both integration and storage cost influence reading difficulty.
Methylphenidate mainly enhances dopamine neurotransmission whereas 3,4-methylenedioxymethamphetamine (MDMA, "ecstasy") mainly enhances serotonin neurotransmission. However, both drugs also induce a weaker increase of cerebral noradrenaline exerting sympathomimetic properties. Dopaminergic psychostimulants are reported to increase sexual drive, while serotonergic drugs typically impair sexual arousal and functions. Additionally, serotonin has also been shown to modulate cognitive perception of romantic relationships. Whether methylphenidate or MDMA alter sexual arousal or cognitive appraisal of intimate relationships is not known. Thus, we evaluated effects of methylphenidate (40 mg) and MDMA (75 mg) on subjective sexual arousal by viewing erotic pictures and on perception of romantic relationships of unknown couples in a double-blind, randomized, placebo-controlled, crossover study in 30 healthy adults. Methylphenidate, but not MDMA, increased ratings of sexual arousal for explicit sexual stimuli. The participants also sought to increase the presentation time of implicit sexual stimuli by button press after methylphenidate treatment compared with placebo. Plasma levels of testosterone, estrogen, and progesterone were not associated with sexual arousal ratings. Neither MDMA nor methylphenidate altered appraisal of romantic relationships of others. The findings indicate that pharmacological stimulation of dopaminergic but not of serotonergic neurotransmission enhances sexual drive. Whether sexual perception is altered in subjects misusing methylphenidate e.g., for cognitive enhancement or as treatment for attention deficit hyperactivity disorder is of high interest and warrants further investigation.
We describe a new dependency parser for English tweets, TWEEBOPARSER. The parser builds on several contributions: new syntactic annotations for a corpus of tweets (TWEEBANK), with conventions informed by the domain; adaptations to a statistical parsing algorithm; and a new approach to exploiting out-of-domain Penn Treebank data. Our experiments show that the parser achieves over 80% unlabeled attachment accuracy on our new, high-quality test set and measure the benefit of our contributions. Our dataset and parser can be found at http://www.ark.cs.cmu.edu/TweetNLP.
According to Tsinghua Chinese Treebank annotation methods, the authors extracted relation words and marked their categories. Then syntax, lexical and position features of automatic syntax tree with and without functional marker were extracted to recognize and classify relation words. Experiment results show that relative recognition accuracy is 95.7%, and relation words classification F1 is 77.2%.
In this paper, we analyze the impact of various dependency representations for various constructions on the general parsing accuracy and on the parsing accuracy of these constructions. We focus on the analysis of coordination constructions, complex predicates, and punctuation mark attachment. We use Latvian Treebank as a dataset, thus, providing insight for an inflective language with a rather free word order. Experiments with MaltParser, a transition-based parser, show clear difference in learnability of various representations for the considered constructions. Future work would include carrying out comparable experiments with a graph-based dependency parser like MSTParser.
This paper mainly introduced the research on constructing Mongolian Treebank based on phrase structure grammar. Having Considered related Mongolian Treebank work and Mongolian words characteristics, we developed a Mongolian syntactic tagset. The tagset includes two kinds of tags. One is syntactic constituent tag and the other is grammatical relation tag. On the basis of the tagset, we developed the Mongolian Treebank auxiliary processing system. Finally, we built a Treebank that contains 3645 sentences and did an experiment on this Treebank.
English. Network theory provides a suitable framework to model the structure of language as a complex system. Based on a network built from a Latin dependency treebank, this paper applies methods for network analysis to show the key role of the verb sum (to be) in the overall structure of the network. Italiano. La teoria dei grafi fornisce un valido supporto alla modellizzazione strutturale del sistema linguistico. Basandosi su un network costruito a partire da una treebank a dipendenze del latino, l’articolo applica diversi metodi di analisi dei grafi, mostrando l’importanza del ruolo rivestito dal verbo sum (essere) nella struttura complessiva del network.
Comunicació presentada al 9th International Conference on Language Resources and Evaluation (LREC'14), celebrat del 26 al 31 de maig de 2014 a Reykjavík, Islàndia.
We investigate whether parsers can be used for self-monitoring in surface realization in order to avoid egregious errors involving "vicious" ambiguities, namely those where the intended interpretation fails to be considerably more likely than alternative ones. Using parse accuracy in a simple reranking strategy for selfmonitoring, we find that with a stateof-the-art averaged perceptron realization ranking model, BLEU scores cannot be improved with any of the well-known Treebank parsers we tested, since these parsers too often make errors that human readers would be unlikely to make. However, by using an SVM ranker to combine the realizer's model score together with features from multiple parsers, including ones designed to make the ranker more robust to parsing mistakes, we show that significant increases in BLEU scores can be achieved. Moreover, via a targeted manual analysis, we demonstrate that the SVM reranker frequently manages to avoid vicious ambiguities, while its ranking errors tend to affect fluency much more often than adequacy.
Automatic text categorisation systems is a type of software that every day it is receiving more interest, due not only to its use in documentaries environments but also to its possible application to tag properly documents on the Web. Many options have been proposed to face this subject using statistical approaches, natural language processing tools, ontologies and lexical databases. Nevertheless, there have been no too many empirical evaluations comparing the influence of the different tools used to solve these problems, particularly in a multilingual environment. In this paper we propose a multi-language rule-based pipeline system for automatic document categorisation and we compare empirically the results of applying techniques that rely on statistics and supervised learning with the results of applying the same techniques but with the support of smarter tools based on language semantics and ontologies, using for this purpose several corpora of documents. GENIE is being applied to real environments, which shows the potential of the proposal.
One way of teaching grammar, namely morphology and syntax, is to visualize sentences as diagrams capturing relationships between words. Similarly, such relationships are captured in a more complex way in treebanks serving as key building stones in modern natural language processing. However, building them is very time consuming, thus we have been seeking for an alternative cheaper and faster way, like crowdsourcing. The purpose of our work is to explore possibility to get sentence diagrams produced by students and teachers. In our pilot study, the object language is Czech, where sentence diagrams are part of elementary school curriculum.
This article analyses the structure of Yoruba numerals and their derivation. Data are collected from the compilation of Yoruba numerals and observation of its use coupled with the researcher's intuitive knowledge of the language. The work dwells on the existing literature on numerals too. The author adopts a descriptive method in analysing the data. The work looks at the roles of affixes in realising odd numbers, multiples of 20, centenary, bicentenary, and so on in their order of increase. It is discovered that the direction of counting in Yoruba is largely progressive. Besides, the language adopts base 5, decimal (base 10) and vigesimal (base 20) systems of counting. It is equally discovered that the choice of either of the two variations is largely dependent on the articulatory parameter of the first vowel (V1) of the root word. It is noted that the Yoruba numeral system offers a suitable linguistic database for both the theoretical and empirical domains of linguistic study especially documentary linguistics. The current study has general pedagogic implications for the teaching and learning of Yoruba numerals.
In the context of processing Bengali words through a computer, there may arise several issues that are directly linked with surface structure of words. These issues may create problems in manual and computer-based counting of number of words in a corpus. They can also create problems in morphological processing of words. These issues come up because there is hardly any consistency in orthographic representation of words in written Bengali texts. The high irregularities in writing of inflected words, proper names, adjectival forms, adverbial forms, compound words, reduplicated words, onomatopoeic words, hyphenated words, etc. present a daunting task before an investigator in normalizing the surface forms of words for generating a lexical database as well as developing a word processing system for the works of language technology.
Prior research suggested the possibility of establishing systematic linkages between some intrinsic features of a presupposition and textual and pragmatic functions that it can carry out with greater probability. This study aims, firstly, to provide an organic view of semantics of presupposition triggers, thanks to a lexical database comprising 19,500 entries. Secondly, the database was used to investigate a corpus of chat conversations including about 200,000 tokens. The results show that triggers occur mainly as non-informative, maintaining an information already known by all participants of the communication; but, depending on their different features, some of them are systematically associated to a function of anaphora and textual cohesion; others to strengthen social conventions and stereotypes. The informative function, although in a minority proportion, is quantitatively significant only in correspondence to a single class of presupposition triggers.
Introduction Previous research has suggested that visual images are more easily generated, more vivid, and more memorable than other sensory modalities. This research examined whether or not imagery is experienced in similar ways by people with and without sight. Specifically, the imageability of visual, auditory, and tactile cue words was compared. The degree to which images were multimodal or unimodal was also examined. Methods Twelve participants who were totally blind from early infancy and 12 sighted participants generated images in response to 53 sensory and nonsensory words, rating imageability and the sensory modality, and describing images. From these 53 items, 4 subgroups of words that stimulated images that were predominantly visual, tactile, auditory, and low-imagery were created. Results T-tests comparing imageability ratings from blind and sighted participants found no differences for auditory and tactile words (both p >.1). Nevertheless, although participants without sight found auditory and tactile images equally imageable, sighted participants found images in response to tactile cue words harder to generate than visual cue words (mean difference: −0.51, p =.025). Participants with sight were also more likely to develop multisensory images than were participants without sight (both U ≥ 15.0, N 1 = 12, N 2 = 12, p ≤.008). Discussion For both the blind and sighted groups, auditory and tactile images were rich and varied, and similar language was used. Sighted participants were more likely to generate multimodal images, and this was particularly the case for tactile words. Nevertheless, cue words that resulted in multisensory images were not necessarily rated as more imageable. The discussion considers whether or not multimodal imagery represents a method of compensating for impoverished unimodal imagery. Implications for practitioners Imagery is important not only as a mnemonic in memory rehabilitation, but also in everyday uses for modes such as autobiographical memory. This research emphasizes the importance of not only auditory and tactile sensory imagery, but also spatial imagery for people without sight.
The feminist movement purports to improve conditions for women, and yet only a minority of women in modern societies self-identify as feminists. This is known as the feminist paradox. It has been suggested that feminists exhibit both physiological and psychological characteristics associated with heightened masculinization, which may predispose women for heightened competitiveness, sex-atypical behaviors, and belief in the interchangeability of sex roles. If feminist activists, i.e., those that manufacture the public image of feminism, are indeed masculinized relative to women in general, this might explain why the views and preferences of these two groups are at variance with each other. We measured the 2D:4D digit ratios (collected from both hands) and a personality trait known as dominance (measured with the Directiveness scale) in a sample of women attending a feminist conference. The sample exhibited significantly more masculine 2D:4D and higher dominance ratings than comparison samples representative of women in general, and these variables were furthermore positively correlated for both hands. The feminist paradox might thus to some extent be explained by biological differences between women in general and the activist women who formulate the feminist agenda.
The corpus for training a parser consists of sentences of heterogeneous grammar usages. Previous parser domain adaptation work has concentrated on adaptation to the shifts in vocabulary rather than grammar usage. In this paper, we focus on exploiting the diversity of training date separately and then accumulates their advantages. We propose an approach that grammar is biased toward relevant syntactic style, and the complementary grammar usage are combined for inference. Multiple grammars with partly complementary points of strength are induced individually. They capture complementary data representation, and we accumulates their advantages in a joint model to assemble the complementary depicting powers. Despite its compatibility with many other methods, out product model achieves 85.20% F <sub xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">1</sub> score on Penn Chinese Treebank, higher than previous systems.
Reviewed by: Récits du corps au Maroc et au Japon ed. by Marc Kober and Khalid Zekri Gaëlle Corvaisier Kober, Marc, et Khalid Zekri, coords. Récits du corps au Maroc et au Japon. Paris: L’Harmattan, 2011. isbn 9782296557208. 200 p. Avec Récits du corps au Maroc et au Japon, Marc Kober et Khalid Zekri posent des questions essentielles pour la littérature francophone contemporaine dans un contexte postcolonial volontairement décentré d’une hégémonie culturelle occidentale, ici européenne. L’un des postulats de cet ouvrage, produit du Centre d’Étude des Nouveaux Espaces Littéraires de l’Université Paris 13 à Villetaneuse en France, est d’observer, de voir et de donner à voir (et à lire) un corps “oriental” afin de le distinguer des habitus nationaux voire régionaux. Par une volonté d’analyse polysémique prudente se déjouant, dans la mesure du possible, d’un européocentrisme prégnant, les auteurs de cet ouvrage questionnent la validité d’une démarche comparatiste entre aires proche-orientale et extrême-orientale aux précédents limités. La révélation d’un corps oriental comme dénominateur commun à un corpus littéraire et visuel (photographie, cinéma, bande dessinée) ne sera néanmoins pas de mise. Il est plutôt question d’analyser comment une réflexion historique, socio-culturelle, politique, religieuse et identitaire affecte, marque et montre des corps hybrides dans une aire culturelle plus globale. Les corps inscrits dans un corpus maroco-japonais ont-ils la possibilité de se parler et de se voir? Qu’ont-ils en commun? Y a-t-il un regard extraeuropéen sur les représentations du corps comme “objet social, historique ou psychanalytique” (7)? En quoi ce regard affecterait-il le travail introspectif et représentatif de l’artiste? Et s’il n’y avait pas de corps oriental à proprement parler, pourrait-on parler de corps (ou de corpus) national? L’existence de rituels similaires (les bains et le hammam; la honte d’être vu nu et la hchouma par exemple) permettrait-elle de concilier des visions du corps féminin intrinsèquement [End Page 217] différentes entre monde arabo-islamique, où son existence en changement est codifiée par la collectivité masculine et religieuse, et espace japonais mythique, religieux et fantastique dans lequel le corps féminin nu (parfois dénué d’érotisme) est omniprésent pour un lecteur occidental qui le quête? L’hétéronormativité fausserait-elle l’impact de la littérature féminine et de la littérature “queer” en les (re)présentant en tant qu’objets marginaux mettant à mal le principe d’appartenance identitaire unique? Comment aborder un corps militaire (principalement masculin) dont l’identité est à jamais marquée par une défaite brutale, et qui personnifie la souffrance de l’échec dans un monde postnucléaire? Et comment envisager le corps corporatif de l’ouvrier et de l’employé qui souffre d’un malaise identitaire dans le Japon des années 1960 et 1970 où modernisation rime avec nouvelle représentation et rejet des traditions? Le corps, cet “objet sémiologique” (15) est un lieu d’enjeux vitaux. C’est un élément perturbateur et perturbé, symptôme de son époque. Il personnifie l’implosion du corps social, il contredit les normes d’hier, il réécrit celles de demain. Il est vu à travers la lunette identitaire, historique et socio-culturelle de celui qui voit d’une manière qui n’est pas sans rappeler l’œuvre visuelle Étant donnés 1e la chute d’eau, 2e le gaz d’éclairage de Marcel Duchamp. Il emprunte à d’autres formats culturels afin d’assurer la survie de son message face à la censure. Il défie les définitions en offrant d’autres mots au champ lexical vernaculaire. Il explore et/ou déjoue les espaces physiologiques dans lesquels il est confiné pour poétiser sur une quête identitaire ambivalente dans laquelle “je est autre” selon la formule consacrée d’Arthur Rimbaud dans sa lettre à Paul Demeny datée du 15 mai 1871. En revisitant de nombreux textes dont des textes mythologiques et...
To solve the problem of lower precision caused by traditional query expansion technology, a new query expansion technique based on semantic context was proposed. The semantic context is constructed by WordNet knowledge base and related feedback documents. Firstly, the query words senses are confirmed by disambiguation with WordNet lexical database. Secondly, the initial expansion words are obtained according to the WordNet semantic hierarchy structure. Finally, the weight of the expansion terms is determined according to the overall correlation of candidate expansion terms and all the query words. These words whose weight is higher than weight threshold will be chosen as the final query expansion word. The experimental results show that the proposed method obviously improves the retrieval precision while preserving higher recall.
We present HamleDT - a HArmonized Multi-LanguagE Dependency Treebank. HamleDT is a compilation of existing dependency treebanks (or depen- dency conversions of other treebanks), transformed so that they all conform to the same annotation style. In the present article, we provide a thorough investigation and discussion of a number of phenomena that are comparable across languages, though their annotation in treebanks often differs. We claim that transformation procedures can be designed to automatically identify most such phenomena and convert them to a unified annotation style. This unification is beneficial both to comparative corpus linguistics and to machine learning of syntactic parsing.
This paper presents a survey of Arabic treebanks to facilitate their reuse for the building of new linguistic resources. In our case, we created from a treebank an automatically induced Property Grammar (GP). So, we discussed characteristics of these treebanks to choose the appropriate one. To build our resource, we adopted an automatic technique, acquiring first a contextfree grammar (CFG) from the chosen treebank, and second, inducing a GP by generating relations between grammatical units described in the CFG.
We present an algorithm and implementation for extracting recurring fragments from treebanks. Using a tree-kernel method the largest common fragments are extracted from each pair of trees. The algorithm presented achieves a thirty-fold speedup over the previously available method on the Wall Street Journal dataset. It is also more general, in that it supports trees with discontinuous constituents. The resulting fragments can be used as a tree-substitution grammar or in classification problems such as authorship attribution and other stylometry tasks.
This book offers an exciting new perspective on the origins of language. Language is conceptualized as a collective invention, on the model of writing or the wheel, and the book places social and cultural dynamics at the centre of its evolution: language emerged and further developed in human communities already suffused with meaning and communication, mimesis, ritual, song and dance, coparenting, new divisions of labour, and revolutionary changes in social relations. The book thus challenges assumptions about the causal relations between genes, capacities, social communication, and innovation: the biological capacities are taken to evolve incrementally on the basis of cognitive plasticity, in a process that recruits previous adaptations and fine-tunes them to serve novel communicative ends. Topics include the ability brought about by language to tell lies, which must have confronted our ancestors with new problems of public trust; the dynamics of social-cognitive co-evolution; the role of gesture and mimesis in linguistic communication; studies of how monkeys and apes express their feelings or thoughts; play, laughter, dance, song, ritual, and other social displays among extant hunter-gatherers; the social nature of language acquisition and innovation; normativity and the emergence of linguistic norms; the interaction of language and emotions; and novel perspectives on the timeframe for language evolution. The contributors are leading international scholars from linguistics, anthropology, paleontology, primatology, psychology, evolutionary biology, artificial intelligence, archaeology, and cognitive science. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
How do children learn to restrict their productivity and avoid ungrammatical utterances? The present study addresses this question by examining why some verbs are used with un- prefixation (e.g., unwrap) and others are not (e.g., *unsqueeze). Experiment 1 used a priming methodology to examine children's (3–4; 5–6) grammatical restrictions on verbal un- prefixation. To elicit production of un-prefixed verbs, test trials were preceded by a prime sentence, which described reversal actions with grammatical un- prefixed verbs (e.g., Marge folded her arms and then she unfolded them). Children then completed target sentences by describing cartoon reversal actions corresponding to (potentially) un- prefixed verbs. The younger age-group's production probability of verbs in un- form was negatively related to the frequency of the target verb in bare form (e.g., squeez/e/ed/es/ing), while the production probability of verbs in un- form for both age groups was negatively predicted by the frequency)
One of the most well studied ecological patterns is Rapoport's rule, which posits that the geographical extent of species ranges increases at higher latitudes. However, studies to date have been limited in their geographic scope and results have been equivocal. In turn, much debate exists over potential links between Rapoport's rule and latitudinal patterns in species richness. Humans collectively speak nearly 7000 different languages, which are spread unevenly across the globe, with loci in the tropics. Causes of this skewed distribution have received only limited study. We analyze the extent of Rapoport's rule in human languages at a global scale and within each region of the globe separately. We test the relationship between Rapoport's rule and the richness of languages spoken in different regions. We also explore the frequency distribution of language-range sizes. The language-range area distribution is strongly right-skewed, with 87% of languages having range areas less than 10,0)
Words are built from smaller meaning bearing parts, called morphemes. As one word can contain multiple morphemes, one morpheme can be present in different words. The number of distinct words a morpheme can be found in is its family size. Here we used Birth-Death-Innovation Models (BDIMs) to analyze the distribution of morpheme family sizes in English and German vocabulary over the last 200 years. Rather than just fitting to a probability distribution, these mechanistic models allow for the direct interpretation of identified parameters. Despite the complexity of language change, we indeed found that a specific variant of this pure stochastic model, the second order linear balanced BDIM, significantly fitted the observed distributions. In this model, birth and death rates are increased for smaller morpheme families. This finding indicates an influence of morpheme family sizes on vocabulary changes. This could be an effect of word formation, perception or both. On a more general level, )
In this paper, we conduct a study about differences between female and male discursive strategies when posting in the microblogging service Twitter, with a particular focus on the hashtag designation process during political debate. The fact that men and women use language in distinct ways, reverberating practices linked to their expected roles in the social groups, is a linguistic phenomenon known to happen in several cultures and that can now be studied on the Web and on online social networks in a large scale enabled by computing power. Here, for instance, after analyzing tweets with political content posted during Brazilian presidential campaign,we found out that male Twitter users, when expressing their attitude toward a given candidate, are more prone to use imperative verbal forms in hashtags, while female users tend to employ declarative forms. This difference can be interpreted as a sign of distinct approaches in relation to other network members: for example, if political ha)
A fundamental principle of brain organization is bilateral symmetry of structures and functions. For spatial sensory and motor information processing, this organization is generally plausible subserving orientation and coordination of a bilaterally symmetric body. However, breaking of the symmetry principle is often seen for functions that depend on convergent information processing and lateralized output control, e.g. left hemispheric dominance for the linguistic speech system. Conversely, a subtle splitting of functions into hemispheres may occur if peripheral information from symmetric sense organs is partly redundant, e.g. auditory pattern recognition, and therefore allows central conceptualizations of complex stimuli from different feature viewpoints, as demonstrated e.g. for hemispheric analysis of frequency modulations in auditory cortex (AC) of mammals including humans. Here we demonstrate that discrimination learning of rapidly but not of slowly amplitude modulated tones is n)
The present study was carried out in the Indo-European speaking tribal population groups of Southern Gujarat, India to investigate and reconstruct their paternal population structure and population histories. The role of language, ethnicity and geography in determining the observed pattern of Y haplogroup clustering in the study populations was also examined. A set of 48 bi-allelic markers on the non-recombining region of Y chromosome (NRY) were analysed in 284 males; representing nine Indo-European speaking tribal populations. The genetic structure of the populations revealed that none of these groups was overtly admixed or completely isolated. However, elevated haplogroup diversity and FST value point towards greater diversity and differentiation which suggests the possibility of early demographic expansion of the study groups. The phylogenetic analysis revealed 13 paternal lineages, of which six haplogroups: C5, H1a*, H2, J2, R1a1* and R2 accounted for a major portion of the Y chro)
Objective: This study aimed to develop a culturally acceptable and valid scale to assess depressive symptoms in older Indigenous Australians, to determine the prevalence of depressive disorders in the older Kimberley community, and to investigate the sociodemographic, lifestyle and clinical factors associated with depression in this population. Methods: Cross-sectional survey of adults aged 45 years or over from six remote Indigenous communities in the Kimberley and 30% of those living in Derby, Western Australia. The 11 linguistic and culturally sensitive items of the Kimberley Indigenous Cognitive Assessment of Depression (KICA-dep) scale were derived from the signs and symptoms required to establish the diagnosis of a depressive episode according to the DSM-IV-TR and ICD-10 criteria, and their frequency was rated on a 4-point scale ranging from ‘never’ to ‘all the time’ (range of scores: 0 to 33). The diagnosis of depressive disorder was established after a face-to-face assessment )
Few quantitative measures of genome architecture or organization exist to support assumptions of differences between microorganisms that are broadly defined as being free-living or pathogenic. General principles about complete proteomes exist for codon usage, amino acid biases and essential or core genes. Genome-wide shifts in amino acid usage between free-living and pathogenic microorganisms result in fundamental differences in the complexity of their respective proteomes that are size and gene content independent. These differences are evident across broad phylogenetic groups–a result of environmental factors and population genetic forces rather than phylogenetic distance. A novel comparative analysis of amino acid usage–utilizing linguistic analyses of word frequency in language and text–identified a global pattern of higher peptide word repetition in 376 free-living versus 421 pathogen genomes across broad ranges of genome size, G+C content and phylogenetic ancestry. This imprint )
The Qiangic languages in western Sichuan (WSC) are believed to be the oldest branch of the Sino-Tibetan linguistic family, and therefore, all Sino-Tibetan populations might have originated in WSC. However, very few genetic investigations have been done on Qiangic populations and no genetic evidences for the origin of Sino-Tibetan populations have been provided. By using the informative Y chromosome and mitochondrial DNA (mtDNA) markers, we analyzed the genetic structure of Qiangic populations. Our results revealed a predominantly Northern Asian-specific component in Qiangic populations, especially in maternal lineages. The Qiangic populations are an admixture of the northward migrations of East Asian initial settlers with Y chromosome haplogroup D (D1-M15 and the later originated D3a-P47) in the late Paleolithic age, and the southward Di-Qiang people with dominant haplogroup O3a2c1*-M134 and O3a2c1a-M117 in the Neolithic Age. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the proper)