Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Sound symbolism is the systematic and non-arbitrary link between word and meaning. Although a number of behavioral studies demonstrate that both children and adults are universally sensitive to sound symbolism in mimetic words, the neural mechanisms underlying this phenomenon have not yet been extensively investigated. The present study used functional magnetic resonance imaging to investigate how Japanese mimetic words are processed in the brain. In Experiment 1, we compared processing for motion mimetic words with that for non-sound symbolic motion verbs and adverbs. Mimetic words uniquely activated the right posterior superior temporal sulcus (STS). In Experiment 2, we further examined the generalizability of the findings from Experiment 1 by testing another domain: shape mimetics. Our results show that the right posterior STS was active when subjects processed both motion and shape mimetic words, thus suggesting that this area may be the primary structure for processing sound symb)
Behavioural evidence suggests that English regular past tense forms are automatically decomposed into their stem and affix (played = play+ed) based on an implicit linguistic rule, which does not apply to the idiosyncratically formed irregular forms (kept). Additionally, regular, but not irregular inflections, are thought to be processed through the procedural memory system (left inferior frontal gyrus, basal ganglia, cerebellum). It has been suggested that this distinction does not to apply to second language (L2) learners of English; however, this has not been tested at the brain level. This fMRI study used a masked-priming task with regular and irregular prime-target pairs (played-play/kept-keep) to investigate morphological processing in native and highly proficient late L2 English speakers. No between-groups differences were revealed. Compared to irregular pairs, regular pairs activated the pars opercularis, bilateral caudate nucleus and the right cerebellum, which are part of the)
The present study investigated the relationship between Chinese reading skills and metalinguistic awareness skills such as phonological, morphological, and orthographic awareness for 101 Preschool, 94 Grade-1, 98 Grade-2, and 98 Grade-3 children from two primary schools in Mainland China. The aim of the study was to examine how each of these metalinguistic awareness skills would exert their influence on the success of reading in Chinese with age. The results showed that all three metalinguistic awareness skills significantly predicted reading success. It further revealed that orthographic awareness played a dominant role in the early stages of reading acquisition, and its influence decreased with age, while the opposite was true for the contribution of morphological awareness. The results were in stark contrast with studies in English, where phonological awareness is typically shown as the single most potent metalinguistic awareness factor in literacy acquisition. In order to account )
Nativists have postulated fundamental geometric knowledge that predates linguistic and symbolic thought. Central to these claims is the proposal for an isolated cognitive system dedicated to processing geometric information. Testing such hypotheses presents challenges due to difficulties in eliminating the combination of geometric and non-geometric information through language. We present evidence using a modified matching interference paradigm that an incongruent shape word interferes with identifying a two-dimensional geometric shape, but an incongruent two-dimensional geometric shape does not interfere with identifying a shape word. This asymmetry in interference effects between two-dimensional geometric shapes and their corresponding shape words suggests that shape words activate spatial representations of shapes but shapes do not activate linguistic representations of shape words. These results appear consistent with hypotheses concerning a cognitive system dedicated to processin)
Three experiments provide evidence of an incipient sense of fairness in preverbal infants. Ten-month-old infants were shown cartoon videos with two agents, the ‘donors’, who distributed resources to two identical recipients. One donor always distributed the goods equally, while the other performed unequal distributions by giving everything to one recipient. In the test phase, a third agent hit or took resources away from either the fair or the unfair donor. We found that infants looked longer when the antisocial actions were directed towards the unfair rather than the fair donor. These findings support the view that infants are able to evaluate agents based on their distributive actions and suggest that the foundations of human socio-moral competence are acquired independently of parental feedback and linguistic experience. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted)
Lewis Carroll's English word game Doublets is represented as a system of networks with each node being an English word and each connectivity edge confirming that its two ending words are equal in letter length, but different by exactly one letter. We show that this system, which we call the Doublets net, constitutes a complex body of linguistic knowledge concerning English word structure that has computable multiscale features. Distributed morphological, phonological and orthographic constraints and the language's local redundancy are seen at the node level. Phonological communities are seen at the network level. And a balancing act between the language's global efficiency and redundancy is seen at the system level. We develop a new measure of intrinsic node-to-node distance and a computational algorithm, called community geometry, which reveal the implicit multiscale structure within binary networks. Because the Doublets net is a modular complex cognitive system, the community geomet)
The article focuses on the changes in affective (emotional) ratings under the influence of the choice of one of the interpretations of ambiguous images and subsequent recognition task performance. Earlier studies showed that recognition decision affects subsequent ratings of the stimulus: the more information is accumulated about the stimulus, the more positive will be its ratings at recognition, and the more negative at non-recognition (Chetverikov, 2014). We hypothesized that a choice of a single interpretation of a stimulus becomes a source of information for subsequent decision concerning recognition or non-recognition of the unambiguous interpretation of that stimulus. Thus, this decision will affect subsequent ratings of stimuli the same way as in the case of initially unambiguous stimuli. The experimental results confirmed our hypothesis. Ratings of unambiguous stimuli corresponding to selected and non-selected interpretation of ambiguous stimuli varied depending on the recognition decision in the same way as did ratings of previously presented and new unambiguous stimuli. When a stimulus is «old» and is recognized, it is liked more, than a recognized «new» stimulus; when it is not recognized, the effect is opposite. Thus, the more information about the stimulus has been accumulated, the higher is the influence of a decision concerning stimulus recognition on subsequent ratings. Similar results were found for confidence ratings. These were higher in the case of recognition than in the case of non-recognition, but the difference between the two situations was more pronounced for «old» stimuli than for «new» ones.
Ever since the publication of À la recherche du temps perdu, the novel's language has attracted critical attention, and the excellent ten-page bibliography of linguistic and stylistic studies on Proust provided at the end of the present collection testifies to the richness of the field. This volume sets out to explore aspects of the linguistic form of À la recherche with a view to unravelling how an idiosyncratic use turns a langue — a set of norms shared by a community — into literature. Close-ups of specific rhetorical, grammatical, and lexical phenomena examined in a synchronic and/or a diachronic perspective allow the authors to highlight how certain stylistic effects are achieved and to challenge clichés. Thus Stéphanie Fonvielle's analyses of Proust's use of maxims and (with Hugues Galli) of the distinguo provide brilliant insights into the rhetorical tradition that comes to serve Proust's textual architectural purposes, while Geneviève Henrot Sostero argues that the various forms of dislocation contribute to the effect of a seamless fondu-enchaîné of the text. Sylvie Bougeard-Pierron's revisiting of the author's allegedly innovative vocabulary, on the other hand, demonstrates that Proust, rather than being an inventor, was an absorptive reader picking up rarities and novelties from all possible sources. Discussing the text's heterogeneity, Michel Sandras concludes that it indicates the integration of different ‘modèle[s] de langue’ (p. 129). Between rhetoric and linguistics, Jacques Dürrenmatt introduces us to the unexpectedly fascinating history of the semicolon and what its use tells us about Proust. Sophie Duval and Stéphane Chaudier move closer to discourse analysis in their respective focus on Jewish humour and irony and the use of the term ‘loi’, while Maribel Peñalver Vicea approaches psychology by examining the potential of autonymy to reveal affects. Isabelle Serça and Davide Vago provide complementary approaches to Proust's use of synonyms, the first as part of a broader interest in lists that ‘met[tent] en scène l’écriture […] dans ses tâtonnements' (p. 191), the second with a more specific focus on pseudo-synonyms, concluding that Proust's practice calls for rethinking established categories. The grammatical issues addressed include the uses and forms of the conditional, confirming the richness of Proust's syntactic toolkit (Danielle Coltier and Patrick Dendale), and the combination of aujourd'hui with the imperfect, reaffirming Hans Robert Jauss's challenge to the interpretation of the imperfect as a marker of timelessness (Anna Isabella Squarzina). The genetic approach is also represented by Yasué Kato's dissection of a description of the ‘petite bande’, which testifies to the gradual expansion, with the reworkings, of the textual network in which the passage inscribes itself. Michele Prandi's general introduction to the relation between linguistics and literature, followed by Henrot Sostero's orientational overview of existing linguistic studies on Proust, and Dürrenmatt's closing summary of the questions that remain to be explored frame the volume. Students of Proust (even without specialized knowledge of linguistics) and linguists intrigued by literary language will find these studies equally rewarding. The index of quoted passages will prove particularly useful for anyone interested in a specific moment in the novel.
AIMS: To explore the enhancing effect of alcohol consumption on attractiveness ratings, in that few studies on the Beer Goggles effect control the stimuli attractiveness level and researchers have seldom considered extending the effect to stimuli other than faces. METHODS: Male and female participants (n = 103) were randomly assigned to alcohol consumption or placebo groups. Both groups were asked to assess the attractiveness of two types of pictures (faces and landscapes) with three levels of attractiveness for each stimulus category (high, moderate and low). RESULTS: We found significant interactions between beverage type and attractiveness level. Attractiveness ratings for moderate- and low-attractiveness faces were significantly higher in the alcohol compared with placebo condition, while there was no significant difference for high-attractiveness stimuli between these two conditions. As for landscapes, only low-attractiveness stimuli were rated significantly higher in the alcohol condition. CONCLUSION: Whether or not alcohol consumption leads to an increase in attractiveness ratings depends on the initial attractiveness of the stimulus materials. Alcohol consumption tends to affect ratings for stimuli with relatively low attractiveness. Furthermore, this effect is not limited to faces; it extends to other types of stimuli like landscapes.
The paper describes a general framework for mining large amounts of text data from a defined set of Web pages. The acquired data are meant to constitute a corpus for training robust and reliable language models and thus the framework needs to also incorporate algorithms for appropriate text processing and duplicity detection in order to secure quality and consistency of the data. As we expect the resulting corpus to be very large, we have also implemented topic detection algorithms that allow us to automatically select subcorpora for domain-specific language models. The description of the framework architecture and the implemented algorithms is complemented with a detailed evaluation section. It analyses the basic properties of the gathered Czech corpus containing more than one billion text tokens collected using the described framework, shows the results of the topic detection methods and finally also describes the design and outcomes of the automatic speech recognition experiments with domain-specific language models estimated from the collected data.
Culture is a phenomenon shared by all humans. Attempts to understand how dynamic factors affect the origin and distribution of cultural elements are, therefore, of interest to all humanity. As case studies go, understanding the distribution of cultural elements in Native American communities during the historical period of the Great Plains would seem a most challenging one. Famously, there is a mixture of powerful internal and external factors, creating-for a relatively brief period in time-a seemingly distinctive set of shared elements from a linguistically diverse set of peoples. This is known across the world as the “Great Plains culture.” Here, quantitative analyses show how different processes operated on two sets of cultural traits among nine High Plains groups. Moccasin decorations exhibit a pattern consistent with geographically-mediated between-group interaction. However, group variations in the religious ceremony of the Sun Dance also reveal evidence of purifying cultural se)
The lexical approach is a method in differential psychology that uses people's estimations of verbal descriptors of human behavior in order to derive the structure of human individuality. The validity of the assumptions of this method about the objectivity of people's estimations is rarely questioned. Meanwhile the social nature of language and the presence of emotionality biases in cognition are well-recognized in psychology. A question remains, however, as to whether such an emotionality-capacities bias is strong enough to affect semantic perception of verbal material. For the lexical approach to be valid as a method of scientific investigations, such biases should not exist in semantic perception of the verbal material that is used by this approach. This article reports on two studies investigating differences between groups contrasted by 12 temperament traits (i.e. by energetic and other capacities, as well as emotionality) in the semantic perception of very general verbal materia)
Obtaining syntactic parses is an important step in many NLP pipelines. However, most of the world’s languages do not have a large amount of syntactically annotated data available for building parsers. Syntactic projection techniques attempt to address this issue by using parallel corpora consisting of resource-poor and resource-rich language pairs, taking advantage of a parser for the resource-rich language and word alignment between the languages to project the parses onto the data for the resource-poor language. These projection methods can suffer, however, when syntactic structures for some sentence pairs in the two languages look quite different. In this paper, we investigate the use of small, parallel, annotated corpora to automatically detect divergent structural patterns between two languages. We then use these detected patterns to improve projection algorithms and dependency parsers, allowing for better performing NLP tools for resource-poor languages, particularly those that may not have large amounts of annotated data necessary for traditional, fully-supervised methods. While this detection process is not exhaustive, we demonstrate that common patterns of divergence can be identified automatically without prior knowledge of a given language pair, and the patterns can be used to improve performance of syntactic projection and parsing.
Background: Evidence from a number of countries in Europe and North America point towards the secular declining trend in menarcheal age with considerable spatial variations over the past two centuries. Similar trends were reported in several developing countries from Asia, Africa and Latin America. However, data corroborating any secular trend in the menarcheal age of the Indian population remained sparse and inadequately verified. Methods: We examined secular trends, regional heterogeneity and association of socioeconomic, anthropometric and contextual factors with menarcheal age among ever-married women (15–49 years) in India. Using the pseudo cohort data approach, we fit multiple linear regression models to estimate secular trends in menarcheal age of 91394 ever-married women using the Indian Human Development Survey. Results: The mean age at menarche among Indian women was 13.76 years (95 % CI: 13.75, 13.77) in 2005. It declined by three months from 13.83 years (95% CI: 13.81, )
Helping behavior as a prosocial action emerges early in childhood and is of interest for psychologists in a broad range of sub disciplines as well as for society. One necessary precondition for active helping is the ability to recognize that somebody needs help. The NeoHelp stimulus set used in this study was developed to enable the assessment and quantification of need-of-help recognition abilities. Previous research with the NeoHelp stimuli has shown that children of different ages are able to recognize their content. Specific effects of age and gender on need-of-help recognition have also been observed. How children subjectively experience these stimuli and thus how they rate depictions of need-of-help and no-need-of-help situations emotionally has not been assessed before. Here we report analyses of valence and arousal ratings for the complete NeoHelp stimulus set obtained from a diverse sample of 46 children. We employed the SAM-scales because they are an established rating instrument validated for diverse populations of adults. However, their use with children still needs further investigation. Thus, there were two main goals of the presented study: 1) Validating that the SAM arousal and valence scales may be used with young children below school age, and 2) investigating children's subjective emotional experience of need-of-help depictions. Our study demonstrates that the SAM scales, if properly explained, may be used reliably with children at and above five years of age. Ratings of younger and older children covered the whole range of the 5-point scales used. There was a linear relationship between arousal and valence ratings across all pictures: the higher the arousal ratings, the lower the valence ratings. Pictures showing a child in need-of-help were rated as lower in valence and higher in arousal than the corresponding no-need-of-help-stimuli regardless of children’s age or gender. With increasing age, arousal ratings for no-need-of-help depictions decreased, but arousal ratings for need-of-help depictions remained on the same higher level across ages. We thus provide first evidence that need-of-help depictions elicit differential subjective emotional responses in children on both, valence and arousal dimensions. This emotional component of need-of-help recognition has to be considered when assessing children's need-of-help recognition abilities.
An intricate history of human dispersal and geographic colonization has strongly affected the distribution of human pathogens. The pig tapeworm Taenia solium occurs throughout the world as the causative agent of cysticercosis, one of the most serious neglected tropical diseases. Discrete genetic lineages of T. solium in Asia and Africa/Latin America are geographically disjunct; only in Madagascar are they sympatric. Linguistic, archaeological and genetic evidence has indicated that the people in Madagascar have mixed ancestry from Island Southeast Asia and East Africa. Hence, anthropogenic introduction of the tapeworm from Southeast Asia and Africa had been postulated. This study shows that the major mitochondrial haplotype of T. solium in Madagascar is closely related to those from the Indian Subcontinent. Parasitological evidence presented here, and human genetics previously reported, support the hypothesis of an Indian influence on Malagasy culture coinciding with periods of early )
Background: Small clinical trials have reported that low-frequency repetitive transcranial magnetic stimulation (rTMS) might improve language recovery in patients with aphasia after stroke. However, no systematic reviews or meta-analyses studies have investigated the effect of rTMS on aphasia. The objective of this study was to perform a meta-analysis of studies that explored the effects of low-frequency rTMS on aphasia in stroke patients. Methods: We searched PubMed, CENTRAL, Embase, CINAHL, ScienceDirect, and Journals@Ovid for randomized controlled trials published between January 1965 and October 2013 using the keywords “aphasia OR language disorders OR anomia OR linguistic disorders AND repetitive transcranial magnetic stimulation OR rTMS”. We used fixed- and random-effects models to estimate the standardized mean difference (SMD) and a 95% CI for the language outcomes. Results: Seven eligible studies involving 160 stroke patients were identified in this meta-analysis. A significa)
ILSP Dependency Parser is a tool trained on the Greek Dependency Treebank, a resource which comprises data annotated at several linguistic levels. Training data at the level of syntax consisted of ~70 KWords annotated using a dependency-based syntactic scheme that includes 25 main relations.
This article introduces iVAR, an R program for imputing missing data in multivariate time series on the basis of vector autoregressive (VAR) models. We conducted a simulation study to compare iVAR with three methods for handling missing data: listwise deletion, imputation with sample means and variances, and multiple imputation ignoring time dependency. The results showed that iVAR produces better estimates for the cross-lagged coefficients than do the other three methods. We demonstrate the use of iVAR with an empirical example of time series electrodermal activity data and discuss the advantages and limitations of the program.
The aim of the demo is threefold. First, it introduces the current version of the annotation tool for discourse relations in the Prague Dependency Treebank 3.0. Second, it presents the discourse relations in the treebank themselves, including new additions in comparison with the previous release. And third, it shows how to search in the treebank, with focus on the discourse relations.
In the adult brain, speech can recruit a brain network that is overlapping with, but not identical to, that involved in perceiving non-linguistic vocalizations. Using the same stimuli that had been presented to human 4-month-olds and adults, as well as adult macaques, we sought to shed light on the cortical networks engaged when human newborns process diverse vocalization types. Near infrared spectroscopy was used to register the response of 40 newborns' perisylvian regions when stimulated with speech, human and macaque emotional vocalizations, as well as auditory controls where the formant structure was destroyed but the long-term spectrum was retained. Left fronto-temporal and parietal regions were significantly activated in the comparison of stimulation versus rest, with unclear selectivity in cortical activation. These results for the newborn brain are qualitatively and quantitatively compared with previous work on newborns, older human infants, adult humans, and adult macaques re)
Bilingual Base Noun Phrase (BaseNP) extraction is one of the key tasks of Natural Language Processing (NLP). This task is more challenging for the pair of English-Vietnamese due to the lack of available Vietnamese language resources such as treebanks, part-of-speech taggers, and parsers. In this paper, we propose a combination model that uses language characteristics based on statistics and the projection method to extract BaseNP correspondences from a bilingual corpus. The language characteristics used in this model include the word segmentation, word order and word classification [1]. Our model overcomes not only the lack of resources of Vietnamese, but also improves the performance of miss-alignment, null-alignment, overlap and conflict projection of the existing methods. The proposed model can be easily applied to other language pairs. Experiment on 66,646 pairs of sentences in the English-Vietnamese bilingual corpus shows that our proposed model is very satisfactory.
Do narratives shape how humans process other minds or do they presuppose an existing theory of mind? This study experimentally investigated this problem by assessing subject responses to systematic alterations in the genre, levels of intentionality, and linguistic complexity of narratives. It showed that the interaction of genre and intentionality level are crucial in determining how narratives are cognitively processed. Specifically, genres that deployed evolutionarily familiar scenarios (relationship stories) were rated as being higher in quality when levels of intentionality were increased; conversely, stories that lacked evolutionary familiarity (espionage stories) were rated as being lower in quality with increases in intentionality level. Overall, the study showed that narrative is not solely either the origin or the product of our intuitions about other minds; instead, different genres will have different—even opposite—effects on how we understand the mind states of others. [AB)
Whether mathematical and linguistic processes share the same neural mechanisms has been a matter of controversy. By examining various sentence structures, we recently demonstrated that activations in the left inferior frontal gyrus (L. IFG) and left supramarginal gyrus (L. SMG) were modulated by the Degree of Merger (DoM), a measure for the complexity of tree structures. In the present study, we hypothesize that the DoM is also critical in mathematical calculations, and clarify whether the DoM in the hierarchical tree structures modulates activations in these regions. We tested an arithmetic task that involved linear and quadratic sequences with recursive computation. Using functional magnetic resonance imaging, we found significant activation in the L. IFG, L. SMG, bilateral intraparietal sulcus (IPS), and precuneus selectively among the tested conditions. We also confirmed that activations in the L. IFG and L. SMG were free from memory-related factors, and that activations in the bi)
Lexical Knowledge base such as WordNet has been used as a valuable tool for measuring semantic similarity in various Information Retrieval (IR) applications. It is a domain independent lexical database. Since, the quality of semantic relationship in WordNet has not upgraded appropriately for the current usage in the modern IR. Building the WordNet from scratch is not an easy task for keeping updated with current terminology and concepts. Therefore, this paper undergoes a different perspective that automatically updates an existing lexical ontology uses knowledge resources such as the Wikipedia and the Web search engine. This methodology has established the recently evolving relations and also aligns the existing relations between concepts based on its usage over time. It consists of three main phases such as candidate article generation, lexical relationship extraction and generalization and WordNet alignment. In candidate article generation, disambiguation mapping disambiguates ambiguous links between WordNet concepts and Wikipedia articles and returns a set of word-article pairings. Lexical relationship extraction phase includes two algorithms, Lexical Relationship Retrieval (LRR) algorithm discovers the set of lexical patterns exists between concepts and sequential pattern grouping algorithm generalizes lexical patterns and computes corresponding weights based on its frequencies. Furthermore, Sequential Minimal Optimization (SMO) selects the suitable good pattern using the optimal combination of weight of lexical patterns and page count based concurrence measures. WordNet alignment phase establishes a new relationship that is not available in WordNet and also aligns the existing patterns based on computed weight. Experimental results illustrate that the proposed approach better than existing mechanisms on benchmark datasets and achieves a correlation value of 0.87. Moreover, the extended WordNet returns high accuracy results in query expansion.
The present paper explored the relationship between emotional facial response and electromyographic modulation in children when they observe facial expression of emotions. Facial responsiveness (evaluated by arousal and valence ratings) and psychophysiological correlates (facial electromyography, EMG) were analyzed when children looked at six facial expressions of emotions (happiness, anger, fear, sadness, surprise and disgust). About EMG measure, corrugator and zygomatic muscle activity was monitored in response to different emotional types. ANOVAs showed differences for both EMG and facial response across the subjects, as a function of different emotions. Specifically, some emotions were well expressed by all the subjects (such as happiness, anger and fear) in terms of high arousal, whereas some others were less level arousal (such as sadness). Zygomatic activity was increased mainly for happiness, from one hand, corrugator activity was increased mainly for anger, fear and surprise, from the other hand. More generally, EMG and facial behavior were highly correlated each other, showing a "mirror" effect with respect of the observed faces.
The theory of embodied language states that language comprehension relies on an internal reenactment of the sensorimotor experience associated with the processed word or sentence. Most evidence in support of this hypothesis had been collected using linguistic material without any emotional connotation. For instance, it had been shown that processing of arm-related verbs, but not of those leg-related verbs, affects the planning and execution of reaching movements; however, at present it is unknown whether this effect is further modulated by verbs evoking an emotional experience. Showing such a modulation might shed light on a very debated issue, i.e. the way in which the emotional meaning of a word is processed. To this end, we assessed whether processing arm/hand-related verbs describing actions with negative connotations (e.g. to stab) affects reaching movements differently from arm/hand-related verbs describing actions with neutral connotation (e.g. to comb). We exploited a go/no-go)
WordNet is an electronic lexical database available on-line as a powerful resource to the researchers in the area of computational linguistics, text processing and other related areas. WordNet for Hindi language has already been developed by IIT, Bombay. The Indian languages WordNets are being created using expansion approach from Hindi WordNet under IndoWordNet project. In expansion approach, semantic relations are borrowed from the reference language, while the lexical relations need to be created for each language, as these relations are language dependent. This paper describes the process of creation of lexical relations like antonym, compounding, conjunction and gradation for IndoWordNet. A lexical creation tool has been presented in this paper with provision to create lexical relations in target language on the basis of relations created in Hindi WordNet and with another provision to create lexical relations in target language without referring to Hindi WordNet. It has been observed that lexical relations for target language can be created easily on the basis of relations created in Hindi WordNet for Hindi in-family languages, while for the languages that do not fall in the same family provision of creation of lexical relation without referring to Hindi WordNet can be used.
Key to fast adaptation of language technologies for any language hinges on the availability of fundamental tools and resources such as monolingual/parallel corpora, annotated corpora, part-of-speech (POS) taggers, parsers and so on. The languages which lack those fundamental resources are often referred as under-resourced\nlanguages.\n\nIn this thesis, we address the problem of cross-lingual dependency parsing of under-resourced languages. We apply three methodologies to induce dependency structures: (i) projecting dependencies from a resource-rich language to under-resourced languages via parallel corpus word alignment links (ii) parsing under-\nresourced languages using parsers whose models are trained on treebanks of other\nlanguages, and do not look at actual word forms, but only on POS categories. Here\nwe address the problem of incompatibilities in annotation styles between source side parsers and target side evaluation treebanks by harmonizing annotations to a common standard; and finally (iii) we add a new under-resourced scenario in which we use machine translated parallel corpora instead of human translated corpora for\nprojecting dependencies to under-resourced languages.\n\nWe apply the aforementioned methodologies to five Indian languages (ILs): Hindi, Urdu, Telugu, Bengali and Tamil (in the order of high to low availability of treebank data). To make the evaluation possible for Tamil, we develop a depen\ndency treebank resource for Tamil from scratch and we use the created data in\nevaluation and as a source in parsing other ILs. Finally, we list out strategies that\ncan be used to obtain dependency structures for target languages under different\nresource-poor scenarios.
Signal detection in clinical trials relies on ratings reliability. We conducted a reliability analysis of site-independent rater scores derived from audio-digital recordings of site-based rater interviews of the structured Brief Psychiatric Rating Scale (BPRS) in a schizophrenia study. "Dual" ratings assessments were conducted as part of a quality assurance program in a 12-week, double-blind, parallel-group study of PF-02545920 compared to placebo in patients with sub-optimally controlled symptoms of schizophrenia (ClinicalTrials.gov identifier NCT01939548). Blinded, site-independent raters scored the recorded site-based BPRS interviews that were administered in relatively stable patients during two visits prior to the randomization visit. We analyzed the impact of BPRS interview length on "dual" scoring variance and discordance between trained and certified site-based raters and the paired scores of the independent raters. Mean total BPRS scores for 392 interviews conducted at the screen and stabilization visits were 50.4±7.2 (SD) for site-based raters and 49.2±7.2 for site-independent raters (t=2.34; p=0.025). "Dual" rated total BPRS scores were highly correlated (r=0.812). Mean BPRS interview length was 21:05±7:47min ranging from 7 to 59min. 89 interviews (23%) were conducted in less than 15min. These shorter interviews had significantly greater "dual" scoring variability (p=0.0016) and absolute discordance (p=0.0037) between site-based and site-independent raters than longer interviews. In-study ratings reliability cannot be guaranteed by pre-study rater certification. Our findings reveal marked variability of BPRS interview length and that shorter interviews are often incomplete yielding greater "dual" scoring discordance that may affect ratings precision.
We investigate the feasibility of aligning Chinese and English parse trees by examining cases of incompatibility between Chinese-English parallel parse trees. This work is done in the context of an annotation project where we construct a parallel treebank by doing word and phrase alignments simultaneously. We discuss the most common incompatibility patterns identified within VPs and NPs and show that most cases of incompatibility are caused by divergent syntactic annotation standards rather than inherent cross-linguistic differences in language itself. This suggests that in principle it is feasible to align the parallel parse trees with some modification of existing syntactic annotation guidelines. We believe this has implications for the use of parallel parse trees as an important resource for Machine Translation models.
Even though auditory stimuli do not directly convey information related to visual stimuli, they often improve visual detection and identification performance. Auditory stimuli often alter visual perception depending on the reliability of the sensory input, with visual and auditory information reciprocally compensating for ambiguity in the other sensory domain. Perceptual processing is characterized by hemispheric asymmetry. While the left hemisphere is more involved in linguistic processing, the right hemisphere dominates spatial processing. In this context, we hypothesized that an auditory facilitation effect in the right visual field for the target identification task, and a similar effect would be observed in the left visual field for the target localization task. In the present study, we conducted target identification and localization tasks using a dual-stream rapid serial visual presentation. When two targets are embedded in a rapid serial visual presentation stream, the target )
Reviewed by: Récits du corps au Maroc et au Japon ed. by Marc Kober and Khalid Zekri Gaëlle Corvaisier Kober, Marc, et Khalid Zekri, coords. Récits du corps au Maroc et au Japon. Paris: L’Harmattan, 2011. isbn 9782296557208. 200 p. Avec Récits du corps au Maroc et au Japon, Marc Kober et Khalid Zekri posent des questions essentielles pour la littérature francophone contemporaine dans un contexte postcolonial volontairement décentré d’une hégémonie culturelle occidentale, ici européenne. L’un des postulats de cet ouvrage, produit du Centre d’Étude des Nouveaux Espaces Littéraires de l’Université Paris 13 à Villetaneuse en France, est d’observer, de voir et de donner à voir (et à lire) un corps “oriental” afin de le distinguer des habitus nationaux voire régionaux. Par une volonté d’analyse polysémique prudente se déjouant, dans la mesure du possible, d’un européocentrisme prégnant, les auteurs de cet ouvrage questionnent la validité d’une démarche comparatiste entre aires proche-orientale et extrême-orientale aux précédents limités. La révélation d’un corps oriental comme dénominateur commun à un corpus littéraire et visuel (photographie, cinéma, bande dessinée) ne sera néanmoins pas de mise. Il est plutôt question d’analyser comment une réflexion historique, socio-culturelle, politique, religieuse et identitaire affecte, marque et montre des corps hybrides dans une aire culturelle plus globale. Les corps inscrits dans un corpus maroco-japonais ont-ils la possibilité de se parler et de se voir? Qu’ont-ils en commun? Y a-t-il un regard extraeuropéen sur les représentations du corps comme “objet social, historique ou psychanalytique” (7)? En quoi ce regard affecterait-il le travail introspectif et représentatif de l’artiste? Et s’il n’y avait pas de corps oriental à proprement parler, pourrait-on parler de corps (ou de corpus) national? L’existence de rituels similaires (les bains et le hammam; la honte d’être vu nu et la hchouma par exemple) permettrait-elle de concilier des visions du corps féminin intrinsèquement [End Page 217] différentes entre monde arabo-islamique, où son existence en changement est codifiée par la collectivité masculine et religieuse, et espace japonais mythique, religieux et fantastique dans lequel le corps féminin nu (parfois dénué d’érotisme) est omniprésent pour un lecteur occidental qui le quête? L’hétéronormativité fausserait-elle l’impact de la littérature féminine et de la littérature “queer” en les (re)présentant en tant qu’objets marginaux mettant à mal le principe d’appartenance identitaire unique? Comment aborder un corps militaire (principalement masculin) dont l’identité est à jamais marquée par une défaite brutale, et qui personnifie la souffrance de l’échec dans un monde postnucléaire? Et comment envisager le corps corporatif de l’ouvrier et de l’employé qui souffre d’un malaise identitaire dans le Japon des années 1960 et 1970 où modernisation rime avec nouvelle représentation et rejet des traditions? Le corps, cet “objet sémiologique” (15) est un lieu d’enjeux vitaux. C’est un élément perturbateur et perturbé, symptôme de son époque. Il personnifie l’implosion du corps social, il contredit les normes d’hier, il réécrit celles de demain. Il est vu à travers la lunette identitaire, historique et socio-culturelle de celui qui voit d’une manière qui n’est pas sans rappeler l’œuvre visuelle Étant donnés 1e la chute d’eau, 2e le gaz d’éclairage de Marcel Duchamp. Il emprunte à d’autres formats culturels afin d’assurer la survie de son message face à la censure. Il défie les définitions en offrant d’autres mots au champ lexical vernaculaire. Il explore et/ou déjoue les espaces physiologiques dans lesquels il est confiné pour poétiser sur une quête identitaire ambivalente dans laquelle “je est autre” selon la formule consacrée d’Arthur Rimbaud dans sa lettre à Paul Demeny datée du 15 mai 1871. En revisitant de nombreux textes dont des textes mythologiques et...
The Qiangic languages in western Sichuan (WSC) are believed to be the oldest branch of the Sino-Tibetan linguistic family, and therefore, all Sino-Tibetan populations might have originated in WSC. However, very few genetic investigations have been done on Qiangic populations and no genetic evidences for the origin of Sino-Tibetan populations have been provided. By using the informative Y chromosome and mitochondrial DNA (mtDNA) markers, we analyzed the genetic structure of Qiangic populations. Our results revealed a predominantly Northern Asian-specific component in Qiangic populations, especially in maternal lineages. The Qiangic populations are an admixture of the northward migrations of East Asian initial settlers with Y chromosome haplogroup D (D1-M15 and the later originated D3a-P47) in the late Paleolithic age, and the southward Di-Qiang people with dominant haplogroup O3a2c1*-M134 and O3a2c1a-M117 in the Neolithic Age. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the proper)
Few quantitative measures of genome architecture or organization exist to support assumptions of differences between microorganisms that are broadly defined as being free-living or pathogenic. General principles about complete proteomes exist for codon usage, amino acid biases and essential or core genes. Genome-wide shifts in amino acid usage between free-living and pathogenic microorganisms result in fundamental differences in the complexity of their respective proteomes that are size and gene content independent. These differences are evident across broad phylogenetic groups–a result of environmental factors and population genetic forces rather than phylogenetic distance. A novel comparative analysis of amino acid usage–utilizing linguistic analyses of word frequency in language and text–identified a global pattern of higher peptide word repetition in 376 free-living versus 421 pathogen genomes across broad ranges of genome size, G+C content and phylogenetic ancestry. This imprint )
Face judgments of dominance play an important role in human social interaction. Perceived facial dominance is thought to indicate physical formidability, as well as resource acquisition and holding potential. Dominance cues in the face affect perceptions of attractiveness, emotional state, and physical strength. Most experimental paradigms test perceptions of facial dominance in individual faces, or they use manipulated versions of the same face in a forced-choice task but in the absence of other faces. Here, we extend this work by assessing whether dominance ratings are absolute or are judged relative to other faces. We presented participants with faces to be rated for dominance (target faces), while also presenting a second face (non-target faces) that was not to be rated. We found that both the masculinity and sex of the non-target face affected dominance ratings of the target face. Masculinized non-target faces decreased the perceived dominance of a target face relative to a feminized non-target face, and displaying a male non-target face decreased perceived dominance of a target face more so than a female non-target face. Perceived dominance of male target faces was affected more by masculinization of male non-target faces than female non-target faces. These results indicate that dominance perceptions can be altered by surrounding faces, demonstrating that facial dominance is judged at least partly relative to other faces.
This web service performs dependency parsing in Spanish using a Malt Parser instance.It parses plain texts introduced by the user and generates linguistically annotated Treebank instances based on a data-driven parsing model.The parsing model is induced from de dependency-annotated IULA Treebank (Marimon et al, 2012) using the language-independent MaltParser 2 system as a dependency model trainer (Nivre et al, 2007).This Treebank contains 589,542 tokens in 42,099.In order to achieve optimal performance, the training corpus was previously analyzed with MaltOptimizer 3 (Ballesteros and Nivre, 2012), a tool developed to set the best parameters for MaltParser. Inputs, outputs and formats InputsThe input to be parsed is a plain text encoded in UTF-8.It can be introduced directly as a text instance in the dialogue box, as a text file or as a URL.The input language available at this moment in the web service is Spanish (es).
Objective: This study aimed to develop a culturally acceptable and valid scale to assess depressive symptoms in older Indigenous Australians, to determine the prevalence of depressive disorders in the older Kimberley community, and to investigate the sociodemographic, lifestyle and clinical factors associated with depression in this population. Methods: Cross-sectional survey of adults aged 45 years or over from six remote Indigenous communities in the Kimberley and 30% of those living in Derby, Western Australia. The 11 linguistic and culturally sensitive items of the Kimberley Indigenous Cognitive Assessment of Depression (KICA-dep) scale were derived from the signs and symptoms required to establish the diagnosis of a depressive episode according to the DSM-IV-TR and ICD-10 criteria, and their frequency was rated on a 4-point scale ranging from ‘never’ to ‘all the time’ (range of scores: 0 to 33). The diagnosis of depressive disorder was established after a face-to-face assessment )
The present study was carried out in the Indo-European speaking tribal population groups of Southern Gujarat, India to investigate and reconstruct their paternal population structure and population histories. The role of language, ethnicity and geography in determining the observed pattern of Y haplogroup clustering in the study populations was also examined. A set of 48 bi-allelic markers on the non-recombining region of Y chromosome (NRY) were analysed in 284 males; representing nine Indo-European speaking tribal populations. The genetic structure of the populations revealed that none of these groups was overtly admixed or completely isolated. However, elevated haplogroup diversity and FST value point towards greater diversity and differentiation which suggests the possibility of early demographic expansion of the study groups. The phylogenetic analysis revealed 13 paternal lineages, of which six haplogroups: C5, H1a*, H2, J2, R1a1* and R2 accounted for a major portion of the Y chro)
In this paper, we report our preliminary efforts in building an English-Turkish parallel treebank corpus for statistical machine translation. In the corpus, we manually generated parallel trees for about 5,000 sentences from Penn Treebank. English sentences in our set have a maximum of 15 tokens, including punctuation. We constrained the translated trees to the reordering of the children and the replacement of the leaf nodes with appropriate glosses. We also report the tools that we built and used in our tree translation task.
A fundamental principle of brain organization is bilateral symmetry of structures and functions. For spatial sensory and motor information processing, this organization is generally plausible subserving orientation and coordination of a bilaterally symmetric body. However, breaking of the symmetry principle is often seen for functions that depend on convergent information processing and lateralized output control, e.g. left hemispheric dominance for the linguistic speech system. Conversely, a subtle splitting of functions into hemispheres may occur if peripheral information from symmetric sense organs is partly redundant, e.g. auditory pattern recognition, and therefore allows central conceptualizations of complex stimuli from different feature viewpoints, as demonstrated e.g. for hemispheric analysis of frequency modulations in auditory cortex (AC) of mammals including humans. Here we demonstrate that discrimination learning of rapidly but not of slowly amplitude modulated tones is n)
Automatic prediction of emotions requires reliably annotated data which can be achieved using scoring or pairwise ranking. But can we predict an emotional score using a ranking-based annotation approach? In this paper, we propose to answer this question by describing a regression analysis to map crowdsourced rankings into affective scores in the induced valence-arousal emotional space. This process takes advantages of the Gaussian Processes for regression that can take into account the variance of the ratings and thus the subjectivity of emotions. Regression models successfully learn to fit input data and provide valid predictions. Two distinct experiments were realized using a small subset of the publicly available LIRIS-ACCEDE affective video database for which crowdsourced ranks, as well as affective ratings, are available for arousal and valence. It allows to enrich LIRIS-ACCEDE by providing absolute video ratings for the whole database in addition to video rankings that are already available.
In this paper, we conduct a study about differences between female and male discursive strategies when posting in the microblogging service Twitter, with a particular focus on the hashtag designation process during political debate. The fact that men and women use language in distinct ways, reverberating practices linked to their expected roles in the social groups, is a linguistic phenomenon known to happen in several cultures and that can now be studied on the Web and on online social networks in a large scale enabled by computing power. Here, for instance, after analyzing tweets with political content posted during Brazilian presidential campaign,we found out that male Twitter users, when expressing their attitude toward a given candidate, are more prone to use imperative verbal forms in hashtags, while female users tend to employ declarative forms. This difference can be interpreted as a sign of distinct approaches in relation to other network members: for example, if political ha)
Words are built from smaller meaning bearing parts, called morphemes. As one word can contain multiple morphemes, one morpheme can be present in different words. The number of distinct words a morpheme can be found in is its family size. Here we used Birth-Death-Innovation Models (BDIMs) to analyze the distribution of morpheme family sizes in English and German vocabulary over the last 200 years. Rather than just fitting to a probability distribution, these mechanistic models allow for the direct interpretation of identified parameters. Despite the complexity of language change, we indeed found that a specific variant of this pure stochastic model, the second order linear balanced BDIM, significantly fitted the observed distributions. In this model, birth and death rates are increased for smaller morpheme families. This finding indicates an influence of morpheme family sizes on vocabulary changes. This could be an effect of word formation, perception or both. On a more general level, )
One of the most well studied ecological patterns is Rapoport's rule, which posits that the geographical extent of species ranges increases at higher latitudes. However, studies to date have been limited in their geographic scope and results have been equivocal. In turn, much debate exists over potential links between Rapoport's rule and latitudinal patterns in species richness. Humans collectively speak nearly 7000 different languages, which are spread unevenly across the globe, with loci in the tropics. Causes of this skewed distribution have received only limited study. We analyze the extent of Rapoport's rule in human languages at a global scale and within each region of the globe separately. We test the relationship between Rapoport's rule and the richness of languages spoken in different regions. We also explore the frequency distribution of language-range sizes. The language-range area distribution is strongly right-skewed, with 87% of languages having range areas less than 10,0)
How do children learn to restrict their productivity and avoid ungrammatical utterances? The present study addresses this question by examining why some verbs are used with un- prefixation (e.g., unwrap) and others are not (e.g., *unsqueeze). Experiment 1 used a priming methodology to examine children's (3–4; 5–6) grammatical restrictions on verbal un- prefixation. To elicit production of un-prefixed verbs, test trials were preceded by a prime sentence, which described reversal actions with grammatical un- prefixed verbs (e.g., Marge folded her arms and then she unfolded them). Children then completed target sentences by describing cartoon reversal actions corresponding to (potentially) un- prefixed verbs. The younger age-group's production probability of verbs in un- form was negatively related to the frequency of the target verb in bare form (e.g., squeez/e/ed/es/ing), while the production probability of verbs in un- form for both age groups was negatively predicted by the frequency)
Human languages are rule governed, but almost invariably these rules have exceptions in the form of irregularities. Since rules in language are efficient and productive, the persistence of irregularity is an anomaly. How does irregularity linger in the face of internal (endogenous) and external (exogenous) pressures to conform to a rule? Here we address this problem by taking a detailed look at simple past tense verbs in the Corpus of Historical American English. The data show that the language is open, with many new verbs entering. At the same time, existing verbs might tend to regularize or irregularize as a consequence of internal dynamics, but overall, the amount of irregularity sustained by the language stays roughly constant over time. Despite continuous vocabulary growth, and presumably, an attendant increase in expressive power, there is no corresponding growth in irregularity. We analyze the set of irregulars, showing they may adhere to a set of minority rules, allowing for i)
We present IceMorph, a semi-supervised morphosyntactic analyzer of Old Icelandic. In addition to machine-read corpora and dictionaries, it applies a small set of declension prototypes to map corpus words to dictionary entries. A web-based GUI allows expert users to modify and augment data through an online process. A machine learning module incorporates prototype data, edit-distance metrics, and expert feedback to continuously update part-of-speech and morphosyntactic classification. An advantage of the analyzer is its ability to achieve competitive classification accuracy with minimum training data. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the co)