Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Development of interpersonal relationships is a fundamental human motivation, and behaviors facilitating social bonding are prized. Some individuals experience enhanced reward from alcohol in social contexts and may be at heightened risk for developing and maintaining problematic drinking. We employed a 3 (group beverage condition) ×2 (genotype) design (N = 422) to test the moderating influence of the dopamine D4 receptor gene (DRD4 VNTR) polymorphism on the effects of alcohol on social bonding. A significant gene x environment interaction showed that carriers of at least one copy of the 7-repeat allele reported higher social bonding in the alcohol, relative to placebo or control conditions, whereas alcohol did not affect ratings of 7-absent allele carriers. Carriers of the 7-repeat allele were especially sensitive to alcohol's effects on social bonding. These data converge with other recent gene-environment interaction findings implicating the DRD4 polymorphism in the development of alcohol use disorders, and results suggest a specific pathway by which social factors may increase risk for problematic drinking among 7-repeat carriers. More generally, our findings highlight the potential utility of employing transdisciplinary methods that integrate genetic methodologies, social psychology, and addiction theory to improve theories of alcohol use and abuse.
A morphological analyser only recognizes words that it already knows in the lexical database. It needs, however, a way of sensing significant changes in the language in the form of newly borrowed or coined words with high frequency. We develop a finite-state morphological guesser in a pipelined methodology for extracting unknown words, lemmatizing them, and giving them a priority weight for inclusion in a lexicon. The processing is performed on a large contemporary corpus of 1,089,111,204 words and passed through a machine-learning-based annotation tool. Our method is tested on a manually-annotated gold standard of 1,310 forms and yields good results despite the complexity of the task. Our work shows the usability of a highly non-deterministic finite state guesser in a practical and complex application. 1
In Spanish verbs associated with three participants – Agent, Theme and Recipient – may appear in alternating constructions, where the 3rd person recipient argument is realized as a prepositional phrase (PP) (Pedro envió una carta a María ‘Peter sent a letter to Mary’) or as one doubled by a clitic (Pedro le envió una carta a María ‘Pedro sent Mary a letter’), the latter being referred to as an indirect object (IO). This paper provides a corpus-based study of the distributional patterns of the two constructions that includes both give-type and send-type verbs. The analysis of PPs and IOs in terms of referential properties shows that both have a strong tendency to be [+definite]. However, the distribution of the IO is more constrained than the PPs in terms of certain referential properties, although there are some lexical differences observed among the verbs. The PP, on the other hand, is free of any restrictions. One important contribution of this study is that it provides empirical evidence that the IO associated with the role of Recipient behaves very differently from the one assuming other roles: clitic doubling, which has become the norm for the latter, is still very restricted for the former, contrary to what has been commonly assumed.
Emotional words--as symbols for biologically relevant concepts--are preferentially processed in brain regions including the visual cortex, frontal and parietal regions, and a corticolimbic circuit including the amygdala. Some of the brain structures found in functional magnetic resonance imaging are not readily apparent in electro- and magnetoencephalographic (EEG; MEG) measures. By means of a combined EEG/MEG source localization procedure to fully exploit the available information, we sought to reduce these discrepancies and gain a better understanding of spatiotemporal brain dynamics underlying emotional-word processing. Eighteen participants read high-arousing positive and negative, and low-arousing neutral nouns, while EEG and MEG were recorded simultaneously. Combined current-density reconstructions (L2-minimum norm least squares) for two early emotion-sensitive time intervals, the P1 (80-120 ms) and the early posterior negativity (EPN, 200-300 ms), were computed using realistic individual head models with a cortical constraint. The P1 time window uncovered an emotion effect peaking in the left middle temporal gyrus. In the EPN time window, processing of emotional words was associated with enhanced activity encompassing parietal and occipital areas, and posterior limbic structures. We suggest that lexical access, being underway within 100 ms, is speeded and/or favored for emotional words, possibly on the basis of an "emotional tagging" of the word form during acquisition. This gives rise to their differential processing in the EPN time window. The EPN, as an index of natural selective attention, appears to reflect an elaborate interplay of distributed structures, related to cognitive functions, such as memory, attention, and evaluation of emotional stimuli.
The study of the Tip of the Tongue phenomenon (TOT) provides valuable clues and insights concerning the organisation of the mental lexicon (meaning, number of syllables, relation with other words, etc.). This paper describes a tool based on psycho-linguistic observations concerning the TOT phenomenon. We've built it to enable a speaker/writer to find the word he is looking for, word he may know, but which he is unable to access in time. We try to simulate the TOT phenomenon by creating a situation where the system knows the target word, yet is unable to access it. In order to find the target word we make use of the paradigmatic and syntagmatic associations stored in the linguistic databases. Our experiment allows the following conclusion: a tool like SVETLAN, capable to structure (automatically) a dictionary by domains can be used sucessfully to help the speaker/writer to find the word he is looking for, if it is combined with a database rich in terms of paradigmatic links like EuroWordNet.
As is well known, Los Angeles is home to a large number of Spanish-speaking immigrants from a variety of different countries. Studies on the Spanish spoken in LA County have demonstrated the existence of a distinct dialect, a product of koineization and dialect leveling. In this paper we seek to explore the dynamics of such leveling as it occurs in a public elementary school in this area. Taking into consideration the fact that use of Spanish is not permited in the public school classroom, we choose to observe children’s use of this language during recess, lunch and a structured after school program designed to help Spanish speakers improve their English skills. These observations are further supplemented by answers to informal interviews with students froom the after school program. The evidence that we find provides further support for the validity of LA Spanish as a distinct dialect and serves to illuminate, at least in part, the role that the public school setting plays in the creation of linguistic norms.
The richness of semantic representations associated with individual words has emerged as an important variable in reading. In the present study we contrasted different measures of semantic richness and explored the time course of their influences during visual word processing as reflected in event-related brain potentials (ERPs). ERPs were recorded while participants performed a lexical decision task on visually presented words and pseudowords. For word stimuli, we orthogonally manipulated two frequently employed measures of semantic richness: the number of semantic features generated in feature-listing tasks and the number of associates based on free association norms. We did not find any influence of the number of associates. In contrast, the number of semantic features modulated ERP amplitudes at central sites starting at about 190 ms, as well as during the later N400 component over centro-parietal regions (300-500 ms). Thus, initial access to semantic representations of single words is fast and word meaning continues to modulate processing later on during reading.
"Age of acquisition is possibly the single most potent variable affecting lexical access. It is also a variable that determines the retention or loss of words in patients who have suffered brain injury, and in patients with Alzheimer´s disease. But the norms of age of acquisition currently available have largely been obtained from university students whereas the ages of acquisition for some words are very different for young people compared with the elderly. The aim of this study was to develop age of acquisition norms for a sample of 500 words with people over 60 years. When these norms were compared with others from young people in predicting the results of a group of Alzheimer patients in a lexical selection task we found that the elderly ratings made a better prediction of the data. We recommend that for studies using older participants appropriate norms should be used in place of those obtained from young adults."
The annotations of explicit and implicit discourse connectives in the Penn Discourse Treebank make it possible to investigate on a large scale how different types of discourse relations are expressed. Assuming an account of the Uniform Information Density hypothesis, we expect that discourse relations should be expressed explicitly with a discourse connector when they are unexpected, but may be implicit when the discourse relation can be anticipated. We investigate whether discourse relations which have been argued to be expected by the comprehender exhibit a higher ratio of implicit connectors. We find support for two hypotheses put forth in previous research which suggest that continuous and causal relations are presupposed by language users when processing consecutive sentences in a text. We then proceed to analyze the effect of Implicit Causality (IC) verbs (which have been argued to raise an expectation for an explanation) as a local cue for an upcoming causal relation.
Palmerston Island is a tiny isolated community in the Pacific. Over the past 140 years it has developed a unique linguistic and cultural identity, influenced by England, the Cook Islands, and more recently New Zealand. The islanders strongly identify with England and consider themselves very different from the rest of the Cook Islands, to which Palmerston Island officially belongs. This paper explores the relationship between Palmerston Islanders� conceptions of themselves and their linguistic ideologies. It is shown that the construction of linguistic and social norms is not entirely subconscious: the community is aware of the different origins of lexical items, and the cultural and social affiliations signalled by different linguistic choices. Subconscious co-evolution of culture and language also takes place and appears likely to be responsible for the substrate influences of Cook Island M�ori in both realms.
Infants begin to segment novel words from speech by 7.5 months, demonstrating an ability to track, encode and retrieve words in the context of larger units. Although it is presumed that word recognition at this stage is a prerequisite to constructing a vocabulary, the continuity between these stages of development has not yet been empirically demonstrated. The goal of the present study is to investigate whether infant word segmentation skills are indeed related to later lexical development. Two word segmentation tasks, varying in complexity, were administered in infancy and related to childhood outcome measures. Outcome measures consisted of age-normed productive vocabulary percentiles and a measure of cognitive development. Results demonstrated a strong degree of association between infant word segmentation abilities at 7 months and productive vocabulary size at 24 months. In addition, outcome groups, as defined by median vocabulary size and growth trajectories at 24 months, showed distinct word segmentation abilities as infants. These findings provide the first prospective evidence supporting the predictive validity of infant word segmentation tasks and suggest that they are indeed associated with mature word knowledge. A video abstract of this article can be viewed at http://www.youtube.com/watch?v=jxzLi5oLZQ8.
This paper focuses on pronominal order in verbal periphrases, taking into consideration data from a research developed by Peterson (2010) in which clitic pronouns produced by newspaper readers from Rio de Janeiro city have been analysed. Developed within the framework of Labovian Sociolinguistics, the study demonstrates that although the readers’ correspondences present the alternant form V1 cl V2 (pode me dizer), highly productive and natural in Brazilian Portuguese, different uses can be observed in relation to proximity or distance to idealized linguistic norms which are proposed by prescriptive manuals. Not only linguistic motivations, but also extralinguistic restrictions (regarding mainly newspapers’ characteristics) have been associated to the pronominal order.
The objective of Web services technology is to facilitate the creation of reusable and network accessible business applications over the Web for automatic discovery and compositions. Automatic discovery of Web Services can be achieved by describing Web service functionality in a richer (in terms of context and semantics) and machine readable form. With a rapid growth in Web services, finding suitable services for the requesters has become a challenging task. In this paper, the authors present an effective Web service discovery mechanism which exploits the strengths of functional semantics, keyword matching, structural matching, syntactic matching and semantic matching in order to explore the relevant Web services which satisfy the requester's functional needs. The proposed discovery mechanism also makes use of the Web service crawler which explores the Web services accessible over the Web and the WordNet lexical database to obtain synonyms of the functional semantic enriched requester's query. The proposed discovery mechanism is implemented and the experimentation proves the effectiveness in terms of Recall and Precision.
Here we describe work on learning the subcategories of verbs in a morphologically rich language using only minimal linguistic resources. Our goal is to learn verb subcategorizations for Quechua, an under-resourced morphologically rich language, from an unannotated corpus. We compare results from applying this approach to an unannotated Arabic corpus with those achieved by processing the same text in treebank form. The original plan was to use only a morphological analyzer and an unannotated corpus, but experiments suggest that this approach by itself will not be effective for learning the combinatorial potential of Arabic verbs in general. The lower bound on resources for acquiring this information is somewhat higher, apparently requiring a a part-of-speech tagger and chunker for most languages, and a morphological disambiguater for Arabic.
Abstract Since the late 1800s, the Uruguayan Government has attempted to enforce cultural and linguistic norms along the border with Brazil through the prohibition of Portuguese, especially in schools, despite the fact that this is the heritage language of most border residents. This research focuses on the differential use of Spanish and Portuguese in Rivera, the largest city on the border. Using self-reported data and metalinguistic commentaries extracted from interviews with 63 Spanish–Portuguese bilinguals, the use of both languages in various domains (home, school, work spaces) and with diverse interlocutors (family, friends, co-workers, superiors) is analyzed. Quantitative and qualitative analysis reveals that Portuguese, which has been marginalized for decades, is more frequently used in the home with relatives and close friends. The use of Portuguese in more formal domains, including schools, is much less frequent. The results from this study corroborate a perception within the community that Portuguese lacks the prestige of Spanish and provide further evidence of its status as a primarily home language. The current research does not show a progressive shift toward Spanish in Rivera nor does it support claims by other researchers that this community is diglossic. Keywords: SpanishPortuguese portuñol language contactlanguage usebilingual education Acknowledgements This research would not have been possible without the generous support of a Tinker Foundation Grant from the Latin American and Iberian Institute of the University of New Mexico, which allowed me to conduct fieldwork in Rivera. I am also extremely grateful for the advice given to me by Ana Maria Carvalho, Adolfo Elizaincín, and Rena Torres-Cacoullos. Thank you also to the people of Rivera, who welcomed me into their community with open arms. Notes 1. Definitions of bilingualism should not be based on language use alone, however, since attitudes toward both languages also shape a bilingual's identity (Ben-Rafael, Olshtain, and Geijst Citation1998; Hoare Citation2001; Joseph Citation2004; Lawson and Sachdev Citation2004). In other words, bilingualism is not merely a linguistic phenomenon involving frequency of use, but rather a social one as well, which is defined largely by its role within the community. Consequently, these attitudes affect the bilingual's choice of one language or another. 2. Five of six of these speakers are of the second generation, while only one is of the third generation (speaker 15, female, professional occupation). The social attributes of each consultant, however, vary. Of the second generation, two males (one with a professional job and one with a nonprofessional job) did not complete the questionnaire. Likewise, the three women who did not complete the questionnaire cover both occupational classes (two professionals and one nonprofessional), thereby maintaining the diversity of the original composition of social characteristics for this generation. 3. Nonprofessionals differ from professionals in that the work they perform does not require formal academic training. These members of the community are taxi drivers, shopkeepers and their employees, hotel owners and their employees, waiters, bartenders, construction workers, etc. 4. Some of the younger consultants indicated percentages of language use with a spouse, which most clarified by writing in the margin novio or novia 'boyfriend' or 'girlfriend.' These percentages were not included in the rate calculations presented in Table 4 since these interlocutors do not belong strictly to home domains.
In this thesis, an implementation of the MAT algorithm for query learning is improved and tested, specifically by attempted simulation of the MAT oracle using large corpora. In order to arrive at suitable test corpora, algorithms for generating positive and negative examples in relation to regular tree languages are developed.
As little is known about the effectiveness of different types of implementation intentions on the regulation of emotions, the present experiments focused on the differential effectiveness of various implementation intentions on the down-regulation of disgust responses. In Experiment 1, an antecedent-focused implementation intention based on cognitive reappraisal allowed participants to rate disgusting pictures as being less unpleasant than participants in the control condition or the goal intention condition, while the reported intensity (arousal) ratings stayed unaffected. In Experiment 2, participants with a response-focused implementation intention, devised to regulate the intensity of the emotional experience, reported a lower evoked arousal after seeing the disgusting slides, while the valence ratings remained unchanged. Thus, implementation intentions were shown to exert differential effects depending on whether they targeted one or another emotional dimension (i.e., valence vs. arousal).
An important goal of addiction research and treatment is to predict behavioural responses to drug-related stimuli. This goal is especially important for patients with impaired insight, which can interfere with therapeutic interventions and potentially invalidate self-report questionnaires. This research tested (i) whether event-related potentials, specifically the late positive potential, predict choice to view cocaine images in cocaine addiction; and (ii) whether such behaviour prediction differs by insight (operationalized in this study as self-awareness of image choice). Fifty-nine cocaine abusers and 32 healthy controls provided data for the following laboratory components that were completed in a fixed-sequence (to establish prediction): (i) event-related potential recordings while passively viewing pleasant, unpleasant, neutral and cocaine images, during which early (400-1000 ms) and late (1000-2000 ms) window late positive potentials were collected; (ii) self-reported arousal ratings for each picture; and (iii) two previously validated tasks: one to assess choice for viewing these same images, and the other to group cocaine abusers by insight. Results showed that pleasant-related late positive potentials and arousal ratings predicted pleasant choice (the choice to view pleasant pictures) in all subjects, validating the method. In the cocaine abusers, the predictive ability of the late positive potentials and arousal ratings depended on insight. Cocaine-related late positive potentials better predicted cocaine image choice in cocaine abusers with impaired insight. Another emotion-relevant event-related potential component (the early posterior negativity) did not show these results, indicating specificity of the late positive potential. In contrast, arousal ratings better predicted respective cocaine image choice (and actual cocaine use severity) in cocaine abusers with intact insight. Taken together, the late positive potential could serve as a biomarker to help predict drug-related choice--and possibly associated behaviours (e.g. drug seeking in natural settings, relapse after treatment)--when insight (and self-report) is compromised.
Morphological segmentation data for the METU-Sabanci Turkish Treebank is provided in this paper. The generalized lexical forms of the morphemes which the treebank previously lacked are added to the treebank. This data maybe used to train POS-taggers that use stemmer outputs to map these lexical forms to morphological tags.
Treebank is a basic language resource for training and testing syntactic parser which forms a key module in various NLP systems like machine translation system. This paper reports an ongoing research of building dependency treebank for Kashmiri (KashTreeBank) and discusses some main annotation issues. The paper is based on the pilot annotation of 500 sentences.
This paper presents our work for participation in the 2012 CIPS-SIGHAN shared task of Traditional Chinese Parsing. We have adopted two multilingual parsing models – a factored model (Stanford Parser) and an unlexicalized model (Berkeley Parser) for parsing the Sinica Treebank. This paper also proposes a new Chinese unknown word model and integrates it into the Berkeley Parser. Our experiment gives the first result of adapting existing multilingual parsing models to the Sinica Treebank and shows that the parsing accuracy can be improved by our suggested approach.
The recent construction of large linguistic treebanks for spoken and written Dutch (e.g. CGN, LASSY, Alpino) has created new and exciting opportunities for the empirical investigation of Dutch syntax and semantics. However, the exploitation of those treebanks requires knowledge of specific data structures and query languages such as XPath. Linguists who are unfamiliar with formal languages are often reluctant towards learning such a language. In order to make treebank querying more attractive for non-technical users we developed GrETEL (Greedy Extraction of Trees for Empirical Linguistics), a query engine in which linguists can use natural language examples as a starting point for searching the Lassy treebank without knowledge about tree representations nor formal query languages. By allowing linguists to search for similar constructions as the example they provide, we hope to bridge the gap between traditional and computational linguistics. Two case studies are conducted to provide a concrete demonstration of the tool. The architecture of the tool is optimised for searching the LASSY treebank, but the approach can be adapted to other treebank lay-outs.
Après un bref rĂŠsumĂŠ de la thĂŠorie de Topic-Focus Articulation (TFA), la prĂŠsente ĂŠtude dĂŠmontre, à l'aide de plusieurs exemples illustrant l'annotation de principaux traits de TFA sur un large corpus (the Prague Dependency Treebank), que l'annotation du corpus apporte une valeur ajoutĂŠe au corpus, si deux conditions sont rĂŠunies: (i) le schĂŠma de l'annotation est basĂŠ sur une thĂŠorie linguistique solide, (ii) le procĂŠdĂŠ d'annotation est ĂŠtabli avec soin (c'est-à-dire de façon systĂŠmatique et cohĂŠrente). Une telle annotation est importante non seulement pour la structure de surface de la phrase mais encore davantage pour la structure phrastique sous-jacente, car elle est susceptible de mettre en ĂŠvidence les phĂŠnomènes cachĂŠs au niveau de la structure de surface, mais incontournables lors de la reprĂŠsentation du sens et du fonctionnement de la phrase.
Among various neural network language models (NNLMs), recurrent neural network-based language models (RNNLMs) are very competitive in many cases. Most current RNNLMs only use one single feature stream, i.e., surface words. However, previous studies proved that language models with additional linguistic information achieve better performance. In this study, we extend RNNLM by explicitly integrating additional linguistic information, including morphological, syntactic, or semantic factors. Our proposed RNNLM is called a factored RNNLM that is expected to enhance RNNLMs. A number of experiments are carried out that show the factored RNNLM improves the performance for all considered tasks: consistent perplexity and word error rate (WER) reductions. In the Penn Treebank corpus, the relative improvements over n-gram LM and RNNLM are 29.0 % and 13.0%, respectively. In the IWSLT-2011 TED ASR test set, absolute WER reductions over RNNLM and n-gram LM reach 0.63 and 0.73 points. Title and Abstract in another language (Chinese) ddddddddddddddd ddddddddddddNNLMddddddddddddddRNNLMddd ddddddddddddddddddddRNNLMddddddddddddddd ddddddddddddddddddddddddddddddddddddddd ddddddddddddddddddRNNLMddddddddddddddddd dddddddddfRNNLMddddddddddddddddfRNNLMdddddd dddddddRNNLMdddddddddddddddddddddWERddddd ddddfRNNLMdddddddddnddddddRNNLMddddd29.0%d13.0%d dIWSLT-2011 TEDdddddddddfRNNLMddddddddddd0.63ddd dnddddddd0.73ddddRNNLMdd
Domain adaptation is an important topic for natural language processing. There has been extensive research on the topic and various methods have been explored, including training data selection, model combination, semi-supervised learning. In this study, we propose to use a goodness measure, namely, description length gain (DLG), for domain adaptation for Chinese word segmentation. We demonstrate that DLG can help domain adaptation in two ways: as additional features for supervised segmenters to improve system performance, and also as a similarity measure for selecting training data to better match a test set. We evaluated our systems on the Chinese Penn Treebank version 7.0, which has 1.2 million words from five different genres, and the Chinese Word Segmentation Bakeoff-3 data.
We examined whether language affects the strength of a visual representation in memory. Participants studied a picture, read a story about the depicted object, and then selected out of two pictures the one whose transparency level most resembled that of the previously presented picture. The stories contained two linguistic manipulations that have been demonstrated to affect concept availability in memory, i.e., object presence and goal-relevance. The results show that described absence of an object caused people to select the most transparent picture more often than described presence of the object. This effect was not moderated by goal-relevance, suggesting that our paradigm tapped into the perceptual quality of representations rather than, for example, their linguistic availability. We discuss the implications of these findings within a framework of grounded cognition. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not )
This study aims at investigating the HLA molecular variation across Switzerland in order to determine possible regional differences, which would be highly relevant to several purposes: optimizing donor recruitment strategies in hematopoietic stem cell transplantation (HSCT), providing reliable reference data in HLA and disease association studies, and understanding the population genetic background(s) of this culturally heterogeneous country. HLA molecular data of more than 20,000 HSCT donors from 9-13 recruitment centers of the whole country were analyzed. Allele and haplotype frequencies were estimated by using new computer tools adapted to the heterogeneity and ambiguity of the data. Nonparametric and resampling statistical tests were performed to assess Hardy-Weinberg equilibrium, selective neutrality and linkage disequilibrium among different loci, both in each recruitment center and in the whole national registry. Genetic variation was explored through genetic distance and hiera)
Many patterns displayed by the distribution of human linguistic groups are similar to the ecological organization described for biological species. It remains a challenge to identify simple and meaningful processes that describe these patterns. The population size distribution of human linguistic groups, for example, is well fitted by a log-normal distribution that may arise from stochastic demographic processes. As we show in this contribution, the distribution of the area size of home ranges of those groups also agrees with a log-normal function. Further, size and area are significantly correlated: the number of speakers p and the area a spanned by linguistic groups follow the allometric relation a ... pz, with an exponent z varying accross different world regions. The empirical evidence presented leads to the hypothesis that the distributions of p and a, and their mutual dependence, rely on demographic dynamics and on the result of conflicts over territory due to group growth. To s)
The small alpine district of East Tyrol (Austria) has an exceptional demographic history. It was contemporaneously inhabited by members of the Romance, the Slavic and the Germanic language groups for centuries. Since the Late Middle Ages, however, the population of the principally agrarian-oriented area is solely Germanic speaking. Historic facts about East Tyrol's colonization are rare, but spatial density-distribution analysis based on the etymology of place-names has facilitated accurate spatial mapping of the various language groups' former settlement regions. To test for present-day Y chromosome population substructure, molecular genetic data were compared to the information attained by the linguistic analysis of pasture names. The linguistic data were used for subdividing East Tyrol into two regions of former Romance (A) and Slavic (B) settlement. Samples from 270 East Tyrolean men were genotyped for 17 Y-chromosomal microsatellites (Y-STRs) and 27 single nucleotide polymorphism)
This study introduced a novel algorithm to compute similarities between natural languages. Using syntactical relationships derived from natural languages, the algorithm proposed a semantic structural model and quantified natural languages using the word similarity method based on WordNet and lexical databases. The experimental results indicated that the algorithm could yield optimal results in semantic recognition when applied to sentences or short texts that are grammatically complex or relatively long (longer than 12 words). The contribution of this study is in its conversion of the grammar of different natural languages into a unified semantic structure, through which the semantic similarity of two sentences or short texts can be obtained by comparison. This study aimed to enhance the capability of computers for fuzzy concept processing, which can be applied to the fields of search engines and artificial intelligence. For instance, in search engines, sentences or short text-based concepts may be semantically structured to replace key-word based queries when executing search tasks. In the field of artificial intelligence, this capability may be applied to intelligent agents to smooth the process of interaction between humans and machines.
Background: Recent advances in automated assessment of basic vocabulary lists allow the construction of linguistic phylogenies useful for tracing dynamics of human population expansions, reconstructing ancestral cultures, and modeling transition rates of cultural traits over time. Methods: Here we investigate the Tupi expansion, a widely-dispersed language family in lowland South America, with a distance-based phylogeny based on 40-word vocabulary lists from 48 languages. We coded 11 cultural traits across the diverse Tupi family including traditional warfare patterns, post-marital residence, corporate structure, community size, paternity beliefs, sibling terminology, presence of canoes, tattooing, shamanism, men's houses, and lip plugs. Results/Discussion: The linguistic phylogeny supports a Tupi homeland in west-central Brazil with subsequent major expansions across much of lowland South America. Consistently, ancestral reconstructions of cultural traits over the linguistic phylogen)
Recently, Kuhlmann (2007, Dependency Structures and Lexicalized Grammars. PhD Thesis, Saarland University) and collaborators have shown how the derivations of generative grammars can be recast as dependency structures. This connection between the generative and dependency traditions opens the door to a fresh perspective on how to formally characterize natural language and what minimal machinery can cover such data. This article draws on both reported properties of structures in dependency treebanks and properties of informant data to determine the complexity of natural language along two dependency measures, gap degree (a measure of discontinuity) and well- versus ill-nestedness (whether interleaving substructures are permitted). We show that natural language includes constructions that require dependency analyses that are ill-nested and/or gap degree > 1, and argue that a grammar formalism on the right track for characterizing natural language should be able to generate such structures. We investigate the adequacy of tree-local multi-component tree adjoining grammar (TL-MCTAG) to cover existent data, examining the relationship between TL-MCTAG derivations and dependency representations. Though focused on TL-MCTAG, this work also advances the larger enterprise of discovering mathematically defined formal systems and testing their adequacies by using both linguistic judgments as well as large bodies of data from annotated corpora.
Findings on song perception and song production have increasingly suggested that common but partially distinct neural networks exist for processing lyrics and melody. However, the neural substrates of song recognition remain to be investigated. The purpose of this study was to examine the neural substrates involved in the accessing "song lexicon" as corresponding to a representational system that might provide links between the musical and phonological lexicons using positron emission tomography (PET). We exposed participants to auditory stimuli consisting of familiar and unfamiliar songs presented in three ways: sung lyrics (song), sung lyrics on a single pitch (lyrics), and the sung syllable 'la' on original pitches (melody). The auditory stimuli were designed to have equivalent familiarity to participants, and they were recorded at exactly the same tempo. Eleven right-handed nonmusicians participated in four conditions: three familiarity decision tasks using song, lyrics, and melod)
Orthographies vary in the degree of transparency of spelling-sound correspondence. These range from shallow orthographies with transparent grapheme-phoneme relations, to deep orthographies, in which these relations are opaque. Only a few studies have examined whether orthographic depth is reflected in brain activity. In these studies a betweenlanguage design was applied, making it difficult to isolate the aspect of orthographic depth. In the present work this question was examined using a within-subject-and-language investigation. The participants were speakers of Hebrew, as they are skilled in reading two forms of script transcribing the same oral language. One form is the shallow pointed script (with diacritics), and the other is the deep unpointed script (without diacritics). Event-related potentials (ERPs) were recorded while skilled readers carried out a lexical decision task in the two forms of script. A visual non-orthographic task controlled for the visual difference between t)
The paper offers an overview of the key issues raised during the 8 years’ activity of the Multilingual Question Answering Track at the Cross Language Evaluation Forum (CLEF). The general aim of the track has been to test both monolingual and cross-language Question Answering (QA) systems that process queries and documents in several European languages, also drawing attention to a number of challenging issues for research in multilingual QA. The paper gives a brief description of how the task has evolved over the years and of the way in which the data sets have been created, presenting also a short summary of the different types of questions developed. The document collections adopted in the competitions are outlined as well, and data about participation is provided. Moreover, the main measures used to evaluate system performances are explained and an overall analysis of the results achieved is presented.
Background: Continuity of care is widely acknowledged as a core value in family medicine. In this systematic review, we aimed to identify the instruments measuring continuity of care and to assess the quality of their measurement properties. Methods: We did a systematic review using the PubMed, Embase and PsycINFO databases, with an extensive search strategy including 'continuity of care', 'coordination of care', 'integration of care', 'patient centered care', 'case management' and its linguistic variations. We searched from 1995 to October 2011 and included articles describing the development and/ or evaluation of the measurement properties of instruments measuring one or more dimensions of continuity of care (1) care from the same provider who knows and follows the patient (personal continuity), (2) communication and cooperation between care providers in one care setting (team continuity), and (3) communication and cooperation between care providers in different care settings (cross)
Corpora with high-quality linguistic annotations are an essential component in many NLP applications and a valuable resource for linguistic research. For obtaining these annotations, a large amount of manual effort is needed, making the creation of these resources time-consuming and costly. One attempt to speed up the annotation process is to use supervised machine-learning systems to automatically assign (possibly erroneous) labels to the data and ask human annotators to correct them where necessary. However, it is not clear to what extent these automatic pre-annotations are successful in reducing human annotation effort, and what impact they have on the quality of the resulting resource. In this article, we present the results of an experiment in which we assess the usefulness of partial semi-automatic annotation for frame labeling. We investigate the impact of automatic pre-annotation of differing quality on annotation time, consistency and accuracy. While we found no conclusive evidence that it can speed up human annotation, we found that automatic pre-annotation does increase its overall quality.
The rat visual system is structured such that the large (>90 %) majority of retinal ganglion axons reach the contralateral lateral geniculate nucleus (LGN) and visual cortex (V1). This anatomical design allows for the relatively selective activation of one cerebral hemisphere under monocular viewing conditions. Here, we describe the design of a harness and face mask allowing simple and noninvasive monocular occlusion in rats. The harness is constructed from synthetic fiber (shoelace-type material) and fits around the girth region and neck, allowing for easy adjustments to fit rats of various weights. The face mask consists of soft rubber material that is attached to the harness by Velcro strips. Eyeholes in the mask can be covered by additional Velcro patches to occlude either one or both eyes. Rats readily adapt to wearing the device, allowing behavioral testing under different types of viewing conditions. We show that rats successfully acquire a water-maze-based visual discrimination task under monocular viewing conditions. Following task acquisition, interocular transfer was assessed. Performance with the previously occluded, “untrained” eye was impaired, suggesting that training effects were partially confined to one cerebral hemisphere. The method described herein provides a simple and noninvasive means to restrict visual input for studies of visual processing and learning in various rodent species.
This paper is to analyze curricular changes of Chongryon Korean schools in Japan. Chongryon Korean schools belong to the category of miscellaneous schools in the classification by the Ministry of Education in Japan. They neither need to follow curricular set by the Japanese government nor receive subsidies from it. They have their own curricula and textbooks. Historically, there have been 6 curricular reforms in Chongryon Korean schools. After characterizing those reforms, this paper compares the ``Korean`` textbooks of the 1993 and 2003 reforms. In general, the political color of advocating socialism and the Juche idea and anti-American and anti-Seoul propoganda gets thinner in new textbooks. The 2003 Korean textbooks emphasize speaking practice and adapt dialogues rich with story-telling. This paper also examines linguistic norms followed in the Korean textbooks. (Osaka University of Economics and Law)
The paper introduces an ongoing project for the development of a parallel treebank for Italian, English and French annotated in the pure dependency format of the Turin University Treebank, i.e. Parallel–TUT. We hypothesize that the major features of this annotation format can be of some help in addressing the typical issues related to parallel corpora, e.g. alignment at various levels. Therefore, benefitting from the tools previously used for TUT, we applied the TUT format to a multilingual sample set of sentences from the JRC-
Statistical machine translation (SMT) is based on alignment models which learn from bilingual corpora the word correspondences between source and target language. These models are assumed to be capable of learning reorderings. However, the difference in word order between two languages is one of the most important sources of errors in SMT. In this paper, we show that SMT can take advantage of inductive learning in order to solve reordering problems. Given a word alignment, we identify those pairs of consecutive source blocks (sequences of words) whose translation is swapped, i.e. those blocks which, if swapped, generate a correct monotonic translation. Afterwards, we classify these pairs into groups, following recursively a co-occurrence block criterion, in order to infer reorderings. Inside the same group, we allow new internal combination in order to generalize the reorder to unseen pairs of blocks. Then, we identify the pairs of blocks in the source corpora (both training and test) which belong to the same group. We swap them and we use the modified source training corpora to realign and to build the final translation system. We have evaluated our reordering approach both in alignment and translation quality. In addition, we have used two state-of-the-art SMT systems: a Phrased-based and an Ngram-based. Experiments are reported on the EuroParl task, showing improvements almost over 1 point in the standard MT evaluation metrics (mWER and BLEU).
Cross-language plagiarism detection deals with the automatic identification and extraction of plagiarism in a multilingual setting. In this setting, a suspicious document is given, and the task is to retrieve all sections from the document that originate from a large, multilingual document collection. Our contributions in this field are as follows: (1) a comprehensive retrieval process for cross-language plagiarism detection is introduced, highlighting the differences to monolingual plagiarism detection, (2) state-of-the-art solutions for two important subtasks are reviewed, (3) retrieval models for the assessment of cross-language similarity are surveyed, and, (4) the three models CL-CNG, CL-ESA and CL-ASA are compared. Our evaluation is of realistic scale: it relies on 120,000 test documents which are selected from the corpora JRC-Acquis and Wikipedia, so that for each test document highly similar documents are available in all of the six languages English, German, Spanish, French, Dutch, and Polish. The models are employed in a series of ranking tasks, and more than 100 million similarities are computed with each model. The results of our evaluation indicate that CL-CNG, despite its simple approach, is the best choice to rank and compare texts across languages if they are syntactically related. CL-ESA almost matches the performance of CL-CNG, but on arbitrary pairs of languages. CL-ASA works best on “exact” translations but does not generalize well.
Research in automatic text plagiarism detection focuses on algorithms that compare suspicious documents against a collection of reference documents. Recent approaches perform well in identifying copied or modified foreign sections, but they assume a closed world where a reference collection is given. This article investigates the question whether plagiarism can be detected by a computer program if no reference can be provided, e.g., if the foreign sections stem from a book that is not available in digital form. We call this problem class intrinsic plagiarism analysis; it is closely related to the problem of authorship verification. Our contributions are threefold. (1) We organize the algorithmic building blocks for intrinsic plagiarism analysis and authorship verification and survey the state of the art. (2) We show how the meta learning approach of Koppel and Schler, termed “unmasking”, can be employed to post-process unreliable stylometric analysis results. (3) We operationalize and evaluate an analysis chain that combines document chunking, style model computation, one-class classification, and meta learning.
We present a novel way of extracting a categorial grammar from annotated data. Using the sentences from the Paris VII annotated treebank [2] as our starting point, we use a tree transducer to convert the annotated trees from the corpus into categorial grammar derivations.We describe both the formal aspects and the implementation of the tree transducer, which is a conservative extension of standard tree transducers allowing a compact specification of the transductions rules relevant for our purposes, and we discuss the specific set of transduction rules we use to convert the corpus into AB grammar derivation trees.Evaluating the resulting tree transducer on the entire corpus, we find that it produces a treebank finds lexical entries for 90,0% of the corpus, though it produces complete derivations for only 75% of all sentence in the corpus.
Reviewed by: From poets to padonki: Linguistic authority and norm negotiation in modern Russian culture Anastassia Zabrodskaja Ingunn Lunde and Martin Paulsen, eds. From poets to padonki: Linguistic authority and norm negotiation in modern Russian culture. Bergen: University of Bergen, 2009. [Slavica Bergensia, 9.] The issues discussed in From poets to padonki became a part of my life in 1999 when I began my university studies and using Russian and Estonian became my everyday reality. Choosing different languages for different purposes, code-switching, and multilingual language play have all been part of my daily language use for the last eleven years. This collection challenges definitions of what can be meant by "language", "standard" language, and the linguistic "norm". It acquaints readers with a wide range of linguistic phenomena in modern Russian culture. Padonki refers to a subculture within the Russian-speaking Internet (Runet), whose representatives use erratic spellings for words, aiming at creating a comic effect. Their nickname padonki itself illustrates such a trend-it is an alteration of podonki 'dregs'. Padonki is also characterized by gratuitous use of profanity and a penchant for obscene subjects. During the last decade, a body of literature has emerged proposing that (socio)linguists direct their attention away from the traditional focus of linguistics, i.e., language as a bounded system, towards broader semiotic resources, to see what is really going on when people use "language" (Stroud 2003, Jacquemet 2005, Shohamy 2006, Makoni and Pennycook 2007, Blommaert 2010). The notion of "language" becomes especially questionable in cases of (multilingual) computer-mediated communication. The last decade has also witnessed rising scholarly interest in language on and of the Internet in general and in e-mails and postings on Internet discussion forums or message boards in particular (e.g., Koutsogiannis and Mitsikopoulou 2003, Palfreyman and al Khalil 2003, Hinrichs 2006, Dorleijn and Nortier 2009, Androutsopoulos 2006, 2009). García (2009: 32) instead of "language" offers a more suitable term for the multiple discursive practices-languaging, i.e., "social practices that are actions performed by our meaning-making selves." For her, dialects, pidgins, creoles, and academic language [End Page 153] are all examples of languaging, as there are differences between language practices at home, in communities, and in academic contexts. I would argue that standard languages are idealized constructs, and none can remain unaffected by language contacts during its entire history. While studying language use by individuals, it is important to shift "from focus on structure to focus on function-from focus on linguistic form in isolation to linguistic form in human context" (Hymes 1974: 77). The volume under review offers fascinating reading for (socio)linguists, who work with different manifestations of real language practices rather than seeking for linguistic norm descriptions. Comprising 17 contributions written in English and Russian, the book summarizes analyses of the norm in modern Russian language culture based on data from multiple sources: "literary fiction, Internet slang, literary criticism and aesthetics, writers' blogs, linguistic play, and various arenas for 'talk about talk,' such as the classroom, blogs, the media, the courtroom, etc." (11). The book opens with Ingunn Lunde and Martin Paulsen's excellent general introduction, an insightful synthesis of how various approaches to the standard Russian language, its norms, and linguistic standards see the history, development, and future of the Russian language culture. The first paper, "Living norms" by Henning Andersen, gives a comprehensive overview of the history and development of the notion of language norms. Andersen reviews historical contributions to the understanding of norms. Making a distinction between declarative and deontic norms and dividing them further into explicit and implicit norms, he highlights the leading ideas in the field of (socio)linguistic studies that concern the position of norms. On the example of excerpts from Soviet grammars and dictionaries, he illustrates how norms can be governed from above by agencies of the state. "The standard norms include only prescribed and permitted forms", and all other forms must be avoided (25). Andersen also comments on spoken language standards, noting that "the [Russian] language continues to be spoken effectively in its numerous variants all over the inherited Russian language area as well as in the diasporas, old and new" (32). Martin Paulsen...
Music processing may be preserved in subjects with Alzheimer disease (AD). It is not known which neural substrates are engaged in music processing, and how music familiarity moderates the engagement of these substrates in AD. We investigated fMRI patterns of brain activation during listening to familiar and non-familiar classical music excerpts in subjects with mild to moderate AD and healthy age-matched controls. We related these patterns to behavioral data on familiarity ratings and musical abilities. Five subjects with AD (age M = 76.2, SD = 6.6, MMSE M = 18.6, SD = 7.7, range 9-26) and five healthy controls (age M72.8, SD = 8.0, MMSE M = 29.6, SD =.6, range 29-30) underwent fMRI with a block design paradigm consisting of 75s of familiar music excerpts followed by 30s white noise, vs. 75s of unfamiliar music excerpts. Participants were instructed to just listen. They received a battery of behavioral tests including repeated familiarity ratings of the presented music excerpts (1=very familiar to 5=very unfamiliar), the Montreal Battery for Evaluation of Amusia (MBEA), and the Seashore test of musical abilities. For fMRI a mixed-model 2x2 ANOVA was used to examine effect of group AD vs. controls) and stimulus (familiar vs. unfamiliar) using a p-value <.01 and a minimum cluster size of 200ÂμL. For behavioural data, independent- and paired-sample t-tests were used with p-value <.05. We found a significant group by stimulus effect in fMRI activation patterns. When familiar to unfamiliar music activations were compared, the following regions showed increases in AD and decreases in controls: left/right lingual gyrus, left inferior parietal, left superior, middle and inferior temporal gyri, left precuneus, left culmen, left/right striatum. Ratings for familiar and unfamiliar excerpts did not differ by group (AD M = 1.2, SD =.2 and M = 1.9, SD =.6; Controls M = 1.2, SD =.3, and M = 1.9, SD =.3). Performance on music ability tests also did not differ by group except for Seashore loudness and rhythm (p <.05). Subjects with AD appear to respond more intensely to non-familiar than familiar music by activating regions associated with recognizing familiar patterns and emotions. They do not differ from controls on behavioral measures. This finding suggests differential neural recruitment with respect to preserved music recognition, and its potential application in diagnosis and treatment.
We describe a new interactive annotation scheme between a human annotator who carries out simplified annotations on CFG trees, and a statistical parser that converts the human annotations automatically into a richly annotated HPSG treebank. In order to check the proposed scheme's effectiveness, we performed automatic pseudo-annotations that emulate the system's idealized behavior and measured the performance of the parser trained on those annotations. In addition, we implemented a prototype system and conducted manual annotation experiments on a small test set.
The study is focused on how to make use of the lexical database Pralex for a processing of nominal entries, how to describe a particular specific phenomenon or a partial issue in an entry form as well as which nominal data are included (mandatory items). To clarify this, the authors follow the order of individual parts of an entry form. Special attention is paid to the issue of homonymy, variation, explanation of a meaning and division of polysemantic entries.