Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
From an evolutionary perspective, environmental threats relevant for survival constantly challenged human beings. Current research suggests the evolution of a fear processing module in the brain to cope with these threats. Recently, humans increasingly encountered modern threats (e.g., guns or car accidents) in addition to evolutionary threats (e.g., snakes or predators) which presumably required an adaptation of perception and behavior. However, the neural processes underlying the perception of these different threats remain to be elucidated. We investigated the effect of image content (i.e., evolutionary vs. modern threats) on the activation of neural networks of emotion processing. During functional magnetic resonance imaging (fMRI) 41 participants watched affective pictures displaying evolutionary-threatening, modern-threatening, evolutionary-neutral and modern-neutral content. Evolutionary-threatening stimuli evoked stronger activations than modern-threatening stimuli in left inferior frontal gyrus and thalamus, right middle frontal gyrus and parietal regions as well as bilaterally in parietal regions, fusiform gyrus and bilateral amygdala. We observed the opposite effect, i.e., higher activity for modern-threatening than for evolutionary-threatening stimuli, bilaterally in the posterior cingulate and the parahippocampal gyrus. We found no differences in subjective arousal ratings between the two threatening conditions. On the valence scale though, subjects rated modern-threatening pictures significantly more negative than evolutionary-threatening pictures, indicating a higher level of perceived threat. The majority of previous studies show a positive relationship between arousal rating and amygdala activity. However, comparing fMRI results with behavioral findings we provide evidence that neural activity in fear processing areas is not only driven by arousal or valence, but presumably also by the evolutionary content of the stimulus. This has also fundamental methodological implications, in the sense to suggest a more elaborate classification of stimulus content to improve the validity of experimental designs.
Since exposure therapy for anxiety disorders incorporates extinction of contextual anxiety, relapses may be due to reinstatement processes. Animal research demonstrated more stable extinction memory and less anxiety relapse due to vagus nerve stimulation (VNS). We report a valid human three-day context conditioning, extinction and return of anxiety protocol, which we used to examine effects of transcutaneous VNS (tVNS). Seventy-five healthy participants received electric stimuli (unconditioned stimuli, US) during acquisition (Day1) when guided through one virtual office (anxiety context, CTX+) but never in another (safety context, CTX-). During extinction (Day2), participants received tVNS, sham, or no stimulation and revisited both contexts without US delivery. On Day3, participants received three USs for reinstatement followed by a test phase. Successful acquisition, i.e. startle potentiation, lower valence, higher arousal, anxiety and contingency ratings in CTX+ versus CTX-, the disappearance of these effects during extinction, and successful reinstatement indicate validity of this paradigm. Interestingly, we found generalized reinstatement in startle responses and differential reinstatement in valence ratings. Altogether, our protocol serves as valid conditioning paradigm. Reinstatement effects indicate different anxiety networks underlying physiological versus verbal responses. However, tVNS did neither affect extinction nor reinstatement, which asks for validation and improvement of the stimulation protocol.
Interpreting is generally recognized as a particularly demanding language processing task for the cognitive system. Dependency distance, the linear distance between two syntactically related words in a sentence, is an index of sentence complexity and is also able to reflect the cognitive constraints during various tasks. In the current research, we examine the difference in dependency distance among three interpreting types, namely, simultaneous interpreting, consecutive interpreting and read-out translated speech based on a treebank comprising these types of interpreting output texts with dependency annotation. Results show that different interpreting renditions yield different dependency distances, and consecutive interpreting texts entail the smallest dependency distance other than those of simultaneous interpreting and read-out translated speech, suggesting that consecutive interpreting bears heavier cognitive demands than simultaneous interpreting. The current research suggests for the first time that interpreting is an extremely demanding cognitive task that can further mediate the dependency distance of output sentences. Such findings may be due to the minimization of dependency distance under cognitive constraints.
This natural language processing toolkit provides language-agnostic 'tokenization', 'parts of speech tagging', 'lemmatization' and 'dependency parsing' of raw text. Next to text parsing, the package also allows you to train annotation models based on data of 'treebanks' in 'CoNLL-U' format as provided at <<a href="https://universaldependencies.org/format.html" target="_top">https://universaldependencies.org/format.html</a>>. The techniques are explained in detail in the paper: 'Tokenizing, POS Tagging, Lemmatizing and Parsing UD 2.0 with UDPipe', available at <<a href="https://doi.org/10.18653%2Fv1%2FK17-3009" target="_top">doi:10.18653/v1/K17-3009</a>>. The toolkit also contains functionalities for commonly used data manipulations on texts which are enriched with the output of the parser. Namely functionalities and algorithms for collocations, token co-occurrence, document term matrix handling, term frequency inverse document frequency calculations, information retrieval metrics (Okapi BM25), handling of multi-word expressions, keyword detection (Rapid Automatic Keyword Extraction, noun phrase extraction, syntactical patterns) sentiment scoring and semantic similarity analysis.
We describe our entry, C2L2, to the CoNLL 2017 shared task on parsing Universal Dependencies from raw text. Our system features an ensemble of three global parsing paradigms, one graph-based and two transition-based. Each model leverages character-level bidirectional LSTMs as lexical feature extractors to encode morphological information. Though relying on baseline tokenizers and focusing only on parsing, our system ranked second in the official end-toend evaluation with a macro-average of 75.00 LAS F1 score over 81 test treebanks. In addition, we had the top average performance on the four surprise languages and on the small treebank subset.
Important advances have recently been made using computational semantic models to decode brain activity patterns associated with concepts; however, this work has almost exclusively focused on concrete nouns. How well these models extend to decoding abstract nouns is largely unknown. We address this question by applying state-of-the-art computational models to decode functional Magnetic Resonance Imaging (fMRI) activity patterns, elicited by participants reading and imagining a diverse set of both concrete and abstract nouns. One of the models we use is linguistic, exploiting the recent word2vec skipgram approach trained on Wikipedia. The second is visually grounded, using deep convolutional neural networks trained on Google Images. Dual coding theory considers concrete concepts to be encoded in the brain both linguistically and visually, and abstract concepts only linguistically. Splitting the fMRI data according to human concreteness ratings, we indeed observe that both models significantly decode the most concrete nouns; however, accuracy is significantly greater using the text-based models for the most abstract nouns. More generally this confirms that current computational models are sufficiently advanced to assist in investigating the representational structure of abstract concepts in the brain.
The literary genre of poetry is inherently related to the expression and elicitation of emotion via both content and form. To explore the nature of this affective impact at an extremely basic textual level, we collected ratings on eight different general affective meaning scales—valence, arousal, friendliness, sadness, spitefulness, poeticity, onomatopoeia, and liking—for 57 German poems (“die verteidigung der wölfe”) which the contemporary author H. M. Enzensberger had labeled as either “friendly”, “sad”, or “spiteful”. Following Jakobson’s (1960) view on the vivid interplay of hierarchical text levels, we used multiple regression analyses to explore the specific influences of affective features from three different text levels (sublexical, lexical, and inter-lexical) on the perceived general affective meaning of the poems using three types of predictors: 1) Lexical predictor variables capturing the mean valence and arousal potential of words; 2) Inter-lexical predictors quantifying peaks, ranges and dynamic changes within the lexical affective content; 3) Sublexical measures of basic affective tone according to sound-meaning correspondences at the sublexical level (see Aryani, Kraxenberger, Ullrich, Jacobs, & Conrad, 2016). We find the lexical predictors to account for a major amount of up to 50 % of the variance in affective ratings. Moreover, inter-lexical and sublexical predictors account for a large portion of additional variance in the perceived general affective meaning. Together, the affective properties of all used textual features account for 43 to 70 % of the variance in the affective ratings and still for 23 to 48 % of the variance in the more abstract aesthetic ratings. In sum, our approach represents a novel method that successfully relates a prominent part of variance in perceived general affective meaning in this corpus of German poems to quantitative estimates of affective properties of textual components at the sublexical, lexical, and inter-lexical level.
Open data standards (e.g. LandXML, TransXML) have been widely recognized as a solution to the interoperability issue in exchanging digital data in the transportation sector. Since these schemas include rich sets of data types covering a wide range of disciplines across all project phases, model view definitions (MVDs) which define subsets of a schema are required to specify what types of data to be shared in accordance with a specific exchange scenario. The traditional method for generating MVDs is time consuming and tedious as developers have to manually search for entities and attributes names that semantically match to the data exchange requirements. This paper presents a computational method that automatically maps users’ keywords to semantics-equivalent data labels (classes and attributes) in LandXML data schema. The study employs a lexical database of civil engineering terms to interpret users’ intention from their keywords. The study also introduces a context-aware entity search algorithm that is able to find equivalent or most similar entities for a given keyword. The developed method has been experimented on a set of keywords extracted from an asset management manual. The experiment results show that the design algorithm is successful in generating partial LandXML branches from keywords.
Evidence suggests that in older adults, positive emotional memories are prioritized in order to enhance emotional well-being. Previous studies have demonstrated that sleep enhances negative emotional memories and preserves aspects of emotional reactivity associated with negative memories in young adults. Given that older adults prioritize positive memories, sleep may not preserve memory and reactivity for negative memories in this age group. Thus, the objective of this study is to investigate the influence of sleep on negative memories and emotional reactivity in older adults. Healthy older (55–80 yrs) adults viewed a mixture of negative and neutral pictures. During a three-hour delay, participants either napped (Nap group) or participated in restful wake activities (Wake group). Afterwards, participants underwent a picture recognition task. Emotional reactivity associated with picture viewing was measured during both sessions using valence and arousal scales, skin conductance response, heart rate deceleration, and corrugator supercilii electromyography. Contrary to what was observed in young adults using this procedure (presented in Jones et al. abstract), preliminary data in older adults suggest no benefit of sleep on negative memories (Nap: M=0.772, SD=0.164; Wake: M=0.868, SD=0.078) and no preservation of valence ratings in the nap group compared to the wake group (Nap: M=0.147, SD=0.087; Wake: M=0.261, SD=0.354). However, there is evidence for the preservation of skin conductance response (Nap: M=0.020, SD=0.231; Wake: M=-0.127, SD=0.192) and heart rate deceleration response (Nap: M=3.167, SD=9.879; Wake: M=-4.087, SD=8.633) in older adults. These initial results suggest that some but not all aspects of emotional reactivity associated with negative memories may be preserved by sleep in older adults. Sleep-dependent consolidation of negative memory contents may decline with age. This work was funded by NIH R01 AG040133 (PI: Spencer).
The extraction of abstract structures from speech (or from gestures in the case of sign languages) has been claimed to be a fundamental mechanism for language acquisition. In the present study we registered the neural responses that are triggered when a violation of an abstract, token-independent rule is detected. We registered ERPs while presenting participants with trisyllabic CVCVCV nonsense words in an oddball paradigm. Standard stimuli followed an ABB rule (where A and B are different syllables). Importantly, to distinguish neural responses triggered by changes in surface information from responses triggered by changes in the underlying abstract structure, we used two types of deviant stimuli. Phoneme deviants differed from standards only in their phonemes. Rule deviants differed from standards in both their phonemes and their composing rule. We observed a significant positivity as early as 300 ms after the presentation of deviant stimuli that violated the abstract rule (Rule dev)
This paper explores how information flow properties of a network affect the formation of categories shared between individuals, who are communicating through that network. Our work is based on the established multi-agent model of the emergence of linguistic categories grounded in external environment. We study how network information propagation efficiency and the direction of information flow affect categorization by performing simulations with idealized network topologies optimizing certain network centrality measures. We measure dynamic social adaptation when either network topology or environment is subject to change during the experiment, and the system has to adapt to new conditions. We find that both decentralized network topology efficient in information propagation and the presence of central authority (information flow from the center to peripheries) are beneficial for the formation of global agreement between agents. Systems with central authority cope well with network top)
This is the first study to examine the effect of phonetic contexts on children’s lexical tone production. Mandarin tones in disyllabic words produced by forty-four 2- to 6-year-old children and twelve mothers were low-pass filtered to eliminate lexical information. Native Mandarin-speaking adults categorized the tones based on the pitch information in the filtered stimuli. All mothers’ tones were categorized with ceiling accuracy. Counter to the findings in most previous studies on children’s tone acquisition and the prevailing assumption in models of speech development that children acquire suprasegmental features much earlier than segmental features, this study found that children as old as six years of age have not mastered the production of Mandarin tones. Children’s tones were judged with significantly lower accuracy than mothers’ productions. Tone accuracy improved, while cross subject variability in tone accuracy decreased, with age. Children’s tone accuracy was affected by the)
We learn language from our social environment. In general, the more sources we have, the less informative each of them is, and the less weight we should assign it. If this is the case, people who interact with fewer others should be more susceptible to the influence of each of their interlocutors. This paper tests whether indeed people who interact with fewer other people have more malleable phonological representations. Using a perceptual learning paradigm, this paper shows that individuals who regularly interact with fewer others are more likely to change their boundary between /d/ and /t/ following exposure to an atypical speaker. It further shows that the effect of number of interlocutors is not due to differences in ability to learn the speaker’s speech patterns, but specific to likelihood of generalizing the learned pattern. These results have implications for both language learning and language change, as they suggest that individuals with smaller social networks might play an )
Collective behaviour is a fascinating and easily observable phenomenon, attractive to a wide range of researchers. In biology, computational models have been extensively used to investigate various properties of collective behaviour, such as: transfer of information across the group, benefits of grouping (defence against predation, foraging), group decision-making process, and group behaviour types. The question ‘why,’ however remains largely unanswered. Here the interest goes into which pressures led to the evolution of such behaviour, and evolutionary computational models have already been used to test various biological hypotheses. Most of these models use genetic algorithms to tune the parameters of previously presented non-evolutionary models, but very few attempt to evolve collective behaviour from scratch. Of these last, the successful attempts display clumping or swarming behaviour. Empirical evidence suggests that in fish schools there exist three classes of behaviour; swarmi)
Recent theories propose that language comprehension can influence perception at the low level of perceptual system. Here, we used an adaptation paradigm to test whether processing language caused color adaptation in the visual system. After prolonged exposure to a color linguistic context, which depicted red, green, or non-specific color scenes, participants immediately performed a color detection task, indicating whether they saw a green color square in the middle of a white screen or not. We found that participants were more likely to perceive the green color square after listening to discourses denoting red compared to discourses denoting green or conveying non-specific color information, revealing that language comprehension caused an adaptation aftereffect at the perceptual level. Therefore, semantic representation of color may have a common neural substrate with color perception. These results are in line with the simulation view of embodied language comprehension theory, which )
Paintings have high cultural and commercial value, so that needs to be preserved. Many techniques have been attempted to analyze properties of paintings, including X-ray analysis and optical coherence tomography (OCT) methods, and enable conservation of paintings from forgeries. In this paper, we suggest a simple and accurate optical analysis system to protect them from counterfeit which is comprised of fiber optics reflectance spectroscopy (FORS) and line laser-based topographic analysis. The system is designed to fully cover the whole area of paintings regardless of its size for the accurate analysis. For additional assessments, a line laser-based high resolved OCT was utilized. Some forgeries were created by the experts from the three different styles of genuine paintings for the experiments. After measuring surface properties of paintings, we could observe the results from the genuine works and the forgeries have the distinctive characteristics. The forgeries could be distinguishe)
We present a new open source software tool called BEASTling, designed to simplify the preparation of Bayesian phylogenetic analyses of linguistic data using the BEAST 2 platform. BEASTling transforms comparatively short and human-readable configuration files into the XML files used by BEAST to specify analyses. By taking advantage of Creative Commons-licensed data from the Glottolog language catalog, BEASTling allows the user to conveniently filter datasets using names for recognised language families, to impose monophyly constraints so that inferred language trees are backward compatible with Glottolog classifications, or to assign geographic location data to languages for phylogeographic analyses. Support for the emerging cross-linguistic linked data format (CLDF) permits easy incorporation of data published in cross-linguistic linked databases into analyses. BEASTling is intended to make the power of Bayesian analysis more accessible to historical linguists without strong programmi)
In decision making, similarity measure and distance between two objects are crucial to be able to determine the relationship between those objects. Many researchers have received much attention for their research on this subject. In this study, we propose two novel similarity measures between hesitant fuzzy linguistic term sets (HFLTSs). In addition, two extensions of Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) are proposed in the hesitant fuzzy linguistic environments. Furthermore, an example of an application concerning traditional Chinese medical diagnosis and an MCDM problem have been given to illustrate the applicability and validation of these similarity measures of HFLTSs. Furthermore, the results of examples demonstrate that the Dice and Jaccard similarity measures are more reasonable than the cosine similarity measure with respect to HFLTSs. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its conten)
Background: To facilitate informed consent, consent forms should use language below the grade eight level. Research Ethics Boards (REBs) provide consent form templates to facilitate this goal. Templates with inappropriate language could promote consent forms that participants find difficult to understand. However, a linguistic analysis of templates is lacking. Methods: We reviewed the websites of 124 REBs for their templates. These included English language medical school REBs in Australia/New Zealand (n = 23), Canada (n = 14), South Africa (n = 8), the United Kingdom (n = 34), and a geographically-stratified sample from the United States (n = 45). Template language was analyzed using Coh-Metrix linguistic software (v.3.0, Memphis, USA). We evaluated the proportion of REBs with five key linguistic outcomes at or below grade eight. Additionally, we compared quantitative readability to the REBs’ own readability standards. To determine if the template’s country of origin or the presenc)
Infants preferentially discriminate between speech tokens that cross native category boundaries prior to acquiring a large receptive vocabulary, implying a major role for unsupervised distributional learning strategies in phoneme acquisition in the first year of life. Multiple sources of between-speaker variability contribute to children’s language input and thus complicate the problem of distributional learning. Adults resolve this type of indexical variability by adjusting their speech processing for individual speakers. For infants to handle indexical variation in the same way, they must be sensitive to both linguistic and indexical cues. To assess infants’ sensitivity to and relative weighting of indexical and linguistic cues, we familiarized 12-month-old infants to tokens of a vowel produced by one speaker, and tested their listening preference to trials containing a vowel category change produced by the same speaker (linguistic information), and the same vowel category produced )
Despite the ongoing growth in the number of published randomized controlled trials (RCTs) and increased quality assessment of RCTs, the association between the quality and characteristics in the text has not been sufficiently studied. We are interested in a specific question: what kind of sentences is a good indicator of high quality RCTs? To help researchers to efficiently screen articles worth reading, this study aims 1) to quantify the linguistic features of articles and 2) to build a document assessment model to evaluate quality of RCTs using only the abstract. All RCTs that were conducted in Japan in 2010 as original articles were included in the analysis. Data were independently assessed by two reviewers using a risk-of-bias tool. Three aspects of linguistic style were quantitatively measured, and a document model was constructed to evaluate the RCTs. A total of 302 RCTs were selected for quality assessment. Of these, 255 articles were assessed as high quality and 47 as low qual)
Background: Most of earlier studies in the field of literature-based discovery have adopted Swanson's ABC model that links pieces of knowledge entailed in disjoint literatures. However, the issue concerning their practicability remains to be solved since most of them did not deal with the context surrounding the discovered associations and usually not accompanied with clinical confirmation. In this study, we aim to propose a method that expands and elaborates the existing hypothesis by advanced text mining techniques for capturing contexts. We extend ABC model to allow for multiple B terms with various biological types. Results: We were able to concretize a specific, metabolite-related hypothesis with abundant contextual information by using the proposed method. Starting from explaining the relationship between lactosylceramide and arterial stiffness, the hypothesis was extended to suggest a potential pathway consisting of lactosylceramide, nitric oxide, malondialdehyde, and arteria)
This opinion paper proposes the use of parallel treebank as learner corpus. We show how an L1-L2 parallel treebank — i.e., parse trees of non-native sentences, aligned to the parse trees of their target hypotheses — can facilitate retrieval of sentences with specific learner errors. We argue for its benefits, in terms of corpus re-use and interoperability, over a conventional learner corpus annotated with error tags. As a proof of concept, we conduct a case study on word-order errors made by learners of Chinese as a foreign language. We report precision and recall in retrieving a range of word-order error categories from L1-L2 tree pairs annotated in the Universal Dependency framework.
In spite of decades of theorizing, the origins of Zipf’s law remain elusive. I propose that a Zipfian distribution straightforwardly follows from the interaction of syntax (word classes differing in class size) and semantics (words having to be sufficiently specific to be distinctive and sufficiently general to be reusable). These factors are independently motivated and well-established ingredients of a natural-language system. Using a computational model, it is shown that neither of these ingredients suffices to produce a Zipfian distribution on its own and that the results deviate from the Zipfian ideal only in the same way as natural language itself does. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied or emailed to multiple sites or posted to a listserv without the copyright holder's express written permission. However, users may print, download, or email articles for individual use. This abstract may be )
l’aéroport et qui découvre qu’elle est amnésique) met en question la liberté et le choix individuel dont le personnage se croit doté. University of Portland (OR) Khadija Khalifé Linguistics edited by Bryan Donaldson Bertucci, Marie-Madeleine, éd. Les français régionaux dans l’espace francophone. Berne: Peter Lang, 2016. ISBN 978-3-631-64650-2. Pp. 251. The fourteen articles in this volume address the question of the linguistic status of les français régionaux. The varieties discussed are found in Europe (le cauchois, le parler du Nord-Pas-de-Calais, French in la Belgique francophone) and beyond (le créole d’Haïti, le français acadien, le français calédonien, le français louisianais). I will highlight two issues. The first is how to define the term français régional. The traditional dialectological definition—a regionally or geographically delimited variety, sometimes called français dialectal, patoisé or d’usance—evolved, by the 1980s, to a variety that is in some way “subordinate to a norm” usually referred to as français standard, français de Paris, or français “des Français.” As researchers studied varieties located in language contact situations where French coexists with endogenous languages, regional varieties were situated along axes related to endogenous and exogenous norms. In this volume, the authors adopt various approaches, such as glottonomic, phenomenologicalhermeneutic, and critical sociolinguistic, in defining their work. They show, convincingly, that regional varieties are characterized by complex arrays of features, but the reader is left with no clear definition of the term français régional. I would point out one simple criterion that is mentioned in several articles: the importance of naming. The name of a variety defines a social space in which speakers are allowed to use linguistic variants that do not always correspond to the variants of the dominant norm; that is, naming gives a variety an existence and a life. The second issue of interest is the linguistic description of regional varieties. In many articles, the emphasis is on cataloging specific lexical items that are unique to a particular variety. However, a sociolinguist reading these lists would like to see some quantitative information about the usage of these variants in different contexts. Furthermore, descriptions of usage in nonlexical areas (phonetics, morphosyntax, discourse) would greatly add to our understanding of variation in these varieties. This information would also inform the question of what is to be taught in schools. There is a general consensus among the authors that schools play an important role in validating endogenous norms and in transmitting culturally relevant values. However, it is not clear that these endogenous 270 FRENCH REVIEW 91.2 Reviews 271 norms have been well described. To conclude, there is no clear synthesis of what has been found in this wide-ranging collection of articles, even though the editor’s overview identifies some important linkages among the approaches and observations. That said, this volume’s original contribution is the fascinating and well-documented information that it presents about different regions of the Francophone world. Its main usefulness as a reference book lies in the glimpses that the individual authors offer into the varieties that they are working on. University of New Brunswick Wladyslaw Cichocki Brunet, Roger. Trésor du terroir: les noms de lieux de la France. Paris: CNRS, 2016. ISBN 978-2-271-08816-1. Pp. 655. Here, at first glance, is an incontournable new compendium of French toponymy covering 25,000 noms de lieux (NL) and lexical families. Brunet’s pioneering onomasiological approach to toponymy asks: Which concepts serve to name the environments (friendly or hostile terrain, dangers, curiosities, etc.) we inhabit? After a short introduction, chapters one through six cover a range of descriptive categories: “Habiter et s’abriter,”“Pays et chemins: le territoire et ses réseaux,”“La vie sociale et ses distinctions,” “Terrains de jeu,” “Eaux, bords d’eaux et météores,” and “Paysages, ressources et travaux.”The next two chapters chronicle the evolution of NLs. Chapter seven,“La vie des noms de lieux,”moves through language change, politics, innovation, and more. Continuing this primarily linguistic analysis, chapter eight, “À distance: pièges...
Previous research has mainly considered the impact of tone-language experience on ability to discriminate linguistic pitch, but proficient bilingual listening requires differential processing of sound variation in each language context. Here, we ask whether Mandarin-English bilinguals, for whom pitch indicates word distinctions in one language but not the other, can process pitch differently in a Mandarin context vs. an English context. Across three eye-tracked word-learning experiments, results indicated that tone-intonation bilinguals process tone in accordance with the language context. In Experiment 1, 51 Mandarin-English bilinguals and 26 English speakers without tone experience were taught Mandarin-compatible novel words with tones. Mandarin-English bilinguals out-performed English speakers, and, for bilinguals, overall accuracy was correlated with Mandarin dominance. Experiment 2 taught 24 Mandarin-English bilinguals and 25 English speakers novel words with Mandarin-like tones,)
Automatic extraction of protein-protein interaction (PPI) pairs from biomedical literature is a widely examined task in biological information extraction. Currently, many kernel based approaches such as linear kernel, tree kernel, graph kernel and combination of multiple kernels has achieved promising results in PPI task. However, most of these kernel methods fail to capture the semantic relation information between two entities. In this paper, we present a special type of tree kernel for PPI extraction which exploits both syntactic (structural) and semantic vectors information known as Distributed Smoothed Tree kernel (DSTK). DSTK comprises of distributed trees with syntactic information along with distributional semantic vectors representing semantic information of the sentences or phrases. To generate robust machine learning model composition of feature based kernel and DSTK were combined using ensemble support vector machine (SVM). Five different corpora (AIMed, BioInfer, HPRD50, )
This paper describes a Romanian Dependency Treebank, built at the Al. I. Cuza University (UAIC), and a special OCR techniques used to build it. The corpus has rich morphological and syntactic annotation. There are few annotated representative corpora in Romanian, and the existent ones are mainly focused on the contemporary Romanian standard. The corpus described below is focused on the nonstandard aspects of the language, the Regional and the Old Romanian. Having the intention to participate at the PROIEL project, which aligns oldest New Testaments, we annotate the first printed Romanian New Testament (Alba Iulia, 1648). We began by applying the UAIC tools for the morphological and syntactic processing of Contemporary Romanian over the books first quarter (second edition). By carefully manually correcting the result of the automated annotation (having a modest accuracy) we obtained a sub-corpus for the training of tools for the Old Romanian processing. But the first edition of the New Testament is written in Cyrillic letters. The existence of books printed in the Old Cyrillic alphabet is a common problem for Romania and The Republic of Moldova, countries where the Romanian is spoken; a problem to solve by the joint efforts of the NLP researchers in the two countries.
We propose new computational models for analyzing self-reported emotional diary texts of pregnant women to support maternal care. We gathered affective ratings outside clinical setting and developed new models to facilitate interpretation and communication of affective expressions between persons representing different affective ratings. Relying on constructed emotion theory, models of dimensional emotion categories and affective ratings of Self Assessment Manikin, we demonstrate our new proposal to analyze linguistic data with computational models exploiting vector space and clustering methods. 35 persons having Finnish as a native language provided affective ratings for 195 emotional adjectives and 16 pregnancy-related nouns in Finnish in dimensions of pleasure, arousal and dominance. We developed new models to represent dependencies and differences of affective ratings between various population subgroup categorizations, including "women without children", "women with children" and "men without children" that we consider important population segments to be addressed in maternal care. Our affective ratings showed significant correlations between pleasure and dominance (like Warriner et al., 2013) and with previous data collections (Söderholm et al., 2013; Eilola & Havelka, 2010; Warriner et al., 2013). Our affective ratings had significant effects on categorizations based on gender, gender-parental role and the time of the day and duration of giving ratings. Our results indicate accordance with significant affectivity differences of gender and age (Warriner et al., 2013) and motherhood (Rosebrock et al., 2015). Our proposed models aim to support health-related communication. Our results suggest gathering next the affective ratings of patients of maternal care in a real clinical setting.
This paper presents a frame annotation scheme for Danish nouns, with VerbNet-derived frames and semantic roles covering both frame arguments and satellites. The scheme was implemented as a new module for a Danish frame tagger and applied to a 90, 000-token Danish treebank with ongoing manual revision. In addition to explicit frames, Constraint Grammar rules are used to map free semantic roles on noun dependents without pre-defined frames, using general syntactic-semantic context clues. We discuss the annotation scheme and present a statistical breakdown and linguistic evaluation of the assigned noun frames and adnominal roles in the corpus.
The article begins with a presentation of a selection of electronic monolingual and bi/multilingual lexicographic resources and corpora available today to contemporary users of Slovene. The focus is on works combined with English and designed for translation purposes which provide information on the meaning of words and wider lexical units, i.e., e-dictionaries, lexical databases, web translation tools and various corpora. In a separate sub-section the most common translation technologies are presented, together with an evaluation of their role in the modern translation process. Sections 2 and 3 provide a brief outline of the changes that have affected classical dictionary planning, compilation and use in the new digital environment, as well as of the relationship between dictionaries and related resources, such as lexical databases. Some stereotypes regarding dictionary use are identified and, in conclusion, the existing corpus-based databases for the Slovenian-English pair are presented, with a view to determining priorities for the future interlingual infrastructure action plans in Slovenia.
Modifying the style of movements will be an important component of robotic interaction as more and more robots move into human-facing scenarios where humans are (consciously or unconsciously) constantly monitoring the motion profile of counterparts in order to make judgments about the state of these counterparts. This thesis includes two main contributions: (1) the development of two MATLAB tools that are designed to aid in the creation and simulation of stylized movement trajectories in varied contexts and (2) three user studies that explore the effects of environmental context on a human’s perception of stylized movement. \n \nFirst and foremost, the results from all of the user studies indicate that environmental contexts and stylized walking sequences both impact affect recognition. In the first two studies, participants were asked to categorize stimuli as one of seven affective labels. The results show that the labels were not applied consistently and so it was concluded that the affect of a multi-dimensional stimuli cannot be adequately categorized using a single affective label. In the third study the stimuli were evaluated on multiple scales and classified using ratings of valence and arousal rather than affective labels. The results were used to create a least squares model for the dataset that decomposed the affect ratings of animations to display the compound effects of stylized walking sequences and environmental contexts on affective ratings.
The present study examined the significance of viewing images of neutral faces versus images of neutral objects on zygomatic muscle activity using facial EMG. Participants (60% women) from a pool of introductory psychology courses had their facial EMG recordings measured in response to images of neutral faces and neutral objects. Participantsâ valence rating of each image was also recorded using the Self-Assessment Manikin (SAM) in order to rate their emotional response to each image. The primary hypothesis was that participants would have greater activity in the zygomatic muscle region when presented with images of neutral faces as opposed to lessor activity when presented with images of neutral objects. It was also hypothesized that if participants preferred seeing images of faces as compared to objects, their positive feelings would produce higher SAM ratings. Results from the present study indicated images of neutral faces showed no significant difference in EMG activity compared to images of neutral objects. Self-report data also showed no significant difference in pleasantness or emotional valence between ratings of neutral faces and ratings of neutral objects.
Measuring the similarity between two sentences is often difficult due to their small lexical overlap. Instead of focusing on the sets of features in two given sentences between which we must measure similarity, we propose a sentence similarity method that considers two types of constraints that must be satisfied by all pairs of sentences in a given corpus. Namely, (a) if two sentences share many features in common, then it is likely that the remaining features in each sentence are also related, and (b) if two sentences contain many related features, then those two sentences are themselves similar. The two constraints are utilized in an iterative bootstrapping procedure that simultaneously updates both word and sentence similarity scores. Experimental results on SemEval 2015 Task 2 dataset show that the proposed iterative approach for measuring sentence semantic similarity is significantly better than the non-iterative counterparts. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the )
In this study we present a novel set of discrimination-based indicators of language processing derived from Naive Discriminative Learning () theory. We compare the effectiveness of these new measures with classical lexical-distributional measures—in particular, frequency counts and form similarity measures—to predict lexical decision latencies when a complete morphological segmentation of masked primes is or is not possible. Data derive from a re-analysis of a large subset of decision latencies from the English Lexicon Project, as well as from the results of two new masked priming studies. Results demonstrate the superiority of discrimination-based predictors over lexical-distributional predictors alone, across both the simple and primed lexical decision tasks. Comparable priming after masked and type primes, across two experiments, fails to support early obligatory segmentation into morphemes as predicted by the morpho-orthographic account of reading. Results fit well with theory, wh)
galegoEste traballo describe o procedemento de deseno e construcion dun corpus ingles- galego lematizado e desambiguado semanticamente con respecto aos sentidos das palabras definidas nunha base de datos lexica. Partese dun conxunto de textos en ingles xa anotados coas etiquetas correspondentes aos nomes, verbos, adxectivos e adverbios; estes textos traducense ao galego e as palabras galegas anotanse co lema e o sentido lexico. O resultado, o corpus SensoGal, representa un recurso util que calquera usuario pode consultar e reutilizar, ao tempo que facilita a presenza do idioma galego no ambito das tecnoloxias. Nas seguintes seccions presentase o proceso de elaboracion en que se identifican as dificultades atopadas nas fases de traducion e anotacion e se rexistran as decisions tomadas por se poden servir como referencia na esperable continuacion do proxecto. Tamen se detalla o sistema de consultas e se fai unha reflexion sobre os resultados obtidos e o posible traballo futuro. EnglishThis paper presents the design and elaboration of an English-Galician corpus lemmatized and semantically disambiguated with respect to the meanings of the words defined in a lexical database. To perform the task, we used a group of texts in English where nouns, verbs, adjectives and adverbs were already tagged; these texts were translated into Galician and the words tagged with their lemma and lexical meaning. The result is the corpus SensoGal: a useful resource for users and linguists that facilitates the presence of Galician language in the field of technology. In the following sections the elaboration process will be described in the phases of translation and labelling by registering the difficulties met and the decisions taken to serve as a reference in the foreseeable continuation of the project. The search system will be also explained. Finally, a reflection about the results and the future work will be done.
The CELEX lexical database (Baayen, Piepenbrock & van Rijn 1995) was developed in the 1990s, providing a database of the syntactic, morphological, phonological and orthographic forms of between 50,000 and 125,000 words of Dutch, English and German. This database was used as the basis for the development of the PolyLex lexicons, which included syntactic, morphological and phonological information for around 3,000 words of Dutch, English and German. Orthographic information was subsequently added in the PolyOrth project. The PolyOrth project was based on the assumption that the underlying, lexical phonological forms could be used to derive the surface orthographic forms by means of a combination of phoneme-grapheme mappings and sets of autonomous spelling rules for each language. One of the complications encountered during the project was the fact that the phonological forms in CELEX were not always genuinely underlying forms which made deriving the orthographic forms tricky. This paper discusses the nature and status of underlying phonological forms, their relation to orthography and the issues of finding this information in databases. (PsycINFO Database Record (c) 2018 APA, all rights reserved)
The paper presents the project for creation of lexical database that is intended to include the Bulgarian and Czech neologisms. It describes the principles for selection of lexical units which should be included in the database as well as the resources of the lexical material. The paper presents also the database structure and the content of its modules and submodules.
In this paper, extensive experiments are conducted to study the impact of features of different categories, in isolation and gradually in an incremental manner, on Arabic Person name recognition. We present an integrated system that employs the rule-based approach with the machine learning (ML)-based approach in order to develop a consolidated hybrid system. Our feature space is comprised of language-independent and language-specific features. The explored features are naturally grouped under six categories: Person named entity tags predicted by the rule-based component, word-level features, POS features, morphological features, gazetteer features, and other contextual features. As decision tree algorithm has proved comparatively higher efficiency as a classifier in current state-of-the-art hybrid Named Entity Recognition for Arabic, it is adopted in this study as the ML technique utilized by the hybrid system. Therefore, the experiments are focused on two dimensions: the standard dataset used and the set of selected features. A number of standard datasets are used for the training and testing of the hybrid system, including ACE (2003–2004) and ANERcorp. The experimental analysis indicates that both language-independent and language-specific features play an important role in overcoming the challenges posed by Arabic language and have demonstrated critical impact on optimizing the performance of the hybrid system.
In this paper, we present the results of searching for long-distance dependencies in an automatically annotated treebank for Dutch. We concentrate on phenomena that have recently been subject to debate, and where conflicting claims have been made regarding the question whether these constructions actually occur with some frequency in spontaneous language use. Long-distance dependencies involving a tensed or infinitival subordinate clause are quite rare and show collocational effects. Resumptive prolepsis and R-pronominal parasitic gaps are outside the scope of the computational grammar. We show that access to syntactic annotation even in such cases helps to find positive examples relatively quickly.
In the field of word recognition and reading, it is commonly assumed that frequently repeated words create more accessible memory traces than infrequently repeated words, thus capturing the word-frequency effect. Nevertheless, recent research has shown that a seemingly related factor, contextual diversity (defined as the number of different contexts [e.g., films] in which a word appears), is a better predictor than word-frequency in word recognition and sentence reading experiments. Recent research has shown that contextual diversity plays an important role when learning new words in a laboratory setting with adult readers. In the current experiment, we directly manipulated contextual diversity in a very ecological scenario: at school, when Grade 3 children were learning words in the classroom. The new words appeared in different contexts/topics (high-contextual diversity) or only in one of them (low-contextual diversity). Results showed that words encountered in different contexts we)
Dyslexia has been claimed to be causally related to deficits in visuo-spatial attention. In particular, inefficient shifting of visual attention during spatial cueing paradigms is assumed to be associated with problems in graphemic parsing during sublexical reading. The current study investigated visuo-spatial attention performance in an exogenous cueing paradigm in a large sample (N = 191) of third and fourth graders with different reading and spelling profiles (controls, isolated reading deficit, isolated spelling deficit, combined deficit in reading and spelling). Once individual variability in reaction times was taken into account by means of z-transformation, a cueing deficit (i.e. no significant difference between valid and invalid trials) was found for children with combined deficits in reading and spelling. However, poor readers without spelling problems showed a cueing effect comparable to controls, but exhibited a particularly strong right-over-left advantage (position effec)
Reviewed by: Corpus Stylistics as Contextual Prosodic Theory and Subtext by Bill Louw, Marija Milojkovic Feng (Robin) Wang (bio) and Philippe Humblé (bio) Bill Louw and Marija Milojkovic. Corpus Stylistics as Contextual Prosodic Theory and Subtext. John Benjamins Publishing Company, 2016. xix + 419 pp. $149. The term corpus stylistics, usually regarded as a near-synonym for stylometry, stylometrics, statistical stylistics, or stylogenetics, is closely related to statistics and corpus linguistics. Despite an increasing number of studies in the field, people still do not attain a clear line of demarcation between corpus linguistics and corpus stylistics. Corpus linguists are typically concerned with “repeated occurrences, generalizations and the description of typical patterns,” while corpus stylistic studies relate to “deviations from linguistic norms that account for the artistic effects of a particular text” (Mahlberg, “Corpus Stylistic Perspective” 19). However, more needs to be known about what new perspectives corpus linguistics can offer to the depiction of stylistic devices and the interpretation of stylistic values. Under these circumstances, Bill Louw and Marija Milojkovic’s Corpus Stylistics as Contextual Prosodic Theory and Subtext is instructive and worthy of reading, for it offers valuable perspectives for interdisciplinary investigations. This volume comprises two parts: the first part (Chapters 1–6) is devoted to the theoretical construction of Contextual Prosodic Theory (CPT), and the second part (Chapters 7–12) applies CPT to literary criticism, translation studies, and foreign language teaching. Chapter 1 revisits the proposal on “language and literature integration” in foreign language teaching. Louw dissolves the doubts from language teachers about “integration” by sufficiently discussing lexical syllabus design and progressive delexicalization. Having critically reviewed different theoretical perspectives on collocation, the authors argue in [End Page 550] Chapter 2 that one objective characteristic of literary devices is that they will demonstrate some evidence of relexicalization through collocation. Chapter 3 focuses on the theoretical interpretation of semantic prosody. Semantic prosody, according to Louw, is the “consistent aura of meaning with which a form is imbued by its collocations” (80). In Chapter 4, the author expounds that data-driven reading will produce a class of negotiator distinct from the intuitive counterparts. Chapter 5 affirms the role of collocation in terms of predicting and grading the potential success of all humorous contexts of situation as well as composition. Moreover, the interaction between collocation and events in the external world is capable of isolating humorous situations that are “waiting to happen” (132). Chapter 6 introduces subtext, a core concept of CPT, and proceeds to explore what these deviations from logical semantic prosody (subtext) can tell us about an author’s text. The second part (Chapters 7–12) is written by Milojkovic and adapts CPT to other disciplines. In this sense, the volume can be considered as a necessary reference for a consortium of scholars. In order to test the applicability and universality of CPT, Milojkovic applies CPT to Slavic languages, namely, Russian and Serbian. Based on a synthesis of the theoretical tools of CPT (i.e., collocation, semantic prosody, and subtext), Milojkovic analyzes the logical construction of literary worlds as well as a hitherto uncharted domain in corpus stylistics: authorial intention, that is, whether the author sincerely means what he or she writes. Chapter 8 reveals the subtext of “in the * of” in a translated poem of Pushkin as a picture of action verging on conflict, which inspires Milojkovic to probe into whether this is an incompatible grammatical pattern to express Pushkin’s call for resignation. Methodologically, the application of CPT in translation studies enriches the theoretical toolkit of corpus-based translation studies. Chapter 9 distinguishes inspired writing from banality by evaluating the deviation from the reference corpus. Chapter 10 puts forward the hypothesis that inspired writing will differ from uninspired in the density of its subtextual and prosodic clashes, and that the clashes themselves will be indicative of the presence of inspiration (274). In order to test this hypothesis, Milojkovic, in Chapter 10, contacts several poets to elicit clear-cut cases of inspired writing. The final two chapters, concerning applications for foreign language teaching, pertain to time-honored pedagogical stylistics. Chapter 11 is a piece [End Page 551] of classroom corpus stylistics research with a twofold purpose: empirically, to verify Louw...
Authorship attribution is to identify the most likely author of a given sample among a set of candidate known authors. It can be not only applied to discover the original author of plain text, such as novels, blogs, emails, posts etc., but also used to identify source code programmers. Authorship attribution of source code is required in diverse applications, ranging from malicious code tracking to solving authorship dispute or software plagiarism detection. This paper aims to propose a new method to identify the programmer of Java source code samples with a higher accuracy. To this end, it first introduces back propagation (BP) neural network based on particle swarm optimization (PSO) into authorship attribution of source code. It begins by computing a set of defined feature metrics, including lexical and layout metrics, structure and syntax metrics, totally 19 dimensions. Then these metrics are input to neural network for supervised learning, the weights of which are output by PSO a)
Norming across many sets of affective pictures in order to compare affect ratings across perceptual and semantic features of images.
Language comprehension involves the simultaneous processing of information at the phonological, syntactic, and lexical level. We track these three distinct streams of information in the brain by using stochastic measures derived from computational language models to detect neural correlates of phoneme, part-of-speech, and word processing in an fMRI experiment. Probabilistic language models have proven to be useful tools for studying how language is processed as a sequence of symbols unfolding in time. Conditional probabilities between sequences of words are at the basis of probabilistic measures such as surprisal and perplexity which have been successfully used as predictors of several behavioural and neural correlates of sentence processing. Here we computed perplexity from sequences of words and their parts of speech, and their phonemic transcriptions. Brain activity time-locked to each word is regressed on the three model-derived measures. We observe that the brain keeps track of t)