Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
We evaluate two cross-lingual techniques for adding enhanced dependencies to existing treebanks in Universal Dependencies. We apply a rule-based system developed for English and a data-driven system trained on Finnish to Swedish and Italian. We find that both systems are accurate enough to bootstrap enhanced dependencies in existing UD treebanks. In the case of Italian, results are even on par with those of a prototype language-specific system.
Developing experimental materials to study metaphor memory is a difficult process, especially in regard to finding a suitable type of control stimulus. The dead metaphors were technically metaphorical but have long since come into common usage; they were thus considered an intermediate type between metaphors and nonmetaphors. The fact that images were reported more often with metaphorical than non-metaphorical sentences suggests that imagery may be a more frequently used mnemonic with metaphors than with more literal language, even though it is not necessarily all that effective. The greater number of counting errors for metaphors may be attributed to differential saliency of the input sentences. The metaphors were novel, while the dead metaphors and nonmetaphors were commonplace. People’s intuitions about metaphors are often highly inaccurate, especially in the direction of thinking them to be more difficult, less informative, and less imageable than they truly are.
When adding enhanced dependencies to an existing UD treebank, one can opt for heuristics that predict the enhanced dependencies on the basis of the UD annotation only. If the treebank is the result of conversion from an underlying treebank, an alternative is to produce the enhanced dependencies directly on the basis of this underlying annotation. Here we present a method for doing the latter for the Dutch UD treebanks. We compare our method with the UD -based approach of Schuster et al. (2018). While there are a number of systematic differences in the output of both methods, it appears these are the result of insufficient detail in the annotation guidelines and it is not the case that one approach is superior over the other in principle.
This paper describes the development of the first syntactically-annotated corpus of Breton. The corpus is part of the Universal Dependencies project. In the paper we describe how the corpus was prepared, some Breton-specific constructions that required special treatment, and in addition we give results for parsing Breton using a number of off-the-shelf data-driven parsers.
In this study, we tested the linguistic relativity hypothesis by studying the effect of grammatical gender (feminine vs. masculine) on affective judgments of conceptual representation in Italian and German. In particular, we examined the within- and cross-language grammatical gender effect and its interaction with participants' demographic characteristics (such as, the raters' age and sex) on semantic differential scales (affective ratings of valence, arousal and dominance) in Italian and German speakers. We selected the stimuli and the relative affective measures from Italian and German adaptations of the ANEW (Affective Norms for English Words). Bayesian and frequentist analyses yielded evidence for the absence of within- and cross-languages effects of grammatical gender and sex- and age-dependent interactions. These results suggest that grammatical gender does not affect judgments of affective features of semantic representation in Italian and German speakers, since an overt coding of word grammar is not required. Although further research is recommended to refine the impact of the grammatical gender on properties of semantic representation, these results have implications for any strong view of the linguistic relativity hypothesis.
Recurrent neural networks (RNNs) are powerful models of sequential data. They have been successfully used in domains such as text and speech. However, RNNs are susceptible to overfitting; regularization is important. In this paper we develop Noisin, a new method for regularizing RNNs. Noisin injects random noise into the hidden states of the RNN and then maximizes the corresponding marginal likelihood of the data. We show how Noisin applies to any RNN and we study many different types of noise. Noisin is unbiased--it preserves the underlying RNN on average. We characterize how Noisin regularizes its RNN both theoretically and empirically. On language modeling benchmarks, Noisin improves over dropout by as much as 12.2% on the Penn Treebank and 9.4% on the Wikitext-2 dataset. We also compared the state-of-the-art language model of Yang et al. 2017, both with and without Noisin. On the Penn Treebank, the method with Noisin more quickly reaches state-of-the-art performance.
This paper presents a treebank for the healthcare domain developed at ezDI. The treebank is created from a wide array of clinical health record documents across hospitals. The data has been de-identified and annotated for constituent syntactic structure. The treebank contains a total of 52053 sentences that have been sampled for subdomains as well as linguistic variations. The paper outlines the sampling process followed to ensure a better domain representation in the corpus, the annotation process and challenges, and corpus statistics. The Penn Treebank tagset and guidelines were largely followed, but there were many syntactic contexts that warranted adaptation of the guidelines. The treebank created was used to re-train the Berkeley parser and the Stanford parser. These parsers were also trained with the GENIA treebank for comparative quality assessment. Our treebank yielded great-er accuracy on both parsers. Berkeley parser performed better on our treebank with an average F1 measure of 91 across 5-folds. This was a significant jump from the out-of-the-box F1 score of 70 on Berkeley parser’s default grammar.
<h3>Introduction</h3><br> DEFT Spanish Treebank was developed by the Linguistic Data Consortium (LDC) and the <a href="http://clic.ub.edu/">Language and Computation Center (CLiC), University of Barcelona</a>. It contains treebank annotation of international Spanish newswire text and Latin American Spanish discussion forum data created for the DARPA Deep Exploration and Filtering of Text (DEFT) program. <br> DEFT aimed to improve state-of-the-art capabilities in automated deep natural language processing with a particular focus on technologies dealing with inference, casual relationships and anomaly detection across several languages. DEFT Spanish Treebank supported the program's goal of deep natural language understanding. <br> <h3>Data</h3><br> Newswire source files were selected from Spanish Gigaword Third Edition (<a href="../../../ldc2011t12">LDC2011T12</a>) and were manually sentence-segmented for DEFT. Discussion forum source files were selected from Spanish discussion forum source data collected by LDC, consisting of continuous multi-posts of 100-1000 words. <br> This release contains 114 files (54,394 tokens) of newswire data and 60 files (55,307 tokens) of discussion forum data all of which were annotated with constituents and syntactic functions. The annotation guidelines for DEFT Spanish Treebank are included in the documentation accompanying this release. <br> Source documents are presented as plain text files with one sentence unit per line. Treebank annotation files are in xml. <br> <h3>Samples</h3><br> Please view this <a href="desc/addenda/LDC2018T01.txt">source sample</a> and <a href="desc/addenda/LDC2018T01.xml">treebank sample</a>. <br> <h3>Updates</h3><br> None at this time. <br> <h3>Acknowledgement</h3><br> This material is based on research sponsored by Air Force Research Laboratory and Defense Advance Research Projects Agency under agreement number FA8750-13-2-0045. The U.S. Government is authorized to reproduce and distribute reprints for Governmental purposes notwithstanding any copyright notation thereon. The views and conclusions contained herein are those of the authors and should not be interpreted as necessarily representing the official policies or endorsements, either expressed or implied, of Air Force Research Laboratory and Defense Advanced Research Projects Agency or the U.S. Government. </br> Portions © 1994-2001, 2004-2009 The Associated Press, © 2002, 2005, 2007, 2009-2010 Xinhua News Agency, © 2006, 2009, 2011, 2018 Trustees of the University of Pennsylvania
In this paper we discuss the project of digitization of the Dictionary of the Serbo-Croatian Standard and Vernacular \nLanguage. Scanning and character recognition were a particular challenge, since various non-standard \ncharacter set encoding was used in the course of the almost 60-year long production of the dictionary. The first \naim of the project was to formalize the micro-structure of the dictionary articles in order to parse the digitized \ntext of and transform it into structured data stored in relational lexical database. This approach is compatible \nwith several standard structured forms and ontologies (TEI, LMF, Ontolex, LexInfo). A lexical database model \nwas designed in compliance with these structured forms, following mostly the lemon model. Mapping of \nthe lexical entry markers to LexInfo and TEI enabled export of the lexical data to the mentioned formats. A \nsoftware solution for the dictionary text analysis, parsing and lexical database population was developed and \ntested on the first and the last published volumes of the dictionary (which contain 27,141 articles in total). An \nevaluation of the results shows that the developed model and software solution can be successfully used for \nthe other volumes as well.
BACKGROUND: Life events (LEs) are associated with future physical and mental health. They are crucial for understanding the pathways to mental disorders as well as the interactions with biological parameters. However, deeper insight is needed into the complex interplay between the type of LE, its subjective evaluation and accompanying factors such as social support. The "Stralsund Life Event List" (SEL) was developed to facilitate this research. METHODS: The SEL is a standardized interview that assesses the time of occurrence and frequency of 81 LEs, their subjective emotional valence, the perceived social support during the LE experience and the impact of past LEs on present life. Data from 2265 subjects from the general population-based cohort study "Study of Health in Pomerania" (SHIP) were analysed. Based on the mean emotional valence ratings of the whole sample, LEs were categorized as "positive" or "negative". For verification, the SEL was related to lifetime major depressive disorder (MDD; Munich Composite International Diagnostic Interview), childhood trauma (Childhood Trauma Questionnaire), resilience (Resilience Scale) and subjective health (SF-12 Health Survey). RESULTS: The report of lifetime MDD was associated with more negative emotional valence ratings of negative LEs (OR = 2.96, p < 0.0001). Negative LEs (b = 0.071, p < 0.0001, β = 0.25) and more negative emotional valence ratings of positive LEs (b = 3.74, p < 0.0001, β = 0.11) were positively associated with childhood trauma. In contrast, more positive emotional valence ratings of positive LEs were associated with higher resilience (b = - 7.05, p < 0.0001, β = 0.13), and a lower present impact of past negative LEs was associated with better subjective health (b = 2.79, p = 0.001, β = 0.05). The internal consistency of the generated scores varied considerably, but the mean value was acceptable (averaged Cronbach's alpha > 0.75). CONCLUSIONS: The SEL is a valid instrument that enables the analysis of the number and frequency of LEs, their emotional valence, perceived social support and current impact on life on a global score and on an individual item level. Thus, we can recommend its use in research settings that require the assessment and analysis of the relationship between the occurrence and subjective evaluation of LEs as well as the complex balance between distressing and stabilizing life experiences.
The author’s attention is drawn to two problem areas of modern linguistics: the study of lexical compatibility and further development of the concept of national variability of the German language. Collocations are chosen as an object of research. They are considered on the one hand as a basic phenomenon of lexical combinatorics, and on the other hand, as a special type of phraseological units, reflecting the features of all system levels of language or language variant. Based on the comparative analysis of Austrian, Swiss and German collocations extracted from lexicographical sources, the author shows that the norms of lexical compatibility in the German language of Germany, Austria and Switzerland in a number of aspects do not coincide. It is noted that national language standards have a significant number of collocations that are unknown and / or uncommon in other regions of the German-speaking area and in many cases have ethnocultural conditionality. As another kind of nationally specific collocations, the author considers the collocations that have common structural and semantic properties in all three national variants of the German language, but differ in the design of the base or collocator. It is argued that the revealed differences are due to either inventory and semantic differences in the lexical content of the phrases, or the variability of the use of general German components. In addition, the author points out that the national features of collocations are also manifested in the specifics of the syntactic organization, especially in the use of prepositions. Thus, it is shown that the peculiarities of Austrian and Swiss collocations are due to the originality of the national variants of the German language at different levels of the language system.
بنك المشجّرات محلّل حاسوبيّ للظّواهر التّركيبيّة في اللّغة العربيّة، استثمر مبادئ نظريّة التّحكّم والرّبط التّوليديّة، وحوسباتها وتصوّراتها للنّحو الكلّيّ، غايته في ذلك بناء نظام حوسبيّ آليّ، يحاكي في اشتغاله النّظام الحوسبيّ اللّغويّ الطّبيعيّ. وقد حقّق بنك المشجّرات نتائج مهمّة في هذا الشّأن، تتمثّل في بلوغه الانتظام والتّناسق في معالجة الأبنية الإعرابيّة، لكنّ العمل لم يخل من هنات، أهمّها عدم اتّسام السّيرورة الاشتقاقيّة بالخاصّيّة التّكراريّة المميّزة للّغة البشريّة، وخرق حوسبة النّقل للقيود الجزبريّة التي أقرّتها النّظريّة اللّسانيّة، وهو ما يجعلنا نشكّك في كفايته الوصفيّة لسانيّا.
Film clips are proven to be one of the most efficient techniques in emotional induction. However, there is scant literature on the effect of this procedure in older adults and, specifically, the effect of using different positive stimuli. Thus, the aim of the present study was to examine emotional differences between young and older adults and to know how a set of film clips works as mood induction procedure in older adults, especially, when trying to elicit attachment-related emotions. To this end, we use this procedure to analyze differences in subjective emotional response between young and older adults. A sample of 57 older adults and 83 young adults watched a film set previously validated in young population. Their responses were studied in an individual laboratory session to elicit 6 target emotions (disgust, fear, sadness, anger, amusement and tenderness) and neutral state. Self-reported emotional experience was measured using the Self-Assessment Manikin (SAM). Our results show that film clips are capable of evoking positive and negative emotions in older adults. Furthermore, older adults experienced more intensely negative emotions than young adults, especially in response to disgust and fear clips. They also reported higher arousal than young adults, especially in the case of sadness, anger and tenderness clips. Nevertheless, the older adults recovered more easily from the effects of the emotion induction. The young adults reported higher arousal ratings than older adults in response to amusement film clips. On the other hand, this study reflects the importance of controlling the baseline state to study the real strength of mood induction. Overall, current data suggests significant differences occur in emotional response in adult age and that film clips are an effective tool for studying positive and negative emotions in aging research.
The article is devoted to a modern problems research of television titles editing of the media addressee by the media addressor. The objectives of work are achieved by application of methods of deductive and inductive logical analysis, descriptive method, content analysis, lexical and semantic, lexical and grammatical, and stylistic analysis, comparative analysis, deep interview and poll of informants. Titles as fragments of media texts of the television program “Time Will Show” act as material of the research. In the article, attention is paid to a problem of television titles editing of the mediaaddressee in which editorial work of the media addressor is often limited to an inscription: “The spelling and a punctuation of the author are kept”. Authors in details analyze television titles of the mass media addressee of the “Time Will Show” program broadcast on Channel 1 of the Russian television; sort a number of examples from other elements of media system subject to influence of the research object and from fiction with justified violation of language norms. Authors offer the answer to a question how the editor has to work with text elements on the screen, support the position with opinions of scientists and results of the comparative analysis of the actual material, deep interview and poll of social networks users. The received results are significant for development of psycholinguistics, pragmalinguistics, cognitive linguistics, linguosemiotics, cultural linguistics, discursive linguistics, in particular, of media discourse theory, influence theory. This is because the peculiarities of verbal self-presentation of the media addressee in television titles characterizing the language personality of the media addressee as the media addressor producing the media message as reaction to a television message of the media addressor, and making influence on the mass addressee – television audience are revealed. The article will be useful to philologists, editors, journalists not only in the theoretical plan, but also in practical work.
<em>The article is devoted to the phenomenon of the household-centred paradigm in H. F. Kvitka-Osnovianenko’s artistic heritage. The verbalization of the Slobozhany’s household hasn’t been analyzed by linguists yet; it makes this research topical. <strong>The aim of the article </strong>is to describe language means of making the concept «Houshold» objective in the lexical-phraseological system of H. F. Kvitka-Osnovianenko’s works. The aim of the article presupposes <strong>solving the following tasks</strong>: outlining the list of lexemes that define household subjects; singling out phraseological units with the «household» component; analyzing the deep culturally meaningful dominant, taking the component «khata» [village house] as an example. <strong>Conclusion. </strong>Language mapping of H. F. Kvitka-Osnovianenko’s works fits the culturally specific household-centred paradigm. In H. F. Kvitka-Osnovianenko’s works household is presented by a wide well-structured lexical-phraseological field as well as numerous proverbs that proves its communicative relevance for Ukrainian, the Slobozhany’s in particular, consciousness. FSF «Household» includes such FSG as «Customs», «Way of Life», «Work», «Household utensils», «Housekeeping» and others. These FSGs contain subgroups (FSSG), microgroups (FSMG) may be realized in the individual, family, social or territorial (urban, rural) aspects of life. Their prototypic representatives are customs, tradition, norms and order, living conditions characteristic for one person, a social group or the society as a whole.</em>
Typing is a ubiquitous daily action for many individuals; yet, research on how these actions have changed our perception of language is limited. One such influence, deemed the QWERTY effect, is an increase in valence ratings for words typed more with the right hand on a traditional keyboard (Jasmin &amp; Casasanto, 2012). Although this finding is intuitively appealing given both right-handed dominance and the smaller number of letters typed with the right hand, extension and replication of the right-side advantage is warranted. The present paper reexamined the QWERTY effect expanding to other embodied cognition variables (Barsalou, 1999). First, we found that the right-side advantage is replicable to new valence stimuli. Further, when examining expertise, right-side advantage interacted with typing speed and typeability (i.e., alternating hand key presses or finger switches) portraying that both skill and our procedural actions play a role in judgment of valence on words.
The article is focusing on the problems of standards, the peculiarities of functioning of modern French language, such as situational variability which influences the choice of certain language unit according to a situation of the communicative act. This research is relevant since the essential value is given to the study of the language norm at all levels of language organization – grammatical, lexical, phonetic. Together with development of the speech, with emergence of new features in it, its norm is modified as well. In other words, the norm needs to be considered not in a statics, but in dynamics, in view of “dynamic aspect of norm”. In linguistics the term “norm” is used in two ways – broad sense and narrow. In a broad sense, the norm is a traditionally and spontaneously established methods of language that distinguish one language system from the others (in this sense, the norm is close to the concept of “usage”, i.e. generally accepted, well-established ways of using this language). In the narrow sense, the norm is the result of purposeful codification of a language. Such understanding of norm is closely connected with a concept of the literary language which otherwise could be called normalized or codified language.
This study examines Swedish morning TV’s framing of the phenomenon of exercising. Morgonstudion in SVT, and Nyhetsmorgon in TV4, is the morning shows that has been investigated. The aim of the study is to enlighten and enhance the understanding of how the phenomenon of exercising are framed in Swedish morning television, as well as to contribute to the theorization of media’s representation of exercise. A qualitative content analysis has been used to capture the language, to see how they present exercising and if they legitimate it. Theoretical framework applied are framing theory, representation, legitimize, healthism and public service vs commercial television. The research fields are health communication, exercise in media and morning television journalism. The result shows that Swedish morning television, through different methods and lexical choices, legitimize the phenomenon of exercise. SVT is more focused on exercise that suits everyone and that viewers can change their lifestyle by making small changes in their everyday life. For example, they have an idea of how to get the pulse up by exercises in the garden while TV4 is aiming at reaching out to those that is already exercising and specifies types of exercising during the show. They give advice on how to find out what training tools you need to complete a triathlon for example. Lexical choices reinforce that exercise is something positive and both the guests and hostess sees exercise as a norm. The elements of exercise differ between the programs, as well as the studio environment and the content. On the other hand, there are similarities such as the language that is used and they are both visited by experts.
We unify recent neural approaches to one-shot learning with older ideas of associative memory in a model for metalearning. Our model learns jointly to represent data and to bind class labels to representations in a single shot. It builds representations via slow weights, learned across tasks through SGD, while fast weights constructed by a Hebbian learning rule implement one-shot binding for each new task. On the Omniglot, Mini-ImageNet, and Penn Treebank one-shot learning benchmarks, our model achieves state-of-the-art results.
The article considers variants of paronyms compatibility and some mistakes of Turkmen students in the use of paronyms – words closed in pronunciation but not identical, and sometimes different in meaning. The vocabulary of foreign students entering the university does not always correspond to the needs of their language practice, often there is no knowledge of the paronyms necessary for it. The study of various paronyms combining variants with other words, semantic connections are as necessary as knowledge of grammatical rules and spelling norms. The correct using of the word in speech suggests, firstly, knowledge of the word structure and the lexical-semantic variants of the polysemy; secondly, analysis of the word-building composition with the individual morphemes meaning identification; thirdly, the ability to choose the word needed for a given context, for which it is necessary to determine its place in the lexical-semantic system, i.e. to find its connection with other words, to identify general and differential features in a given lexical group making up a certain semantic unity. As a result of observations of Turkmen students“ oral and written speech, the main reasons for the erroneous paronyms substitution are revealed, involving roots consonance. The same root paronyms are also confused because of inaccurate understanding of the prefixes, suffixes meaning and difference they bring to the word meaning. Considering the same root-words, there are many problems due to the fact that they are heterogeneous in the semantic and word-formation aspect. Paronyms like the same-root words, similar to synonyms. They have a number of common features which leads to the similarity of these two language units, and finally to the error occurrence while using paronyms in speech. The article also contains examples of training exercises for speech skills developing in the paronyms using at Russian language classes in the Turkmen audience.
The article is focused on the principal theoretical and methodological concepts related to the education of children with speech impairment in the conditions of the inclusive environment of elementary school. In particular, the author analyzes the following terminological concepts: language, speech, norms of speech, disorders in oral and written speech, anomalies in speech development, phonetic and phonemic speech underdevelopment, general speech underdevelopment, stuttering, lexical andgrammatical aspect of speech. It is noted that teachers, educators and teachers’ assistants as the main specialists who should create appropriate conditions for inclusiveeducation for children with speech impairment should be able to distinguish the norms for children’s speech from their disorders, to differentiate types of speech development disorders on the basis of psychological and pedagogical classification, to predict the possible difficulties in teaching children with special educational needs related to the formation of reading, writing and mathematical skills. Besides, the article provides with short characteristics and signs of the above types of speech development disorders in children, which will help primary school teachers with an inclusive form of training to choose appropriate methods and means of teaching children with speech impairment during the organization of educational process.
The article deals with the of terms with international components in Ukrainian language. The essence of the term 'codification' and its connections with the concepts of 'standardization' and 'normalization' have been found out. It has also been stressed that the term codification has several interpretations. These are: 1) the systematization of norms, 2) the kind of norm-setting, and 3) the means of regulating something. Common for these meanings the written fixation, which is of a recommending nature. Standardization means that only one term must survive, and any other is deleted. Normalization pertains the process in any terminology that happens on two levels: lexical and word-building. The essence of codification is the objective assessment of language doublets, variants, innovations, trends of development. Codification can be absolute and relative. Absolute is related to the established norm, since the use of many words and terms with the MC does not cause any warnings. These words got into the language, became theirs own, without them we cannot provide our daily and scientific communication. Relative is associated with tensile shaky norm. Actually, it requires a careful attitude towards invariants and variants in favor of one in them. With the adoption of Ukrainian language status as a state, the question for purpose of discussion terminology base has been raised. Today we observe a tendency to introduce purely Ukrainian terms into the sphere of scientific broadcasting. Based on the dictionaries by M.Vakulenko and P.Stepa, in which foreign words (including terms with international components) are offered in Ukrainian, we try to understand the speed of replacing terms with MC in pure Ukrainian words. Key words: Ukrainian language, terminology, words with international components,
The creation of real texts in Latin can be regarded not only as the intellectual activity of the pastime of a limited number of enthusiasts, but also a long linguistic experiment. In this study, the product of such linguistic activity serves as a source of materials, primarily lexical and derivational innovations, for analyzing events that may arise in a "restored" language system. To reveal trends in the development of linguistic material and its potential, the data obtained from new Latin texts were compared with the results of studying Latin borrowings in English (in the natural conditions of a living language). Obviously, the choice of vocabulary and new terms to denote modern realities in these new Latin news texts are subject to the preferences of individual researchers and are sometimes arbitrary to a greater extent than in the case of creating text in a naturally developing language, where the speaker / user strongly dominates the usage norms. In addition, when developing innovations, the authors of the texts in question inevitably follow their native language of the L2 experience. As a result of innovation, the "New Latin Scheme" shows some features more typical of modern European languages, in addition, the main development trends in the group of Latin borrowings in English were different from those found in the lexicon of the new Latin.
We demonstrate that replacing an LSTM encoder with a self-attentive architecture can lead to improvements to a state-of-the-art discriminative constituency parser. The use of attention makes explicit the manner in which information is propagated between different locations in the sentence, which we use to both analyze our model and propose potential improvements. For example, we find that separating positional and content information in the encoder can lead to improved parsing accuracy. Additionally, we evaluate different approaches for lexical representation. Our parser achieves new state-of-the-art results for single models trained on the Penn Treebank: 93.55 F1 without the use of any external data, and 95.13 F1 when using pre-trained word representations. Our parser also outperforms the previous best-published accuracy figures on 8 of the 9 languages in the SPMRL dataset.
Web 2.0 has brought with it numerous user-produced data revealing one's thoughts, experiences, and knowledge, which are a great source for many tasks, such as information extraction, and knowledge base construction. However, the colloquial nature of the texts poses new challenges for current natural language processing techniques, which are more adapt to the formal form of the language. Ellipsis is a common linguistic phenomenon that some words are left out as they are understood from the context, especially in oral utterance, hindering the improvement of dependency parsing, which is of great importance for tasks relied on the meaning of the sentence. In order to promote research in this area, we are releasing a Chinese dependency treebank of 319 weibos, containing 572 sentences with omissions restored and contexts reserved.
Annotation corpus for discourse relations benefits NLP tasks such as machine translation and question answering. In this paper, we present SciDTB, a domain-specific discourse treebank annotated on scientific articles. Different from widely-used RST-DT and PDTB, SciDTB uses dependency trees to represent discourse structure, which is flexible and simplified to some extent but do not sacrifice structural integrity. We discuss the labeling framework, annotation workflow and some statistics about SciDTB. Furthermore, our treebank is made as a benchmark for evaluating discourse dependency parsers, on which we provide several baselines as fundamental work.
Cognitive variation due to language and culture has been shown in a range of domains, including visual perception,emotions, theory of mind, economic strategies, decision making, and categorization. While such patterns are robust,individuals within a given culture are affected by these cultural patterns differentially. One possible cause for theseindividual differences is personality (e.g. extroversion or agreeableness). The personality traits of individuals will affecthow they interact with and adopt cultural patterns. To explore this possibility, we perform analyses on online data fromindividuals with self-identified Myers-Briggs personality types (a popularized personality measure that is widely self-reported in social media). In particular, we examine how personality type predicts the rate at which individuals adopt novellexical items and conform to the linguistic norms of their surrounding community. The results make explicit predictionsabout which individuals will be more affect by cultural and linguistic patterns.
Anger is considered a unique high-arousal and approach-related negative emotion. The influence of individual differences in trait anger on the processing of visual stimuli is relevant to questions about emotional processing and remains to be explored. Using functional magnetic resonance imaging (fMRI), we explored the neural responses to standardized images, selected based on valence and arousal ratings in a group of men with high trait anger compared to those with normative to low anger scores (controls). Results show increased activation in the left-lateralized ventral fronto-parietal attention network to unpleasant images by individuals with high trait anger. There was also a group by arousal interaction in the left thalamus/pulvinar such that individuals with high trait anger had increased pulvinar activation to the high-arousal (versus low arousal) unpleasant images as compared to controls. Thus, individual differences in trait anger in men are associated with brain regions subserving executive attentional and sensory integration during the processing of unpleasant emotional stimuli, particularly to high arousal images.
This paper describes a method of creating synthetic treebanks for cross-lingual dependency parsing using a combination of machine translation (including pivot translation), annotation projection and the spanning tree algorithm. Sentences are first automatically translated from a lesser-resourced language to a number of related highly-resourced languages, parsed and then the annotations are projected back to the lesser-resourced language, leading to multiple trees for each sentence from the lesser-resourced language. The final treebank is created by merging the possible trees into a graph and running the spanning tree algorithm to vote for the best tree for each sentence. We present experiments aimed at parsing Faroese using a combination of Danish, Swedish and Norwegian. In a similar experimental setup to the CoNLL 2018 shared task on dependency parsing we report state-of-the-art results on dependency parsing for Faroese using an off-theshelf parser.
Abstract. Affective science calls for methods to induce mood in an engaging and ecologically valid way. We present a method employing a naturally occurring scenario that fits these criteria: a job interview. Participants got positive or negative feedback from a fictive expert to induce positive or negative mood. After mood induction, we assessed participants’ decision making behavior in the so-called information sampling task (IST). Results show that our mood induction successfully changed valence, dominance, and state self-esteem ratings, while there were no differences in arousal ratings. Decision making in the IST was not influenced by the induced mood. Effect sizes of mood induction were equally high for positive and negative mood concerning valence ratings (d =.8) with participants scoring high on self-control showing smaller mood induction effects. We conclude that our mood induction technique is an effective and natural way to induce mood in the laboratory, meeting current criteria of affective science.
[full article and abstract in Russian; abstract in English]
 This article investigates the informal verbalizations of urban place names; it also questions the hypothesis of their potential untranslatability. The concept of slang is based on the assumption that slang acts as a linguistic channel for liberating carnivalesque laughter. The most numerous lexical fields (presumably most important for the carnivalesque worldview) are those denoting entertainment establishments (brothels, night clubs, cafes), prisons and hospitals. Humankind resorts to substandard communication to state the predominance of primeval instincts, a negative relation to the dominant value system with its restrictions. Mocking values and norms, physiological deficiency, gender or racial peculiarities is the essence of slang as a constituent of humour culture. Slang is a revolt against hierarchy. It relieves tension without endangering the stability of the society. Lexical semantic analysis of slang toponyms shows the potential possibility of adequate comparing and translating of place names in both directions.
The Effective Set-Size model has been used to describe uncertainty in various signal detection experiments. The model regards images as if they were an effective number (M*) of searchable locations, where the observer treats each location as a location-known-exactly detection task with signals having average detectability d'. The model assumes a rational observer behaves as if he searches an effective number of independent locations and follows signal detection theory at each location. Thus the location-known-exactly detectability (d') and the effective number of independent locations M* fully characterize search performance. In this model the image rating in a single-response task is assumed to be the maximum response that the observer would assign to these many locations. The model has been used by a number of other researchers, and is well corroborated. We examine this model as a way of differentiating imaging tasks that radiologists perform. Tasks involving more searching or location uncertainty may have higher estimated M* values. In this work we applied the Effective Set-Size model to a number of medical imaging data sets. The data sets include radiologists reading screening and diagnostic mammography with and without computer-aided diagnosis (CAD), and breast tomosynthesis. We developed an algorithm to fit the model parameters using two-sample maximum-likelihood ordinal regression, similar to the classic bi-normal model. The resulting model ROC curves are rational and fit the observed data well. We find that the distributions of M* and d' differ significantly among these data sets, and differ between pairs of imaging systems within studies. For example, on average tomosynthesis increased readers’ d' values, while CAD reduced the M* parameters. We demonstrate that the model parameters M* and d' are correlated. We conclude that the Effective Set-Size model may be a useful way of differentiating location uncertainty from the diagnostic uncertainty in medical imaging tasks.
International audience
Previous studies have demonstrated differential perception of body expressions between males and females. However, only two recent studies (Kret et al., 2011; Krüger et al., 2013) explored the interaction effect between observer gender and subject gender, and it remains unclear whether this interaction between the two gender factors is gender-congruent (i.e., better recognition of emotions expressed by subjects of the same gender) or gender-incongruent (i.e., better recognition of emotions expressed by subjects of the opposite gender). Here, we used event-related potentials (ERPs) to investigate the recognition of fearful and angry body expressions posed by males and females. Male and female observers also completed an affective rating task (including valence, intensity, and arousal ratings). Behavioral results showed that male observers reported higher arousal rating scores for angry body expressions posed by females than males. ERP data showed that when recognizing angry body expressions, female observers had larger P1 for male than female bodies, while male observers had larger P3 for female than male bodies. These results indicate gender-incongruent effects in early and later stages of body expression processing, which fits well with the evolutionary theory that females mainly play a role in care of offspring while males mainly play a role in family guarding and protection. Furthermore, it is found that in both angry and fearful conditions male observers exhibited a larger N170 for male than female bodies, and female observers showed a larger N170 for female than male bodies. This gender-incongruent effect in the structural encoding stage of processing may be due to the familiarity of the body configural features of the same gender. The current results provide insights into the significant role of gender in body expression processing, helping us understand the issue of gender vulnerability associated with psychiatric disorders characterized by deficits of body language reading.
High quality communication between health care providers (HCPs) and adolescents and young adults (AYAs) with type 1 diabetes (T1D) may contribute to better diabetes self-care and health outcomes. Health communication reflects both informational content and how information is conveyed, including affect and tone. The aim of this study was to assess HCP affective communication and the relationship between HCP affective communication and glycemic control in AYAs with T1D. As part of a larger study of AYA-HCP health communication, routine clinic visits for 69 AYAs with T1D (M age 17.81 years; 56.5% female) and 8 HCPs (88% female) were audiorecorded. Clinic visits were coded using the Roter Interaction Analysis System (RIAS), a validated coding structure assessing verbal and non-verbal exchanges in a medical encounter. HCP global affective ratings were used to create two composite variables—positive HCP affect (e.g., attentiveness; respectfulness; Cronbach’s a = 0.82) and negative HCP affect (e.g., anger; dominance; Cronbach’s a = 0.75). Hemoglobin A1c (A1c) was taken from the medical chart. The mean A1c was 8.97% (±2.30). Descriptive analyses of positive and negative HCP affect indicated that HCPs expressed a high level of positive affect (M = 4.21) and a relatively low level of negative affect (M = 2.80). Negative affect was positively associated with HbA1c. After controlling for salient covariates (e.g., HCP, race, regimen), A1c accounted for a significant portion of the variance in negative affect during the clinic visit (Adj R2 =.36, ß = 0.57, p &lt; 0.001). This sample of HCPs predominantly exhibited positive affect during routine T1D visits. Glycemic control was not associated with positive affect, but higher A1c was associated with more negative affect. This finding suggests elevated A1c levels may elicit more negative affect in routine diabetes care. Future research should examine these associations over time, including how AYA-HCP health communication quality predicts long-term glycemic control. Disclosure K. Homma: None. F.R. Cogen: None. R. Streisand: None. M. Monaghan: Research Support; Self; American Diabetes Association, National Institutes of Health.
Świgra is a parser of Polish generating constituency trees using a DCG style grammar stemming from Marek Świdzinski’s grammar “Gramatyka formalna jezyka polskiego” (1992). The grammar was heavily rewritten for the purpose of annotating the Skladnica treebank. The structure of trees was simplified with respect to Świdzinski’s version, many new types of constructions were included (in particular various forms of coordinated structures), a statistical disambiguating component was added. Moreover, the Clarin version of Świgra uses the valency dictionary Walenty developed within Clarin.
Please join the LAII and University Libraries for a presentation with Richard E. Greenleaf Visiting Library Scholar Jonathan Steuck, a doctoral candidate in Hispanic Linguistics and Language Science at Penn State University. Steuck’s main research interests include code-switching and bilingualism, the prosodic-syntactic interface, and dialectical variation of intonation. Much of his work focuses on Spanish and English in the US. Steuck will discuss his research findings regarding the tendency of speakers of Spanish and English in New Mexico to fluidly alternate between languages in the same conversation (i.e., code-switch). Research suggests that this is a skilled behavior, reflective of a high degree of language proficiency in both languages and the linguistic norms of a speech community. Recent studies have also found that bilinguals may utilize the phonetic features and quantitative patterns present in code-switching to anticipate an upcoming language switch (e.g. Fricke et al., 2016; Tamargo et al. 2016). Previous studies are constrained, however, by the exclusion of factors at the interface of prosody and syntax. This talk will help address this issue by comparing intra-sentential multi-word code-switches (MWCS) with i) English-only and Spanish-only prosodic sentences (see Chafe, 1994) from the same bilingual speakers and ii) Spanish prosodic sentences from more monolingual speakers of an older variety. A consideration of pause expression (i.e. no pause versus (un)filled pause(s)), length measures (e.g. words and seconds), and other prosodic-syntactic factors will illuminate the extent to which MWCS differ from speech where code-switches are absent. Data derive from the New Mexico Spanish-English Bilingual corpus (Torres Cacoullos & Travis, in prep) and the New Mexico-Colorado Spanish Survey (Bills & Vigil 2008). While MWCS are found to have prosodic-syntactic properties that are distinct from Spanish-only and English-only productions, MWCS are no less fluid than speech produced in only one language. Overall, this provides new insight into where code-switches tend to occur in natural discourse and indicates that the previously proposed constraints of code-switching will need to be reexamined in tandem with prosodic junctures. These findings also set the stage for experimental research exploring the prosodic-syntactic cues speakers may utilize to comprehend code-switches. Each year the LAII partners with University Libraries to offer Richard E. Greenleaf Visting Library Scholar awards to support scholars who work with UNM’s nationally-acclaimed Latin American library holdings. The award honors Dr. Richard E. Greenleaf, distinguished scholar of colonial Latin American history, and his extensive career in teaching, research, and service. Event sponsored by Latin American and Iberian Institute, University Libraries.