Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Text segmentation is a fundamental task in natural language processing. Depending on the levels of granularity, the task can be defined as segmenting a document into topical segments, or segmenting a sentence into elementary discourse units (EDUs). Traditional solutions to the two tasks heavily rely on carefully designed features. The recently proposed neural models do not need manual feature engineering, but they either suffer from sparse boundary tags or cannot efficiently handle the issue of variable size output vocabulary. In light of such limitations, we propose a generic end-to-end segmentation model, namely <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula>, which first uses a bidirectional recurrent neural network to encode an input text sequence. <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> then uses another recurrent neural networks, together with a pointer network, to select text boundaries in the input sequence. In this way, <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> does not require any hand-crafted features. More importantly, <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> inherently handles the issue of variable size output vocabulary and the issue of sparse boundary tags. In our experiments, <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula> outperforms state-of-the-art models on two tasks: document-level topic segmentation and sentence-level EDU segmentation. As a downstream application, we further propose a hierarchical attention model for sentence-level sentiment analysis based on the outcomes of <inline-formula><tex-math notation="LaTeX">${\mathrm{S}\scriptstyle{\mathrm{EG}}}{\mathrm{B}\scriptstyle{\mathrm{OT}}}$</tex-math></inline-formula>. The hierarchical model can make full use of both word-level and EDU-level information simultaneously for sentence-level sentiment analysis. In particular, it can effectively exploit EDU-level information, such as the inner properties of EDUs, which cannot be fully encoded in word-level features. Experimental results show that our hierarchical model achieves new state-of-the-art results on the Movie Review and Stanford Sentiment Treebank benchmarks.
In linguistic pragmatics norms can be seen as traditions that guide verbal interaction. In order to pin down the notion of tradition, we use a model of linguistic pragmatics that goes back to Eugenio Coseriu's system of linguistic competence and to the concept of tradition elaborated by Ramon Menendez Pidal, thus bringing together linguistics and philology. The functioning of norms as traditions is illustrated with two examples: with a routine of verbal politeness and with a narration style that is based on the aspect system of Romance languages and functions as a cultural tradition.
This chapter explains discourse analytical and linguistic approach to the study of the linguistic norm. According to this theoretical viewpoint, the standard is seen as the factual linguistic result obtained by means of an evaluative discourse regarding language use. Discourse linguistics tries to uncover the hidden mechanisms of power and persuasion by which certain agents (persons, institutions) manage to impose their linguistic ideology on the speech community. Hence, there are many points of contact between discourse linguistics and sociolinguistics, language planning theory and the history of linguistic thought.
Abstract Two research traditions explain the way we deal with emotional situations: emotional intelligence (EI) and emotion regulation (ER). EI refers to the individual differences in the knowledge, identification, and regulation of emotions. ER describes processes in which emotions are experienced, expressed, and altered. Our study examined the EI-ER link and their moderating role on affective responses. We used self-report questionnaires and a cognitive reappraisal (CR) task, in which subjective affective responses were registered. We found that higher levels of ER difficulties correlated with lower EI. Gender had an overall impact on affective changes, indicating a more unpleasant and more arousing affective state for women compared with men. Regarding the moderating role of EI and ER difficulties, the ability to utilize emotions (Utilization) decreased the valence into a more unpleasant direction, similar to the effect of the inability to identify and differentiate emotions (Clarity). A weak control over emotions (Impulse), however, increased the valence into a more pleasant direction. The lack of attention to emotional signals (Awareness) marginally decreased the initial intensity (i.e., lower level of arousal). We demonstrated that EI and ER have distinctive routes and a different influence on the affective outcome defined by valence and arousal ratings: (1) EI has an impact through the utilization of emotions mainly on the valence dimension; and (2) individual differences in ER have a moderating effect on both valence and arousal dimensions. This study provided evidence on how individual differences contribute to a successful ER process when using a CR strategy.
Treebanks are an essential resource for syntactic parsing. The available Paninian dependency treebank(s) for Telugu is annotated only with inter-chunk dependency relations and not all words of a sentence are part of the parse tree. In this paper, we automatically annotate the intra-chunk dependencies in the treebank using a Shift-Reduce parser based on Context Free Grammar rules for Telugu chunks. We also propose a few additional intra-chunk dependency relations for Telugu apart from the ones used in Hindi treebank. Annotating intra-chunk dependencies finally provides a complete parse tree for every sentence in the treebank. Having a fully expanded treebank is crucial for developing end to end parsers which produce complete trees. We present a fully expanded dependency treebank for Telugu consisting of 3220 sentences. In this paper, we also convert the treebank annotated with Anncorra part-of-speech tagset to the latest BIS tagset. The BIS tagset is a hierarchical tagset adopted as a unified part-of-speech standard across all Indian Languages. The final treebank is made publicly available.
Text discourse parsing plays an important role in understanding information flow and argumentative structure in natural language. Previous research under the Rhetorical Structure Theory (RST) has mostly focused on inducing and evaluating models from the English treebank. However, the parsing tasks for other languages such as German, Dutch, and Portuguese are still challenging due to the shortage of annotated data. In this work, we investigate two approaches to establish a neural, cross-lingual discourse parser via: (1) utilizing multilingual vector representations; and
This article is an investigation of teachers’ corrections and comments on graduation essays from the 1870s. Our aim is to learn more about the linguistic norms in Swedish secondary schools at the time, such as they appear in this specific context. The study is based on 13 graduation essays (in total 7,184 words) from one secondary school (Högre Allmänna Läroverket i Falun). There are 532 corrections and comments in the texts, corresponding to an average of 7.4 corrections/comments per 100 words. Corrections and comments are thus very common in the texts that we have investigated. Many of the corrections focus on textual details, such as the choice of words and sentence structure, but by no means are all of the corrections due to incorrect language use. Rather, it seems that many of the corrections target the style of essays. Corrections or comments concerning the structure of the texts are very rare. Most of the comments made by teachers are in the category called focusing, i.e. co-appearing with a correction in the text. None of the essays contain positive feedback from the teacher, which we believe could be significant for the correction norm at the time.
Dans cet article, nous proposons un modele de representations vectorielles de paire de mots, obtenues a partir d’une adaptation du modele Skip-gram de Word2vec. Ce modele est utilise pour generer des vecteurs de paires de verbes, entrainees sur le corpus de textes anglais Ukwac. Les vecteurs sont evalues sur les donnees ConceptNet & EACL, sur une tâche de classification de relations lexicales. Nous comparons les resultats obtenus avec les vecteurs paires a des modeles utilisant des vecteurs mots, et testons l’evaluation avec des verbes dans leur forme originale et dans leur forme lemmatisee. Enfin, nous presentons des experiences ou ces vecteurs paires sont utilises sur une tâche d’identification de relation discursive entre deux segments de texte. Nos resultats sur le corpus anglais Penn Discourse Treebank, demontrent l’importance de l’information verbale pour la tâche, et la complementarite de ces vecteurs paires avec les connecteurs discursifs des relations.
Mindful meditation, an exercise which encourages its practitioners to be present in the moment and to be aware of their current emotions, thoughts, and sensations, has been shown to affect the processing of emotional information (Sobolewski et al., 2011) and to increase empathy (Tan, Lo, & Macrae, 2014). The prefrontal cortex has been implicated in these processes (Seitz, Nickel, & Azari, 2006). We sought to investigate the neural mechanisms which underlie how mindful meditation affects emotional processing and to determine whether any changes in brain activity could be linked to changes in empathy. Participants in our experimental group practiced mindful meditation for ten minutes. Those in the control groups either listened to an unexciting newscast or sat quietly for ten minutes. Next, participants in all conditions viewed a series of emotionally valenced images from the International Affective Picture System (IAPS) (Lang, Bradley, & Cuthbert, 2008). As they viewed the images, activity in the prefrontal cortex was monitored with functional near-infrared spectroscopy (fNIRS). Following the presentation of the IAPS images, participants completed questionnaires and inventories that measured empathy, previous experience with meditation, and personality traits. As well as differences in empathy between participants who meditated and participants who did not, we expect to discover any variations in patterns of prefrontal cortical activity during the viewing of positive, negative, and neutral imagery. We hope that our results will elucidate how mindful meditation affects emotional processing and the development of empathy, and to determine whether certain personality traits, social conservatism, or lack of sleep may predict neural and behavioral responses to mindful meditation. Lang, P.J., Bradley, M.M., & Cuthbert, B.N. (2008). International affective picture System (IAPS): Affective ratings of pictures and instruction manual. Technical Report A-8. University of Florida, Gainesville, FL. Seitz, R.J., Nickel, J., & Azari, N.P. (2006). Functional modularity of the medial prefrontal cortex: Involvement in human empathy. Neuropsychology, 20(6), 743-751. Sobolewski, A., Holt, E., Kublik, E., & Wróbel, A. (2011). Impact of meditation on emotional processing—A visual ERP study. Neuroscience Research, 71(1), 44–48. doi: 10.1016/j.neures.2011.06.002 Tan, L.B.G., Lo, B.C.Y., & Macrae, C.N. (2014). Brief Mindfulness Meditation Improves Mental State Attribution and Empathizing. PLoS ONE, 9(10). doi: 10.1371/journal.pone.0110510
The COVID-19 crisis resulted in a large proportion of the world's population having to employ social distancing measures and self-quarantine. Given that limiting social interaction impacts mental health, we assessed the effects of quarantine on emotive perception as a proxy of affective states. To this end, we conducted an online experiment whereby 112 participants provided affective ratings for a set of normative images and reported on their well-being during COVID-19 self-isolation. We found that current valence ratings were significantly lower than the original ones from 2015. This negative shift correlated with key aspects of the personal situation during the confinement, including working and living status, and subjective well-being. These findings indicate that quarantine impacts mood negatively, resulting in a negatively biased perception of emotive stimuli. Moreover, our online assessment method shows its validity for large-scale population studies on the impact of COVID-19 related mitigation methods and well-being.
The paper presents a short introduction to several electronic resources for Ukrainian language, namely, two treebanks: the Gold standard (ab. 130 thousand tokens), manually annotated in the Universal Dependencies flavour (https://universaldependencies.org/), which comprises the training data for a machine-trained syntactic parser, and a big (near 3 billion tokens),
Meishan Zhang, Yue Zhang, Guohong Fu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Reference corpus for word alignment is an important resource for developing and evaluating word alignment methods. For Myanmar-English language pairs, there is no reference corpus to evaluate the word alignment tasks. Therefore, we created the guidelines for Myanmar-English word alignment annotation between two languages over contrastive learning and built the Myanmar-English reference corpus consisting of verified alignments from Myanmar ALT of the Asian Language Treebank (ALT). This reference corpus contains confident labels sure (S) and possible (P) for word alignments which are used to test for the purpose of evaluation of the word alignments tasks. We discuss the most linking ambiguities to define consistent and systematic instructions to align manual words. We evaluated the results of annotators agreement using our reference corpus in terms of alignment error rate (AER) in word alignment tasks and discuss the words relationships in terms of BLEU scores.
Background \n \nBack pain originating from multiple spinal disorders is the leading cause of disability worldwide. Adult spinal deformity (ASD) is a complex three-dimensional entity of multiple anatomical and functional disorders and predisposes to decreased health-related quality of life (HRQoL). With rising life expectancy and a growing elderly population it is predicted that an increasing number of symptomatic ASD patients will need surgical treatment. It is noteworthy that the prevalence of different grades of ASD, impact on HRQoL and patient-reported outcome (PRO) of ASD surgery has not earlier been reported in the Finnish population. \n \nThe aims of this study were to assess the reliability and repeatability of radiographic diagnostic imaging of ASD and to produce a culturally adapted and valid Finnish version of the Scoliosis Research Society (SRS) Questionnaire version 30 for spinal deformities. Thereafter the prevalence of ASD, applicability of a simplified version of the SRS-Schwab ASD classification and the Finnish SRS-30 were evaluated in a symptomatic adult patient cohort with prolonged degenerative spinal disorders. Finally, the long-term outcome, complications, patient satisfaction and predictive factors for poor outcome of ASD surgery, were investigated. \n \nMethods \n \nOver one year, a consecutive cohort of adult patients was recruited to Studies I-III after referral to the Central Hospital of Central Finland spine clinic due to prolonged degenerative spinal disease. 637 patients returned the completed HRQoL questionnaires and digital full spine radiographs were obtained. The radiographs of 49 patients were randomly selected for the reliability assessment, and a repeatability study of sagittal spinopelvic measurements with basic software tools was performed by three raters differing in their experience of image rating. The SRS-30 underwent translation and cross-cultural adaptation into Finnish and was subsequently validated and psychometrically tested among 274 patients. The SRS-Schwab ASD classification was graded and simplified dividing into mild, moderate and marked groups. The division was tested along with the Finnish SRS-30 questionnaire during evaluation of the prevalence and HRQoL of patients with sagittal malalignment among symptomatic adult patients with spinal degenerative disease but no pre-known deformity. The 79 patients in Study IV were operated during 2007-2016 in our clinic. The clinical and radiographic outcome, patient satisfaction, predictive factors for poor outcome and complications were analysed using the diagnostic tools renovated and tested in Studies I-III. \n \nResults \n \nThe intra-and interrater reliability of the sagittal spinopelvic measurements proved reliable and repeatable with intraclass correlation coefficients (ICC) between 0.78-0.99 and standard error of measurement (SEM) of 0.80-6.2° or 2.2-5.8mm. Greater rater experience in performing the radiographic measurements decreases, and greater complexity of the measurement landmarks increases, intra- and inter-rater bias. \n \nThe reproducibility and internal consistency (ICC 0.905, SEM 0.17, Cronbach α 0.885) of the Finnish version of the SRS-30 was good. The SRS-30 had discriminative validity in the pain, self-image and satisfaction with management domains compared with other questionnaires. A statistically significant difference between the moderate and marked deformity groups in the SRS-30 domains of function/activity (p=0.022) and self-image/appearance (p=0.016) was found. \n \nOf the 637 patients in the consecutive cohort, 25% had moderate and 11% marked spinal deformities. The patients with marked deformity were significantly older, more overweight and more physically inactive than the others in the study population. The 3-class categorization of the SRS-Schwab ASD classification determined well the severity of sagittal deformity and concomitant loss of function, activity (p=0.004), and self-image/appearance (p=0.030) measured with the SRS-30, and disability with the ODI (p=0.033). \n \nASD operation decreased disability (ODI) and pain (VAS) significantly (p=0.001). Postoperative improvement in radiographic sagittal parameters was significant and maintained at 4-5 years of follow-up (p≤0.001). The mechanical failure of instrumentation of bone resulted in reoperation risk of 13.9% within the first and 29.8% during the 5-year follow-up. According to SRS-30, 49 (62.0%) patients were satisfied or very satisfied with the treatment and 57 (72.1%) would have the same operation again. Depression predicted poor outcome with an odds ratio of 6.97 (p=0.018). \n \nConclusions \n \nThe study comprised an unselected consecutive cohort of adult patients with prolonged degenerative spinal diseases, and thus the results can be generalized. Rater experience had a positive influence on the otherwise good reliability and repeatability of the spinopelvic measurements taken from full spine radiographs. The deformity-specific Finnish SRS-30 translation proved reliable and valid among the study cohort. The simplified categories of the SRS-Schwab ASD classification can detect different grades of deformity and related loss of HRQoL. Long-term radiographic and patient-reported clinical outcomes after the ASD surgery remained significantly better than preoperative scores. Risk for reoperation was highest during the first postoperative year. However good patient satisfaction and outcomes could be achieved irrespective of adverse effects. Depression was the only significant predictive factor for poor outcome after ASD surgery. \n \n \n \nKeywords: adult spinal deformity, ASD, scoliosis, kyphosis, full-spine radiograph, reliability, repeatability, validation, outcome, health-related quality of life, Scoliosis Research Society questionnaire 30, SRS-30, SRS-Schwab ASD classification, spine surgery, pelvic incidence, pelvic tilt, sagittal vertical axis, lumbar lordosis, thoracic kyphosis
The article considers the modernstate of computer (electronic) lexicography, which is traditionally divided into corpus and electronic ones. At first glance theincreasing global computerization has greatly facilitated the work of lexicographers and linguists, however, a number ofproblems have arisen at once. Among the first ones are the principles of corpora compilation, the requirements of which haveconsiderably expanded since the appearance of the first ones, and at present stage of corpus technologies development corporashould be able to answer a wide range of possible inquiries and meet the needs of different users. Most of the principlesdeveloped up till now are standardized and unified for the convenience of both scientific and non-scientific research. The nextkey issue that exists nowadays is not only the inclusion of the main lexicographic postulates, but also consideringtechnological and operational features of the new media (smartphones and tablets). Provided practical lexicography isprimarily intended to meet the needs of the end user, the procedures for determining the needs of users are becomingincreasingly relevant today by involving them in surveys and analysis of logs. It is noted that more and more non-specialists ofthe field are involved in the actual development of lexicographic projects. They perform simple tasks, mainly as volunteers.Such practices also require well-balanced and clear rules in order to avoid creating a low-quality final product. There exist afew online platforms where you can “hire” a number of volunteers to perform such tasks. Another aspect of the developmentof new lexicography, mainly in the English-speaking world, is a special type of lexicography - lexicography for fun. The mainobject of research here is the vocabulary associated with certain hobbies, books, films or computer games. The articleemphasizes the importance and relevance of this subject with anthropocentric cognitive paradigm of modern linguistics. Thestate of research of this phenomenon abroad and in Ukraine is clarified and disclosed. The novelty of the work is to analyzethe current state of research, development and methods of organization of lexicographic work, as well as to indicate theperspectives of its development in Ukraine. References Burkhanov, Igor. 1999. Linguistic Foundations of Ideography: Semantic Analysis and IdeographicDictionaries. Rzeszow: WSP. Cibej, Jaka, Darja Fise, Kosem Iztok. “The role of crowdsourcing in lexicography” (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015). Down, Ellie. 2005. The Unofficial Guide to Harry Potter. Chichester: Summersdale Publishers Ltd. Gao, Yongwei “The Appification of Dictionaries: From a Chinese Perspective”. (paper presented at eLex 2013 conference in Tallinn (Estonia) from 17 to 19 October 2013). Garside Roger, Geoffrey Leech, and Tony McEnery. 1997. Corpus Annotation: Linguistic Information from Computer Text Corpora. London: Longman. Gouws, Rufus Hjalmar, Ulrich Heid, Wolfgang Schweickard, et al. 2013. An International Encyclopedia ofLexicography. Supplementary. Volume: Recent Developments with Focus on Electronic and Computational Lexicography. Berlin; Boston: De Gruyter Mouton. Granger, Sylviane. 2012. “Introduction: Electronic lexicography-from challenge to opportunity”. Electronic Lexicography, edited by Sylviane Granger, and Magali Paquot. Oxford: Oxford University Press. Hidalgo, Pablo. 2017. Star Wars: The Last Jedi. The Visual Dictionary. N. Y.: DK Publishing. Holmer Louise, Martens von, Monica, and Skoldberg Emma. Making a dictionary app from a lexical database: the case of the Contemporary Dictionary of the Swedish Academy. (paper presented at eLEX 2015 conference in Herstmonceux Castle (UK) from 11 to 13 August 2015) Horot, Yevheniya, Lesya Malimon. 2015. “Suchasna leksykohrafiya: problemy y perspektyvy”. Aktualnipytannya inozemnoyi filolohiyi 3: 42–48. Danchevska, Yuliya. 2014. “Korpusy tekstiv u linhvodydaktytsi: zdobutky ta perspektyvy”. Novapedahohichna dumka 1: 58–60. Dubichynskyy, Volodymyr. 2013. “Ukraynskaya leksykohrafyya: istoriya i sovremennost”. Slavyanskayaleksykohrafyya. Moskva: Azbukovnik. Kulchytska, Tetyana. 1999. Ukrayinska leksykohrafiya XIX – XX st.: bibliohraf. pokazhchyk. NAN Ukrayiny: Lviv. nauk. b-ka im. V. Stefanyka. Kulchytskyy Ihor, Yulia Danchevska, and Ihor Likhnyakevych. 2013. “Deyaki aspekty stvorennya tavykorystannya paralelnykh korpusiv”. Naukovyy visnyk VNU im. Lesi Ukrayinky 20: 48–52. Kulchytskyy Ihor. 2017. “Informatsiyna tekhnolohiya poperedn’oho opratsyuvannya pryrodomovnykh tekstiv. Informatsiyni tekhnolohiyi ta vzayemodiyi”. Kulchytskyy, Ihor. 2002. Kompyuterno-tekhnolohichni aspekty stvorennya suchasnykh leksykohrafichnykh system. – K.: Nats. b-ka Ukrayiny im. V. I. Vernadskoho NAN Ukrayiny Kulchytskyy, Ihor. 2015. “Tekhnolohichni aspekty ukladannya korpusiv tekstiv”. Dani tekstovykh korpusiv u linhvistychnykh doslidzhennyakh. Lviv: Vydavnytstvo Lvivskoi politekhniky. Kupriyanov, Yevhen. 2008. “Kompyuterna leksykohrafiya yak problema suchasnoho movoznavstva(istorychnyy aspekt)”. Visnyk Kharkivskoho natsionalnoho universytetu im. V.N. Karazina 53: 12–16. Levchenko Olena, Ihor Kulchytskyy. 2013. “Tekhnolohiya peretvorennya p’atymovnoho slovnyka porivnyan u elektronnu formu”. Visnyk Natsionalnoho universytetu Lvivska politekhnika 770: 129–138. Lew, Robert. Space restrictions in paper and electronic dictionaries and their implications for the design of production dictionaries. Accessed February 25, 2019. https://repozytorium.amu.edu.pl/bitstream/10593/799/1/Lew_space_restrictions_in_paper_and_electronic_dictionaries.pdf McArthur, Tom. 1986. Worlds of Reference. Cambridge: Cambridge University Press. Meyer, Christian, Andrea Abel. 2017. “User participation in the Internet era”. The Routledge Handbook of Lexicography. Abingdon; New York: Routledge. Perebyynis, Valentyna, Viktor Sorokin. 2009. Tradytsiyna ta kompyuterna leksykohrafiya. Kyiv: Vyd. tsentr KNLU. Polyuha, Lev. 2006. “Ukrayinske slovnytstvo na perelomi tysyacholit”. Ukrayinoznavchi studiyi 6-7:17–25. Reynolds, David. 2000. Star Wars: the Visual Dictionary. The Ultimate Guide to Star Wars Characters and Creatures. N.Y.: Dorling Kindersley. Rundell, Michael. Redefining the dictionary: From print to digital. Accessed February 20, 2019.https://www.kdictionaries.com/kdn/kdn21_2013.pdf Rusanivskyy, Vitaliy, Volodymyr Shyrokov. 2002. “Informatsiyno-linhvistychni osnovy suchasnoyitlumachnoyi leksykohrafiyi”. Movoznavstvo 6: 7–48. Shyrokov, Volodymyr. 2011. Kompyuterna leksykohrafiya. Kyiv: Nauk. dumka. Syvokozova, Viktoriya. 2013. “Do pytannya periodyzatsiyi ukrayinskoyi tlumachnoyi leksykohrafiyi”. Movni i kontseptualni kartyny svitu 43: 61–62. Snizhko, Nataliya. 2017. “Nova dzherelna baza ukrayinskoyi leksykohrafiyi v systemi intehralnykhlinhvistychnykh doslidzhen”. Lyudyna. Kompyuter. Komunikatsiya. Lviv: Vydavnytstvo Lvivskoyipolitekhniky 47–51. Starko, Vasyl. 2017. “Kompyuterni linhvistychni proekty hurtu r2u: stan ta zastosuvannya”. Ukrayinka mova. 3: 86–97. Svensen, Bo. 1993. Practical Lexicography: Principles and Methods of Dictionary-Making. Translated by John Sykes and Kerstin Schofield. Oxford: Oxford University Press. Tarp, Sven. 2012. “Online dictionaries: today and tomorrow”. Lexicographica. De Gruyter. 28: 253–268. Trap-Jensen, Lars. Lexicography between NLP and Linguistics: Aspects of Theory and Practice (paperpresented at eLEX 2017 conference in Leiden (the Netherlands) from 19 to 21 September 2017.
This dissertation investigates the frequency, semantic, and functional characteristics of recurring discontinuous formulaic language in a learner corpus of argumentative and literary essays. Discontinuous sequences of words, or ‘frames’, are recurrent sequences of words that have one or more variable slots. For example, in the * of and it is * to, where the asterisks represent variable slots in the sequences of words. A corpus of English argumentative essays authored by native speakers of English, Japanese, and Spanish is analyzed using modern methods in corpus linguistics to determine which frames are used frequently. Frequent frames are compared between the L1 groups to investigate trends in structure, frequency, and variability.\nThe focus of analysis then narrows to a group of 30 recurring frames comprised of only function words, that is, function word frames such as in the * of, the * of the, and to the * that. The 30 frames are grouped based on structural characteristics, such as noun and preposition-based frames, for further analysis to better understand their semantic and functional characteristics. A lexical database is used to explore the semantic characteristics of fillers of the frames. To determine discourse functions of all instances of each of the 30 target frames, a well-known taxonomy previously applied to continuous sequences of words is adapted and applied to the present context. Discourse functions of the frames are then used as dependent variables in a multinomial logistic regression conducted with four distinct predictor variables: (1) L1, (2) proficiency level, (3) topic of essay, and (4) specific frame. The purpose of the regression is to see which predictor best accounts for discourse function fulfilled by frames from the structural groups.\nFindings from the various analyses first indicate that Japanese learners of English use function word frames at far lower rates than the L1 English and Spanish speakers. Secondly, fillers of the structural groups of function word frames tend to be abstract nouns and the frames largely serve the discourse function of intangible framing of a following noun phrase. In terms of predictor variables, the frames themselves, particularly prepositions, best predict discourse function. The results lend support to the idea that function words, despite carrying little meaning in isolation, are semantically motivated and systematically contribute meaning to larger sequences of words. Pedagogical implications of this study include the teaching of function words from a more phraseological perspective as well as highlighting the connection between frames, fillers, and discourse functions.
Objects with sharp contours are preferred less than objects with smooth contours, as sharp angles are thought to be an indicator of threat. In two experiments, we probe the link between low level visual features, such as contour curvature, and affective ratings. In Experiment 1, we used artist-traced line drawings of all images from the International Affective Picture System (IAPS) image set. We computationally extracted the contour curvature, length, and orientation statistics of all images, and explored whether these features are predictive of emotional valence scores. Our results replicate previous research, finding a significant negative relationship between high curvature (i.e., angularity) and emotional valence (p = 0.012). Additionally, we find that length is positively related to valence, such that images containing long contours are rated as more positive (p = 0.049). In Experiment 2, we composed new, content-free line drawings of contours with different combinations of length, curvature, and orientation values. Sixty-seven participants were presented with these images on Amazon Mechanical Turk (MTurk) and had to categorize them as positive or negative. A linear mixed effects model revealed that low curvature, long, horizontal contours predicted participants’ positive responses, while short, high curvature contours predicted participants’ negative responses. Taken together, these findings have implications for theories of threat detection such as the Snake Detection Theory, which posits that humans evolved to fear snakes and thus our thalamic nuclei are able to detect snakes rapidly and automatically. It is unlikely, however, that the thalamus has a true representation of “snake”. Our findings suggest a more plausible scenario, whereby visual features associated with threatening stimuli, such as snakes, are quickly detected and passed on to visual cortex for further processing. We have also identified the low-level contour features that are associated with positive valence.
Abstract This chapter discusses the positioning of Belarus in the international context of socioeconomic development based on an assessment of the country's dynamics in world rankings. The country's presence in the recognized world rankings and its holding high positions in them is an obvious advantage for achieving a favorable investment image. Ratings characterize the country's comparative position at the international level in a number of areas: from credit capacity to human capital development. There has been analyzed the position of the Republic of Belarus in several recognized international comparisons, such as Human Development Index, Doing Business, ICT Development Index, Global Innovation Index, Sustainable Development Goals Index, Corruption Perceptions Index, Rule of Law Index, Worldwide Governance Indicators, and others. However, Belarus is not yet participating in the international competitiveness assessment through such popular international ratings as Global Competitiveness Index and Global Entrepreneurship Monitor. The research findings show that the strongest aspects of the socioeconomic development of Belarus are in place due to the high educational level of the human capital development, gender equality, and the implementation of the UN sustainable development goals. The analysis also shows that the weaknesses of institutional environment and public administration do not enable the full implementation of the planned goals of socioeconomic development.
Sentiment Analysis is an application of Natural Langue Processing to analyze social media corpora to extract insights of corpora. Sentiment analytical results are the real feedback of the customers, which enables the organizations and companies to take appropriate decision on their products and business policies. Stemming plays in-evitable and vital role in sentiment analysis. Stemming is one of the phase of preprocessing the social media corpora. Today most of the researches uses strong stemmers to identify stem words of social media corpora. The most popular stemming algorithms such as Lancaster and Porter stemming algorithms causes prejudiced the meaning of the words. The over-stemmed words mislead the sentiment classification process. To prevent the over-stemming the Unprejudiced lighter stemming algorithm is proposed to sustain the meaning of the stemmed words. The propose Un-prejudiced algorithm uses lexical database and Parts of speech of Python Natural Language Tool Kit. There are a few stemming algorithm accuracy evaluation methods, in this paper we focused on Paice Error-rate relative to truncation (ERRT) measure to evaluate the accuracy of Lancaster, Porter and Unprejudiced stemming algorithms. The experiments were conducted on 25,758 source words and results were evaluated using Paice stem evaluation method and Sirsat method. The Paice Evaluation ERRT values 0.47209, 0.28703, 0.15502 of Lancaster, Porter, Unprejudiced respectively are proved that the Unprejudiced stemmer is more accurate than Lancaster and Porter. Sirsat’s stem evaluation method Average Words Conflation Factor (AWCF) results 10310.31, 14031.17, 23349.87 of Lancaster, Porter, Unprejudiced respectively are also proved the Unprejudiced stemming algorithm is more accurate than Lancaster and Porter stemming algorithms.
The present study explores the images of standard Lithuanian of young people at a gymnasium located in the area of the standard language. The data were obtained from a questionnaire based on the methodological principles of perceptual dialectology. The image of the standard language in the consciousness of the respondents emerges from the analysis of the questionnaire data: the frequency of linguistic codes, the mental maps of the standard language areas, and associations of the standard language. The analysis of the data shows that the gymnasium students tend to distance themselves from the regional linguistic code. The respondents’ disassociation from the local variety and their stronger preference for the code of the standard language is probably related to their sense of language security in the area of the linguistic homeland (including that of the standard language). The mental maps show that the young people associate the standard Lithuanian with the larger or smaller area of central Lithuania, which includes cities (Kaunas, Vilnius), adjacent non-dialect areas (Jonava, Kaišiadorys), and one or two dialect zones; it nearly overlaps the area of the standard language delineated in the second decade of the twenty-first century. Vilnius is the part of this image – probably of its status of the capital city and a significant social, cultural, and urban centre of attraction. The gymnasium students think that speakers of the standard language are city dwellers first and foremost, while the mental connection between the code of standard language and education occurs less often. Such views might have emerged due to the location of the city – hence that of the respondents’ linguistic homeland. Identifying the standard language user as an ordinary person or a Lithuanian could most likely be explained by the fact that the standard language is not only a national language to the young people: it is also an equivalent of their linguistic code. The gymnasium students do not associate the standard language with linguistic norms (the correct use of language). The consistency of the young generation’s attitudes (both those visualised on the maps and verbalised in the questionnaire answers) suggests the high value of the variety spoken in the area they associate with the standard language. The results of the study provide insights into the functioning, vitality, and continuity of the standard language in this area.
The research features syntactic structure in the twentieth century "The Catcher in the Rye" by J. D. Salinger and the twenty first century "Fates and Furies" by L. Groff. The research objective was to study the nature of syntactic relations expressed by word order in speech of narrators and characters. The paper outlines the rules of word order in the English sentence and reviews related studies in the field of syntax. The author analyzed the syntactic structure of sentences in the speech of narrators and characters in the two novels. The analysis was based on the descriptive method and techniques of observation, interpretation, comparison, and generalization. There were numerous examples of omission of auxiliary verbs in interrogative sentences in characters' speech, as well as interrogative sentences with affirmative structure. In "The Catcher in the Rye", affirmative sentences obecame interrogative with the help of interjections eh and ah. Both novels contained sentences where adverbial modifiers, objects, or attributes preceded the main parts – in the narrators' speech. A lot of one-member and contextually incomplete sentences were used to describe events and personages in both novels. In "The Catcher in the Rye", the narrator's speech revealed few cases of violations of word order rules, mostly in sentences with direct word order. The characters' speech appeared to contain much more cases of word order violations, since the novel features colloquial speech of twentieth century American teenagers. The speech of adult personages was characterized by correct word order. In "Fates and Furies", the narrator's speech demonstrated a significant number of elliptical sentences where auxiliary verb to be was omitted in simple verbal predicate with the verb in Present Continuous, as well as in compound nominal predicate and in passive voice. A comparative study of syntactic structure contributed to a deeper understanding of the nature of syntactic relations reflected by word order in the English sentence, grammatical structure of the English language, and popular types of sentences. In addition, the study showed the way native speakers express their ideas and thoughts by linguistic means and violate linguistic norms. The results can be used in various grammar courses and compiling textbooks.
Neuroimaging investigations in non-exercise contexts have shown that the dorsolateral prefrontal cortex (dlPFC), medial PFC and anterior cingulate, are engaged when individuals attempt to cognitively control negative affect. Moreover, there are indications that aversive interoceptive stimuli preferentially activate the right hemisphere. We theorized that affective responses to incremental exercise would be regulated by the same prefrontal network implicated in non-exercise affect regulation. We hypothesized that there would be preferential right-dlPFC activation, among individuals with low tolerance to exercise intensity and, therefore, less positive affective responses to challenging intensities of exercise (i.e., above ventilatory threshold, VT). PURPOSE: To investigate dlPFC activation and affective responses during incremental exercise. METHODS: Thirty-eight participants (15M, 21F, Age: 23.7 ± 6.9 y; BMI: 24.0 ± 4.8 kg·m-2; VO2max: 32.8 ± 7.8 ml·kg-1·min-1) completed an incremental cycling test to volitional termination. They were divided into low- and high-Tolerance groups based on a median split of their Tolerance scores (Preference for and Tolerance of the Intensity of Exercise Questionnaire). Near-infrared spectroscopy was used to assess changes from rest in the Tissue Oxygenation Index (ΔTOI) in the left (AF3) and right (AF4) dlPFC. Affective valence ratings (Feeling Scale; FS) were collected each min. RESULTS: Tolerance scores were positively correlated with FS ratings above VT (r = 0.33, p =.04), such that lower-Tolerance individuals reported lower FS ratings. For ΔTOI, a significant interaction was found between Tolerance group (low-high) and Hemisphere (left-right), p =.02, ηp2 =.129. ΔTOI in the right dlPFC was larger for low- vs high-Tolerance individuals (p =.03). CONCLUSION: Low self-reported tolerance for exercise intensity is associated with lower ratings of affective valence above VT and larger increases in right-dlPFC oxygenation from rest. These results suggest that the prefrontal regulation of negative affective responses to increasing exercise intensity may exhibit similarities to the regulation of negative affective responses in non-exercise contexts.
The recognition of the Canadian language standard means that in Canada, continuous historical interaction of the British and American English standards brought about a special norm which is no longer British or American. The distinguishing features of Canadian prosodic rules serve as the evidence for divergent characteristics of the Canadian linguistic standard in the paradigm of the national variations of the multinational English language. Therefore, Canadian English is an amalgam of American, British, and Canadian English. It includes the literary norm of its own, based on the local linguistic norm, distributed through the system of education and media in the Canadian areas, where the local speech distinctiveness is particularly pronounced and is close to the dialect status.The body of the experimental material, arising from the research aim and the tasks, included 20 episodes of speech from representatives of 4 Canadian areas: provinces Ontario, British Columbia, the Islands of Newfoundland and Labrador, Island of Prince Edward. All experimental texts are quasi-spontaneous monologues of improvised, informal, and laidback nature. These characteristics of a communicative act are the key extralingual factors to feature colloquial speech. This gives the groundings to study these monologue texts as samples of natural colloquial informal speech. The perceptive and instrumental analysis revealed the key speech units of the speech melodies (scales and terminal tones), common and distinguishing features of speech prosody as well as frequent speech parameters, taking into consideration gender and territorial belonging of the Canadians. The bilingual situation in the country, influence of British and American English on Canadians’ pronunciation resulted in emerging of a distinguishing Canadian English with its peculiarities at the prosodic level of the phonological system. The divergent features of Canadian English are associated with re-distribution of functions of American and British prosodic parameters with a due consideration of the new language system requirements, revealed at the prosodic level as a high frequency of descending stepping and level types of scale, ascending and fall-rising terminal tone within the variable nature of basic frequency. The specific national features of the prosodic system of Canadian English is the high frequency of descending and level stepping scales.
On the one hand, use and spread of slang as an important part of the mode of communication has been mostly associated with teenagers' language who by passing the accepted linguistic norms seek proving themselves. Moreover, origination of slang is said to be one of the properties of teenagers at the turn of the century. On the other hand, although language of genders has been studied by several scholars, it gets more interesting when it is paired with this entertaining word play, which due to its very nature-being coarse and direct-has mostly been attributed to males. Females, however, are getting inventive and trying to come up with some slang of their own and it seems that for teenage girls’ slangs are evolving at an even faster clip. Are there really any differences between male and female teenagers in their use of the slang words? If there are, what kind of difference may be found in societies which are perceived to be male dominant? Being concern of the present study, the researchers used both empirical and ethnographic elements of research to find relationships, if any, between use of slang words and the gender of Iranian teenagers, in a society which is perceived to be mostly male-dominant. Two high schools were therefore randomly chosen and six male and six female high schoolers who aged between 16 and 17 and came from upper middle-class families were randomly selected and interviewed using sociolinguistic interview protocols. Preliminary results, based on Chi-Square tests, revealed significant differences between the linguistic behavior of Iranian male and female teenagers. Females were interestingly found to be both more direct and more creative in their use of slang words. Our findings shed doubt on the generally accepted view of scholars on the linguistic (and social) dominance of males in modern Iranian society.
Many government schemes were unsuccessful because lack of proper feedback on the ongoing schemes, where billion dollars investment is going to be in vain. Sentiment analysis is one of best approach to analyse opinions of the peoples on various government schemes. Sentiment analysis and machine learning techniques emerged to analyse huge social media corpora to track people's views on government policies, products and services. Sentiment analysis process consists of various phases which include data discovery, data collection, data pre-processing, and data analysis. Stemming is a process to generate the morphemes in natural language sentences for various applications such as sentiment analysis, information retrieval, and domain analysis. The stemming process involved two major errors, which are over-stemming and under-stemming errors. Most of sentiment analysis natural languages processing applications used Lancaster and Porter stemming algorithms where more than one word inflected into same morpheme, which causes the etymology behaviour of the stemming word and prone to classify the tweets false positives and false negative. The proposed un-prejudice light stemming algorithm prevent etymology behaviour of morpheme and sustain its meaning during stemming process by selecting a word which has maximum number of synonyms in lexical database.
The similarity between words constitutes significant support to tasks in natural language processing. Several works use Lexical resources such as WordNet for semantic similarity and synonym identification. Nevertheless, words out-of-vocabulary or missing links between senses are perceived problems of this approach. Distributional-based proposals like word embeddings have successfully been used to meet such problems, but the lack of contextual information can prevent the achievement of even better results. The distributional models that include contextual information can bring advantages to this area, but these models are still scarcely explored. Therefore, this work studies the advantages of incorporating syntactic information in the distributional models, fostering for better results in semantic similarity approaches. For that purpose, the current work explore existing lexical and distributional techniques regarding the measurement of word similarity in Brazilian Portuguese. Experiments were carried out with the lexical database WordNet, using different techniques over a standard dataset. The results indicate that word embeddings can cover words out of vocabulary and have better results in comparison with lexical approaches. The main contribution of this article is a new approach to apply syntactic context in the training process of word embeddings to a Brazilian Portuguese corpus. The comparison of this model with the outcome of the previous experiments shows sound results and presents relevant complementary aspects.
When biblicalhumanities.org launched in 2013, we naively assumed that software developers would know how to take advantage of high quality, innovative data if it were published in open formats under free licenses. We understood the importance of building communities that know how to take advantage of this data, software and frameworks that make it easily accessible, and tailoring data to specific use cases, but we assumed that this would happen if we just made the data freely available. Freely licensed base texts, morphologies, contextual glosses, syntax treebanks, discourse analyses, lexicons, textual variants, images, conjectural emendations, and grammars are available now, and some are at least as good as commercial resources, but they are not yet widely used. In the last year or so, Jupyter Notebooks and academic conferences have started to bring this data into academic study, but we have barely begun to tap its potential for translation, language learning, and scripture engagement for those who know the languages. To realize this potential, we need to understand the use cases and mindset of potential users and create software and other resources tailored to their needs. Over time, we need to gather feedback from these same users, walk with them as they learn and grow, and welcome them into communities that will create the next generation of resources.
This contains data files needed for FinMeter. This data is complementary for FinMeter Python library described in: Mika Hämäläinen and Khalid Alnajjar (2019). Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing. In <em>the Proceedings of The 12th International Conference on Natural Language Generation</em>. Sources: The pretrained vectors for Finnish (es - I know) and English (en) are from E. Grave, P. Bojanowski, P. Gupta, A. Joulin, T. Mikolov, <em>Learning Word Vectors for 157 Languages. Creative Commons Attribution-Share-Alike License 3.0</em>. See https://fasttext.cc/docs/en/crawl-vectors.html The word2vec model trained on the Finnish Internet ParseBank is from Kanerva, Jenna; Luotolahti, Juhani; Laippala, Veronika; Ginter, Filip: Syntactic N-gram Collection from a Large-Scale Corpus of Internet Finnish. Proceedings of the Sixth International Conference Baltic HLT. 2014. paper. Creative Commons Attribution-ShareAlike 4.0 International License. See http://bionlp.utu.fi/finnish-internet-parsebank.html The Finnish concreteness data has been automatically translated from Brysbaert, Marc, Amy Beth Warriner, and Victor Kuperman. "Concreteness ratings for 40 thousand generally known English word lemmas." <em>Behavior research methods</em> 46.3 (2014): 904-911. Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. see http://crr.ugent.be/archives/1330
The present study examined the effects of labeling on pain tolerance, sensation, and affect for individuals who are high or low pain catastrophizers, as measured through the Pain Catastrophizing Scale (PCS). Pasticipants completed the PCS and were randomly assigned to 1 of 3 labeling conditions: A maximizing, a minimizing, and a neutral label condition. All participants then took part in a cold-pressor test. The cold-pressor measure of pain tolerance, as well as visual analog scales of sensory and affective ratings of pain, provided the dependent measures. Participants also completed the Anxiety Sensitivity Index (ASI) and the Somatic Amplification Questionnaire (SAQ). Results indicated that high pain catastrophizers have significantly reduced pain tolerance, increased pain sensations, and increased pain unpleasantness compared with low pain catastrophizers. In addition, significant correlations were found between the dependent measures. Main effects for labeling, and interaction effects between pain catastrophizing and labeling, were not supported.
This contains data files needed for FinMeter. This data is complementary for FinMeter Python library described in: Mika Hämäläinen and Khalid Alnajjar (2019). Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing. In <em>the Proceedings of The 12th International Conference on Natural Language Generation</em>. Sources: The pretrained vectors for Finnish (es - I know) and English (en) are from E. Grave, P. Bojanowski, P. Gupta, A. Joulin, T. Mikolov, <em>Learning Word Vectors for 157 Languages. Creative Commons Attribution-Share-Alike License 3.0</em>. See https://fasttext.cc/docs/en/crawl-vectors.html The word2vec model trained on the Finnish Internet ParseBank is from Kanerva, Jenna; Luotolahti, Juhani; Laippala, Veronika; Ginter, Filip: Syntactic N-gram Collection from a Large-Scale Corpus of Internet Finnish. Proceedings of the Sixth International Conference Baltic HLT. 2014. paper. Creative Commons Attribution-ShareAlike 4.0 International License. See http://bionlp.utu.fi/finnish-internet-parsebank.html The Finnish concreteness data has been automatically translated from Brysbaert, Marc, Amy Beth Warriner, and Victor Kuperman. "Concreteness ratings for 40 thousand generally known English word lemmas." <em>Behavior research methods</em> 46.3 (2014): 904-911. Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. see http://crr.ugent.be/archives/1330
English Abstract: The article analyses the relationship between etymology of political neologisms, ways of their formation and methods of rendering them from English into Russian. Translation strategy is directed at the recipient of the translated text and should therefore be pragmatic and based on the functional and stylistic norms of the translation language. The author concludes that the most efficient way of word-formation in the English language is compounding (morphological neologisms) and derivation of new meaning for the already existing words (semantic neologisms). The translation methods demonstrating high potential are transliteration, calquing (loan translation) and descriptive translation. Adequate translation is based on the following criteria: conciseness, single interpretation, conformity with the linguistic norms of the translation language. Russian Abstract: В работе затрагиваются вопросы взаимосвязи между этимологией, способами образования политических неологизмов и приемами их передачи с английского языка на русский. Выбор переводческой стратегии ориентирован на получателя текста перевода, поэтому должен обеспечивать прагматические задачи в соответствии с функционально-стилистическими нормами языка перевода. В работе делается вывод о том, что наиболее продуктивным способом образования новых слов является процесс словосложения (морфологические неологизмы) и наделения уже существующих в языке слов новым значением (семантические неологизмы); самыми эффективными приемами перевода неологизмов являются транслитерация, калькирование и описательный перевод. Критериями адекватного перевода политических неологизмов служат краткость, однозначность толкования, соответствие нормам языка перевода.
It is common knowledge and prescribed in all normative Portuguese grammars that the verb must agree in number and person with its subject, whether the latter is superposed or postponed to the verb. The lack of agreement between subject and verb (concordância verbal varíavel) is seen by users with a scholastic education as being wrong and linked to the poorest social strata, that is, with a low or zero level of education. However, cases of lack of agreement are not uncommon in informal speech of users with medium or high levels of education. This dichotomy between linguistic norms and orality, and its perception by the Brazilian population (i.e. lack of agreement as a sign of a lower educational and social level) can be verified also in the artistic reproduction, namely in the filmic dialogues of Brazilian national cinema, where verbal agreement variation is used to typify characters with little or no education. This paper will attempt to analyse how this linguistic phenomenon is interpreted by film discourse, i.e. a reproduction of orality. First, the linguistic issue will be presented, that is, the agreement and the lack of it in some registers of Brazilian Portuguese, and a brief presentation of the type of data on which this research was conducted (filmic dialogues from Brazilian films). Then, cases of variable verbal agreement (hereafter CVV) of the first plural person (hence 1PP) will be presented. In this perspective, the analysis of the alternation of use between two pronominal forms in subject function for the 1PP will be considered as a possible cause for the variation of verbal agreement between verb and subject with the 1PP. The approach adopted in this research brings together the variational studies and tools of Corpora Linguistics in an attempt to offer a critical view of the linguistic choices involved in film production.
The article discusses the current changing of linguistic norms in English as a lingua franca of global communication nowadays. It aims at both determining the causes of language deviations and analyzing language errors as well as their impact on the effectiveness of the English language communication. Based on the analysis of abundant empirical material, we prove that language innovations are caused by the immanent structural, functional, and pragmatic variability / instability of the English language; they are also associated with cognitive and sociocultural evolution. The research methodology includes: a corpus-based analysis of speech errors; interpretative, context and discourse analyses of the sources of language errors, as well as their distribution, adaptation, habitualization, legitimization, and regulation. We discuss the degree of influence of these processes on native and non-native speakers. Special attention is paid to multilingual interference and the Internet language creation. The findings show that it is impossible to separate language errors from language innovations today. Such conventional governing principles of error normalization as credibility, codification, and approval are still playing an important role while the demographic and geographic principles are losing their significance. The Internet communication often proliferates error normalization processes, which result in evaluating (accepting or rejecting) any innovation according to the principle of “virtual validity”. In conclusion, the English language status as a language of the international communication significantly transforms its norms, rules, and traditions. We think, this will not worsen it, but allow people of different nationalities to communicate in English more effective using their “English variant”, which is the most adapted one to their cognitive, functional and pragmatic needs.
Cet article presente la creation d’un treebank journalistique serbe, ParCoJour. Il est compose de 30K tokens et dote de trois couches d’annotation: etiquetage morphosyntaxique, lemmatisation et annotation syntaxique. Une fois construit, ParCoJour a ete utilise dans trois experiences afin d’evaluer l’impact du domaine textuel sur le parsing du serbe en comparant les performances de Talismane, un systeme par apprentissage automatique, sur deux types de corpus, journalistique et litteraire: 1) parsing du corpus journalistique avec un modele entraine sur le corpus journalistique; 2) parsing du corpus journalistique avec un modele entraine sur le corpus litteraire; 3) parsing du corpus litteraire avec un modele entraine sur le corpus journalistique. Les resultats sont compares a ceux ou les deux corpus relevaient du domaine litteraire. Le changement de domaine textuel dans la deuxieme et la troisieme experience entraine une baisse des performances, mais les resultats de parsing restent satisfaisants.
Like odor identification, remote odor memory, reflected in familiarity ratings, is impaired in AD (Murphy, Nature Reviews Neurology, 2019). We investigated the relative abilities of standard screening (MMSE), odor identification and remote odor memory to predict transitions from amnestic MCI (aMCI) to AD in a sample from the UCSD ADRC. The sample contained 170 controls, 210 AD, 26 aMCI converters to AD and 42 aMCI non-converters. A receiver operating characteristic (ROC) curve plots the trade-off between sensitivity and specificity. The area under the curve (AUC) indicates how well a marker discriminates patients from controls. Analyses showed higher predictive value for converting from aMCI to AD in ApoE ε4+ carriers for odor familiarity, odor identification and for the combination than for the MMSE. ROC/AUCs for the conversion from aMCI to AD have ranged from.63 -.67 for CSF biomarkers. Odor familiarity and odor identification had similar AUC values; however, combining odor familiarity and odor identification produced an ROC/AUC value of 1.0 in ε4 carriers, appreciably higher than for MMSE alone (.58). Olfactory biomarkers show real promise as early, non-invasive indicators of disease, particularly in samples enriched with ε4 carriers. Although odor identification has been the focus of olfactory biomarker work, the results suggest that other measures of olfactory function have the potential to enhance prediction. Combining odor familiarity and odor identification produced a predictive value of 1.0 in ε4 carriers. The results warrant further investigation into the potential for enhancing drug trials and clinical screening. Supported by NIH grants R01AG004085-26 (CM) and P50AG005131 (UCSD ADRC). We thank the UCSD ADRC and particularly Drs. Douglas Galasko and David Salmon.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
Death is a vague, frightening and abstract concept, mostly considered as a taboo. The current study investigates the ways of conceptualizing death among Persian language speakers in a representative corpus of Persian texts. It seeks answer to the question: What kind of linguistic tools are used by Persian language speakers in order to conceptualize the phenomenon of death and what cultural elements are involved in this regard.To answer the research question, we used the Persian Linguistic Database (PLDB) and a selection of obituaries and epitaphs as data and attempted to identify and extract all the expressions which were directly or indirectly related to the concept of death. Using the conceptual metaphor theory (CMT), by Lakoff and Johnson (1987), and cultural conceptualization by Sharifian (2011), the ways death is conceptualized were identified on the basis of a cognitive-cultural approach.
Abstract: Temperatures above 20° Celsius have shown to adversely impact human behavior, leading to increased aggression and violence. Climate change will contribute to both the magnitude and severity of this pattern as temperatures continue their rise. Contributions to this field of research have only recently begun to analyze online behavior and language as a proxy for hedonic state, or well-being. From a development perspective this study is relevant since the poor tend to live in some of the warmest regions on earth, and would thus be disproportionately impacted by increased temperatures. We use several sources of data; U.S. based daily statewide temperature data from 2016 through 2017, as well as localized viewer chat data from a live video streaming website. We will sort chatting comments looking for key words (i.e. hate speech, swearing, etc.), and with the use of a word rating system we then assess the overall mood of the chatters contingent on high temperature readings on the precise day of the communications. After controlling for spatiotemporal fixed effects, we find strong evidence that hedonic state decreases above 20°c.
Antiaddictive social advertising is a special speech genre of modern communication with specific features determined by the target setting and the chosen strategy. Advertising can be considered as a special functional style, within which separate genres are distinguished, first of all commercial advertising and social advertising. Social advertising, which refers to ethical categories, has become especially popular because society is faced with such problems, the solution of which depends on mass behavior. Advertising has become an integral part of the daily life of a person. Since advertising is mass-replicated in the media, it can enter the consciousness of the addressee, even against his/her will and desire, no wonder advertising is sometimes defined as the “fifth power” after media, whose power is considered the “fourth power”. Therefore, it is so important for advertising to follow the moral principles of society, it is so important to comply with modern ethical and linguistic norms. If commercial advertising is widespread, the antiaddictive one is little known. Many specialists in advertising argue for the need of the strategy of “shock” advertising, a remarkable feature of which is hyperbolization, used to cause fear. In antiaddictive advertising there is often a morphological imperative in the meaning of categorical motivation. It is justified as such advertizing has to be as much as possible appellate, has to be understood unambiguously. Anti-drug social advertising should show positive, motivation to a healthy lifestyle.
In two experiments, the influence of inducing negative mood on cognitive performance was explored by analyzing physical arm reaching movements as indicators of mind wandering. Mood was induced by viewing a series of six photos per mood condition that were previously established for their emotionally valenced and arousal ratings. A reach tracking device recorded three metrics of arm movement that were expected to reflect instances of mind wandering: initiation latency, movement time, and arm curvature. In the first experiment, 29 participants were randomly assigned into one of two induced-mood groups, negative mood (n = 15) or neutral mood (n = 14). Participants performed a simple Go/No-go task in which arm movements were detected by the reach tracker. The first experiment indicated that the mood inducement was successful but the effect of negative mood on either self-reported mind wandering or variances in arm movement were not significant. Thus, the second experiment prompted the change to a visual-search target-selection task in which variances in initiation latency, movement time, and curvature were expected to be more pronounced. The second experiment consisted of 23 participants who were also randomly assigned to either negative (n = 12) or neutral (n = 11) mood condition. The second experiment revealed that the mood induction was still successful but that there were still no significant effects observed between mood and indicators of mind wandering. Though the results of this study did not reflect initial predictions, it may suggest that low-arousing negative moods in healthy individuals are not associated with increased mind wandering.
We recently introduced the EmojiGrid as an intuitive graphical self-report tool to measure food&shy;evoked valence and arousal. The EmojiGrid is a Cartesian grid, labeled with facial icons (emoji) expressing different degrees of valence and arousal. The lack of verbal labels makes it a valuable, language-independent tool for cross-cultural research. Users can efficiently report their subjective ratings of both valence and arousal with a single click on the location of the grid that best represents their affective state after perceiving a stimulus. The EmojiGrid has previously been validated for the assessment of emotions evoked by food images. In this study we validated the EmojiGrid for the affective appraisal of odors. Observers (N=56, 24 males, mean age=24.3±4.6) smelled 40 randomly presented odors (27 food and 13 non&shy;food smells), ranging from very unpleasant and arousing (e.g., feces, fish), via pleasant and calming (e.g., clove, cinnamon), to very pleasant and stimulating (e.g., peach, caramel). The odor samples consisted of felt pens, with tips that were impregnated with 4 ml of fluid odorant substance. Each pen was presented once, for about 5 seconds at 2 cm below both nostrils. The participants sniffed following a verbal command. Immediately after sniffing the pen was removed, and participants were given at least 30 s to smell fresh air. The participants reported their affective appraisal of each odor using the EmojiGrid. The resulting mean valence and arousal ratings closely agree with those from previous studies in the literature that were obtained with alternative rating tools. In addition, we find that the EmojiGrid yields the typical universal U-shaped relation between mean valence and arousal that is commonly observed for a wide range of affective sensory (visual, auditory, tactile, gustatory) stimuli. We conclude that the EmojiGrid is also a valid affective self-report tool for the assessment of odor evoked emotions.
Mirror-sensory synesthetes mirror the pain or touch that they observe in other people on their own bodies. This type of synesthesia has been associated with enhanced empathy. We investigated whether the enhanced empathy of people with mirror-sensory synesthesia influences the experience of situations involving touch or pain and whether it affects their prosocial decision making. Mirror-sensory synesthetes (<i>N</i> = 18, all female), verified with a touch-interference paradigm, were compared with a similar number of age-matched control individuals (all female). Participants viewed arousing images depicting pain or touch; we recorded subjective valence and arousal ratings, and physiological responses, hypothesizing more extreme reactions in synesthetes. The subjective impact of positive and negative images was stronger in synesthetes than in control participants; the stronger the reported synesthesia, the more extreme the picture ratings. However, there was no evidence for differential physiological or hormonal responses to arousing pictures. Prosocial decision making was assessed with an economic game assessing altruism, in which participants had to divide money between themselves and a second player. Mirror-sensory synesthetes donated more money than non-synesthetes, showing enhanced prosocial behaviour, and also scored higher on the Interpersonal Reactivity Index as a measure of empathy. Our study demonstrates the subjective impact of mirror-sensory synesthesia and its stimulating influence on prosocial behaviour.This article is part of the discussion meeting issue ‘Bridging senses: new developments in synaesthesia’.
Classic natural language processing resources such as the Penn Treebank (Marcus et al. 1993) have long been used both as evaluation data for many linguistic tasks and as training data for a variety of off-the-shelf language processing tools. Recent work has highlighted a gender imbalance in the authors of this text data (Garimella et al. 2019) and hypothesized that tools created with such resources will privilege users from particular demographic groups (Hovy and Søgaard 2015). Domain adaptation is typically employed as a strategy in machine learning to adjust models trained and evaluated with data from different genres. However, the present work seeks to evaluate whether domain adaptation to demographic groups such as age or gender may be an effective strategy to ameliorate the effects of biased or outdated training corpora in linguistic preprocessing tasks. We find adaptation to demographic groups to be an effective strategy for improving preprocessing performance across all demographic groups.
In this paper, we discuss constituent ordering generalizations in Japanese. Japanese has SOV as its basic order, but a significant range of argument order variations brought about by ‘scrambling’ is permitted. Although scrambling does not induce much in the way of semantic effects, it is conceivable that marked orders are derived from the unmarked order under some pragmatic or other motivations. The difference in the effect of basic and derived order is not reflected in native speaker’s grammaticality judgments, but we suggest that the intuition about the ordering of arguments may be attested in corpus data. By using the Keyaki treebank (a proper subset of which is NINJAL Parsed Corpus of Modern Japanese (NPCMJ)), it is shown that the naturally-occurring corpus data confirm that marked orderings of arguments are less frequent than their unmarked ordering counterparts. We suggest some possible motivations lying behind the argument order variations.
We present a recurrent neural network memory that uses sparse coding to\ncreate a combinatoric encoding of sequential inputs. Using several examples, we\nshow that the network can associate distant causes and effects in a discrete\nstochastic process, predict partially-observable higher-order sequences, and\nenable a DQN agent to navigate a maze by giving it memory. The network uses\nonly biologically-plausible, local and immediate credit assignment. Memory\nrequirements are typically one order of magnitude less than existing LSTM, GRU\nand autoregressive feed-forward sequence learning models. The most significant\nlimitation of the memory is generalization to unseen input sequences. We\nexplore this limitation by measuring next-word prediction perplexity on the\nPenn Treebank dataset.\n
In this paper, a classifier of emails by level of urgency in service companies is presented, using a natural language processing algorithm to grammatically label the unstructured text of messages to give it meaning and structure before analyzing it. The proposed classifier uses a lexical database that is composed of unigrams classified into eight basic emotions (anger, fear, anticipation, confidence, surprise, sadness, joy and disgust) and two feelings (positive and negative). This allows to compare the text of the grammatically labeled messages with the unigrams of the database, in order to conduct the analysis of emotions and feelings. The analysis determines the percentage of each emotion and the polarity of feeling in order to classify the most negative emails as urgent and channel them to the corresponding departments to be attended to. The research has been implemented using a set of data taken from a company dedicated to electronic invoicing in Mexico.