Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The emergence of China as a global economic power in the 21st Century has brought about surging needs for cross-lingual and cross-cultural mediation, typically performed by translators. Advances in Artificial Intelligence and Language Engineering have been bolstered by Machine learning and suitable Big Data cultivation. They have helped to meet some of the translator's needs, though the technical specialists have not kept pace with the practical and expanding requirements in language mediation. One major technical and linguistic hurdle involves words outside the vocabulary of the translator or the lexical database he/she consults, especially Multi-Word Expressions (Compound Words) in technical subjects. A further problem lies in the multiplicity of renditions of a term in the target language. This paper discusses a proactive approach following the successful extraction and application of sizable bilingual Multi-Word Expressions (Compound Words) for language mediation in technical subjects, which do not fall within the expertise of typical translators, who have inadequate appreciation of the range of new technical tools available to help him/her. Our approach draws on the personal reflections of translators and teachers of translation and is based on the prior R&D efforts relating to 300,000 comparable Chinese-English patents. The subsequent protocol we have developed aims to be proactive in meeting four identified practical challenges in technical translation (e.g. patents). It has broader economic implication in the Age of Big Data (Tsou et al, 2015) and Trade War, as the workload, if not, the challenges, increasingly cannot be met by currently available front-line translators. We shall demonstrate how new tools can be harnessed to spearhead the application of language technology not only in language mediation but also in the “teaching” and “learning” of translation. It shows how a better appreciation of their needs may enhance the contributions of the technical specialists, and thus enhance the resultant synergetic benefits. © 2019 Incoma Ltd. All rights reserved.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
Previous studies have shown that hoarding behavior usually starts at a subclinical level in early adolescence and gradually worsens; however, a limited number of studies have examined the prevalence of hoarding behavior and its association with developmental disorders in young adults. The aims of this study were to estimate the prevalence of hoarding behavior and to identify correlations between hoarding behavior and developmental disorder traits in university students. The study participants included 801 university students (616 men, 185 women) who completed questionnaires (ASRS: Adult ADHD Self-Report Scale version 1.1, AQ16: Autism-Spectrum Quotient with 16 items, and CIR: Clutter Image Rating). Among 801 participants, 27 (3.4%) exceeded the CIR cut-off score. Moreover, the participants with hoarding behavior had a significantly higher percentage of ADHD traits compared to participants without hoarding behavior (HB(+) vs HB(−), 40.7% vs 21.7%). In addition, 7.4% of HB(+) participants had autism spectrum disorder (ASD) traits, compared to 4.1% of HB(−) participants. A correlation analysis revealed that the CIR composite score had a stronger correlation with the ASRS inattentive score than with the hyperactivity/impulsivity score (CIR composite vs ASRS IA, r = 0.283; CIR composite vs ASRS H/I, r = 0.147). The results showed a high prevalence of ADHD traits in the university students with hoarding behavior. Moreover, we found that the hoarding behavior was more strongly correlated with inattentive symptoms rather than with hyperactivity/impulsivity symptoms. Our results support the concept of a common pathophysiology behind hoarding behavior and ADHD in young adults.
Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider the correlation between parameters and therefore can lead to unstable and slow convergence in training. Second-order optimization methods provide a possible solution to this issue. However these methods are either computationally heavy or do not have competitive performance. In this paper, a novel optimization method - stochastic natural gradient based on minimum variance assumption (SNGM) is proposed for training RNNLMs. It allows the natural gradient method to operate at a comparable training efficiency to the SGD method. By modifying the gradient according to the local curvature of the KL-divergence between current and updated probabilistic distributions, the proposed SNGM approach is shown to outperform both the SGD and limited memory BFGS methods across three tasks: Penn Treebank, Switchboard conversational speech recognition and AMI meeting room transcription in terms of both perplexity and word error rate.
espanolResumen en castellano: Esta tesis trata de la morfologia verbal del ingles antiguo para identificar y lematizar los verbos debiles de esta lengua en un corpus al que se accede a traves de una base de datos lexica. La lematizacion es una de las tareas mas importantes a la hora de construir un diccionario. Sin embargo, es una de las tareas pendientes en el campo de la linguistica historica debido a que no existen corpora exhaustivos y lematizados de esta lengua. El enfoque de esta tesis doctoral esta en la lematizacion de las tres clases de verbos debiles del ingles antiguo, aunque las areas de la Lexicografia y la Linguistica de Corpus son tambien relevantes para esta investigacion. Las fuentes principales de esta investigacion son las formas flexivas que estan atestiguadas en el Dictionary of Old English Corpus (DOEC) y que estan disponibles en el lematizador Norna, las fuentes lexicograficas que existen publicadas sobre esta lengua, principalmente el Dictionary of Old English (DOE), y otras fuentes textuales como el York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) y una indexacion de fuentes secundarias del ingles antiguo. El objetivo principal supone la identificacion de las flexiones de los verbos debiles y de su lematizacion con uno de los lemas propuestos en las listas de referencia. Conseguir este objetivo implica manejar las fuentes disponibles en ingles antiguo para poder lematizar y validar los resultados del analisis y el diseno de un metodo que combine busquedas automaticas en la base de datos lexica Nerthus y la revision manual de los resultados. La metodologia incluye cuatro pasos sucesivos con diversas tareas en cada paso. El primero de estos pasos tiene como objetivo la lematizacion de las formas canonicas de los verbos debiles lanzando cadenas de busquedas especificas para cada clase de verbos debiles en el lematizador Norna, donde esta disponible un indice de tipos del DOEC, la fuente de informacion mas fiable de la que se dispone en ingles antiguo. Despues, los resultados se validan con el DOE y se anaden las formas no-canonicas de los verbos debiles entre las letras A y H. El tercer paso tiene como objetivo identificar las formas no-canonicas de las terminaciones flexivas y de las vocales de los radicales que aparecen con mas frecuencia en los verbos debiles para generar patrones de lematizacion. La busqueda de estos patrones y de la lista de prefijos no-canonicos que esta disponible en Norna culmina en la lematizacion de las formas flexivas no transparentes de los verbos debiles. La validacion de los resultados de las letras I a la Y supone el ultimo paso de la metodologia, donde se comparan los datos obtenidos con el analisis sintactico del YCOE y con los datos que se obtienen de una base de datos de indexacion de las fuentes secundarias del ingles antiguo. Los problemas que surgen a lo largo del proceso de lematizacion tienen que ver principalmente con las peculiaridades del ingles antiguo y las limitaciones de la lematizacion de tipos que esta investigacion sigue. La discusion de los resultados del analisis concluye esta tesis. Las principales aportaciones de esta tesis son las listas de lemas y sus formas flexivas, especialmente las de los verbos entre las letras I y la Y ya que no estan disponibles todavia, y el metodo que se ha disenado para identificar estas formas, incluyendo los patrones de lematizacion generados para lematizar las formas con terminaciones no comunes y vocales no canonicas en el radical. EnglishThis thesis deals with the verbal morphology of the Old English language in order to identify and lemmatise weak verbs in a corpus accessed through a lexical database. Lemmatisation is a pending task in the field of historical linguistics given the lack of comprehensive and lemmatised corpora in this language. The focus of this doctoral dissertation is on the lemmatisation of the three classes of weak verbs, although the linguistic fields of Lexicography and Corpus Linguistics are also relevant to this research. The main aim involves the identification of the canonical and non-canonical realisations of the Old English weak verbs and their lemmatisation with a lemma from a reference list of weak verbs. Achieving this goal involves, firstly, the use of the available sources of the Old English language in order to lemmatise and validate the results and, secondly, the design of a semi-automatic research methodology that combines automatic searches in the lexical database Nerthus and the manual revision of the results in order to achieve this task. The sources for this investigation are the inflectional forms that are attested in the Dictionary of Old English Corpus (DOEC) which are available in the lemmatiser Norna, the lexicographical sources published on the Old English language, mainly the Dictionary of Old English (DOE), and other textual sources such as the York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) and an index of secondary sources of Old English. The methodology comprises four successive steps and several tasks within each step. The first step aims at the lemmatisation of the transparent forms of weak verbs with the search of specific query strings for each subclass of weak verbs in the lemmatiser Norna, where an index type of the DOEC, the most reliable source of information regarding the Old English language, is available. Then, the second step validates the results with the DOE and adds to the analysis the non-canonical attestations for the weak verbs from the letter A-H. Thirdly, the identification of the most recurrent non-canonical inflectional endings and stem vowels attested in weak verbs gives rise to lemmatisation patterns. The search of these sets of correspondences and the list of non-canonical prefixes that is available in Norna results in the lemmatisation of the non-canonical inflections of weak verbs. The validation of the results from the letter I-Y concludes the research methodology with the syntactic parsing provided by the YCOE and the data retrieved from the index of secondary sources of Old English Freya. The issues that arise throughout the lemmatisation process mainly concern the idiosyncrasy of the Old English language writing system and the limitations of the lemmatisation by type that this investigation follows. The quantitative and qualitative discussion of the results of the analysis concludes this thesis. The main contributions of this thesis are the lists of weak lemmas and their lemmatised inflectional forms, specially those of the verbs I-Y which are not available yet and the designed research methodology to identify these forms, including the sets of lemmatisation patterns of the non-canonical inflectional endings and stem vowels of weak verbs.
Recent research on discourse relations has found that they are cued not only by discourse markers (DMs) but also by other textual signals and that signaling information is indicative of genres. While several corpora exist with discourse relation signaling information such as the Penn Discourse Treebank (PDTB, Prasad et al. 2008) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada 2018), they both annotate the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. 1993), which is limited to the news domain. Thus, this paper adapts the signal identification and anchoring scheme (Liu and Zeldes, 2019) to three more genres, examines the distribution of signaling devices across relations and genres, and provides a taxonomy of indicative signals found in this dataset.
According to World Intellectual Property Organization (2017) report, over 3 million patents exist in the patent database, but only certain numbers have commercial potential. Generally, to assess the commercial potential of patent, it consumes time and requires various expertise. Currently, several models have been developed to address this matter, which to assess using questionnaire tool for portfolios by human. So that occurs bias any limitation exists that our research will address by artificial intelligence. Hence, this research applies a Natural Language Programming to assess for commercial potential of patent, consisting of five steps - (i) Morphological analysis based on the Lexical database, (ii) Syntactic analysis of sentence to check syntax sentence patterns, (iii) Sematic analysis to interpret the meaning of words derived from the previous step, (iv) Discourse integration from context of domain together with the main sentence providing more accurate sentence analysis, and (v) Pragmatic analysis to ensure the correct meaning of interpretation. Then, the obtained data is used to determine criterion factors and formulate the model for assessing commercial potential of patent using Natural Language programming. This finding should deliver an alternative effective patent assessment system, which addresses some current deficiency in patent's assessment for commercial potential.
The intelligent information processing of the standard Zhuang language spoken mainly in Southern China is presently in its infancy, and lacks a well-defined language corpus and automatic part-of-speech tagging methods. Therefore, this study proposes an adversarial part-of-speech tagging method based on reinforcement learning, which solves the problems associated with a lack of a language corpus, time-consuming laborious manual marking, and the low performance of machine marking. Firstly, we construct a markup dictionary based on the grammatical characteristics of standard Zhuang and the Penn Chinese Treebank. Secondly, a dependency syntax analysis is applied for constructing the semantic information feature vectors of sentences, and long short-term memory is adopted as the policy network architecture to enhance available information using recurrent memory, and a conditional random field is employed as the discriminant network to perform label inference with global normalization. Finally, we use reinforcement learning as the model framework, target parts of speech as the feedback of the environment, and then obtain the optimal policy through adversarial learning. The results show that the combination of reinforcement learning and adversarial network alleviates the dependence of the model on the training corpus to some extent, and can quickly and effectively expand the scale of the annotation dictionary for the Zhuang language, thereby obtaining better labeling results.
Lexical Markup Framework (LMF) or ISO 24613 [1] is a de jure standard that\nprovides a framework for modelling and encoding lexical information in\nretrodigitised print dictionaries and NLP lexical databases. An in-depth review\nis currently underway within the standardisation subcommittee,\nISO-TC37/SC4/WG4, to find a more modular, flexible and durable follow up to the\noriginal LMF standard published in 2008. In this paper we will present some of\nthe major improvements which have so far been implemented in the new version of\nLMF.\n
Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Speech-to-Text (STT) and Text-to-Speech (TTS) systems, the question remains how to best leverage state-of-the-art language models (which capture rich semantic features, but are trained on only written text) on inputs with ASR errors. In this paper, we present Telephonetic, a data augmentation framework that helps robustify language model features to ASR corrupted inputs. To capture phonetic alterations, we employ a character-level language model trained using probabilistic masking. Phonetic augmentations are generated in two stages: a TTS encoder (Tacotron 2, WaveGlow) and a STT decoder (DeepSpeech). Similarly, semantic perturbations are produced by sampling from nearby words in an embedding space, which is computed using the BERT language model. Words are selected for augmentation according to a hierarchical grammar sampling strategy. Telephonetic is evaluated on the Penn Treebank (PTB) corpus, and demonstrates its effectiveness as a bootstrapping technique for transferring neural language models to the speech domain. Notably, our language model achieves a test perplexity of 37.49 on PTB, which to our knowledge is state-of-the-art among models trained only on PTB.
Paper is dedicated to the testing of the concept of literacy, based on the questionnaire, carried out in the school year of 2018/2019 among the students of two secondary vocational schools in Vrsac, Belgrade and Grammar School in Vrsac (200 respondents). The primary hypothesis of the research was that detection and detailed study of high school students conceptosphere on literacy identify the fields to improve the teaching of Serbian as a mother tongue in secondary schools and the aim of work that, based on the collected and then processed data in analytical, cognitive and descriptive method, is to (a) isolate the dominant concepts of (non)literacy, (b) look at the tendencies of spreading and shaping the notion of literacy induced by the needs of modern life, and also that, in order to improve linguistic culture in all domains and all educational levels -(c) point to the possibility of improving the teaching of the Serbian language as a mother tongue. According to results of the survey secondary school students experience literacy in the 21 st century as a complex concept; from the one who is literate expecting linguistic knowledge, what are the basic, traditionally accepted parameters, and recognize illiteracy as the lack of ability to apply knowledge in the field of language. They also demonstrated that it is necessary to improve the efficiency of teaching approaches designed to improve functional literacy in a variety of communicative situations; increase the number of hours and exercises in the field of spelling, or nurture and acquire more comprehensive and knowledge in use and skills of different forms of literacy needed for managing in 21 st century; more attention should be paid to including relevant language handbooks in teaching; more explicit, on frequent and more familiar examples to students, point to the advantages of knowing and respecting the linguistic norm, paving the way for a better linguistic culture and enrichment of the mother tongue.
يعد القرآن الكريم من مصادر المعرفة، وقد تولدت منه فروع واسعة؛ إذ نُزل القرآن الكريم باللغة العربية، ولا يوجد خيار آخر لإتقان المعرفة الواردة فيه إلا من خلال تعلم اللغة العربية. تهدف هذه الدراسة إلى بيان مفهوم المدونة العربية القرآنية ومكوناتها، والكشف عن علاقة تعلم اللغة العربية بالقرآن الكريم، وبيان كيفية تعليم وتعلم القواعد العربية الأساسية عبر المدونة العربية القرآنية، وستتبع الدراسة المنهج الوصفي والتحليلي. إن وجود العلاقة بين اللغة العربية والقرآن الكريم، يدفع الطلبة المتخصصين في اللغة العربية أن يربطوا اللغة العربية بالقرآن؛ لذلك نرى أن المدونة العربية القرآنية تساعدهم على فهم القواعد القرآنية بطريقة مثيرة للاهتمام. في نظرة شاملة يمكن أن نستنتج أن المدونة العربية القرآنية هي واحدة من أهم الأدوات الحسابية التي تم إنتاجها في خدمة اللغة العربية؛ حيث توفر للمتعلمين ما يحتاجون إليه في مجال اللغة واللغويات والدراسات الحاسوبية، كما تمهد الطريق للباحثين لدراسة الهياكل المورفولوجية والنحوية من خلال دراسات الحوسبة العميقة للقرآن.
 الكلمات المفتاحية: المصرف القرآني، نموذج حاسوبي، المدونة العربية القرآنية، المعجم القرآني.
 Abstract 
 The Holy Quran is a source of knowledge and it has generated wide branches of knowledge. The Holy Quran was revealed in Arabic. Hence, there is no other option to master its knowledge except by learning the Arabic language. This study aims at explaining the concept of the Arabic Quranic Corpus and its components, revealing the relationship between learning the Arabic language and the Holy Quran, and showing how to teach and learn basic Arabic grammar through the Quranic Arabic Corpus. The study will follow the descriptive and analytical approach. The existence of the relationship between the Arabic language and the Holy Quran prompts Arabic learners to associate Arabic with the Qur'an. Therefore, we see that the Quranic Arabic Corpus helps them to understand Quranic rules in an interesting way. In a comprehensive view, we can conclude that the Arabic Quranic Corpus is one of the most important web-based medium produced to serve the Arabic language. It provides learners with what they need in the field of language, linguistics and computer studies, and paves the way for researchers to study morphological and grammatical structures through technology with detail description of grammars.
 Keywords: Quranic Treebank, Computational Model, Arabic Quranic Corpus, Qur’anic Dictionary.
The relationship between words in a sentence often tell us more about the underlying semantic content of a document than its actual words individually. Natural language understanding has seen an increasing effort in the formation of techniques that try to produce non-trivial features, in the last few years, especially after robust word embeddings models became prominent, when they proved themselves able to capture and represent semantic relationships from massive amounts of data. These new dense vector representations indeed leverage the baseline in natural language processing, but they still fall short in dealing with intrinsic issues in linguistics, such as polysemy and homonymy. Systems that make use of natural language at its core, can be affected by a weak semantic representation of human language, resulting in inaccurate outcomes based on poor decisions. In this subject, word sense disambiguation and lexical chains have been exploring alternatives to alleviate several problems in linguistics, such as semantic representation, definitions, differentiation, polysemy, and homonymy. However, little effort is seen in combining recent advances in token embeddings (e.g. words, documents) with word sense disambiguation and lexical chains. To collaborate in building a bridge between these areas, this work proposes a collection of algorithms to extract semantic features from large corpora as its main contributions, named MSSA, MSSA-D, MSSA-NR, FLLC II, and FXLC II. The MSSA techniques focus on disambiguating and annotating each word by its specific sense, considering the semantic effects of its context. The lexical chains group derive the semantic relations between consecutive words in a document in a dynamic and pre-defined manner. These original techniques' target is to uncover the implicit semantic links between words using their lexical structure, incorporating multi-sense embeddings, word sense disambiguation, lexical chains, and lexical databases. A few natural language problems are selected to validate the contributions of this work, in which our techniques outperform state-of-the-art systems. All the proposed algorithms can be used separately as independent components or combined in one single system to improve the semantic representation of words, sentences, and documents. Additionally, they can also work in a recurrent form, refining even more their results.
The practice of assessing brand management in construction in Ukraine is in a passive stage, but due to the entry into the Ukrainian market of foreign companies for which regular evaluation of their brand — the need for survival in a competitive environment, Ukrainian companies are beginning to pay more attention to the creation and formation of their competitive trading because of its high business image / rating. A modern toolkit based on appropriate approaches is used to form organizational and economic foundations. The article analyzes modern scientific approaches to assessing the economic potential of an enterprise, identifies the main trends and factors that affect the assessment of construction enterprises. The article is devoted to the study of theoretical and methodological assessments of branding of construction enterprises. The role and importance of innovation in ensuring the efficient operation of modern enterprises is emphasized. It is found that construction, especially innovative, is of great social importance and has a significant economic effect. The importance of determining the potential of innovative development of construction in general and construction enterprises in particular is substantiated. The purpose of this article is to systematically investigate the interpretation of the potential of innovative development of construction enterprises. The theoretical basis of the research is the scientific works of foreign and domestic scientists on the problems of identifying the essence of innovative development potential. A systematic study of the general characteristics of the construction company brand was conducted and the priority directions for choosing the development of the economic potential of the enterprises were determined. The necessity to understand the potential of the enterprise in the unity of all its elements, which are subject to the achievement of the overall goals of the enterprise, is substantiated. The weight of the component of the brand in the potential of the construction industry enterprises is substantiated. Some aspects of the development of theoretical and methodological approaches to the estimation of the intellectual capital of construction enterprises are formulated. Existing theoretical and methodological approaches to the brand assessment of construction enterprises are analyzed.
Nonsuicidal self-injury (NSSI; e.g., cutting or burning the skin without suicidal intent) is a dangerous and increasingly prevalent health-risk behavior. Despite advances in NSSI research over the past decade, many aspects of NSSI remain poorly understood. In particular, there are few strong predictors of NSSI, it is unclear how positive attitudes toward NSSI develop, and there are no empirically supported treatments for NSSI. In the present study, I addressed these topics with a multi-method, experimental, and longitudinal approach. For Aim 1 of the study, I examined baseline differences between NSSI (n = 58) and control (n = 86) adult participants on NSSI-themed versions of five measures that cover different aspects of attitudes: the implicit association test (IAT); the affect misattribution procedure (AMP); explicit affective ratings; startle eyeblink reactivity; and startle postauricular reactivity. Compared to the control group, the NSSI group displayed significantly more positive attitudes on all five measures. Moreover, AMP scores and explicit ratings prospectively predicted self-cutting frequency over the ensuing six months. For Aim 2, I employed pain offset relief conditioning in an attempt to induce more positive implicit attitudes toward NSSI in the control group. This conditioning significantly diminished startle eyeblink reactivity in the context of NSSI images, but did not significantly affect any other measures. For Aim 3, I tested the ability of aversive conditioning in the NSSI group to reverse positive implicit attitudes toward NSSI and to reduce NSSI behaviors over the subsequent six months. Aversive conditioning normalized startle eyeblink and postauricular reactivity, but did not significantly affect any other measures. Results also provided preliminary support for the hypothesis that aversive conditioning prospectively reduces self-cutting. In conjunction with my other recent studies (Franklin et al., 2010; 2011, 2012, 2013), these findings have prompted a new theoretical framework called the Benefits and Barriers model of NSSI.
This paper suggests annotation guidelines to build a Universal Dependencies (UD) treebank for Korean. We discuss the part-of-speech annotation of Korean specific-categories such as prenouns, numeral classifiers, and (pre)final endings, and propose how to implement UD scheme in Korean regarding selecting a head and assigning dependency relations to dependents. UD prioritizes content words over functional words since the former exhibits less cross-linguistic variations. In a noun phrase, for instance, a core noun is always a head of the entire noun phrase independently of a language. The rest are treated as a dependent: not only a modifier such as an adjective but also a functional category such as an article, numeral quantifier, demonstrative, and so on. However, when it comes to head-less constructions such as coordination or predicate ellipsis, UD firmly advocates the head-initial strategy. The present application of UD to Korean tries to follow UD’s principles as much as possible. Korean is a head-final language, so that headed constructions are analyzed head-finally. In contrast, head-less ones are tagged head-initially. This might disregard language-specific characteristics from a linguistic perspective, but the strategy allows us to build up a set of treebanks in a cross-linguistically consistent way (i.e., the fundamental purpose of UD).
This thesis studies the connections between parsing friendly representations and interlingua grammars developed for multilingual language generation. Parsing friendly representations refer to dependency tree representations that can be used for robust, accurate and scalable analysis of natural language text. Shared multilingual abstractions are central to both these representations. Universal Dependencies (UD) is a framework to develop cross-lingual representations, using dependency trees for multlingual representations. Similarly, Grammatical Framework (GF) is a framework for interlingual grammars, used to derive abstract syntax trees (ASTs) corresponding to sentences. The first half of this thesis explores the connections between the representations behind these two multilingual abstractions. The first study presents a conversion method from abstract syntax trees (ASTs) to dependency trees and present the mapping between the two abstractions – GF and UD – by applying the conversion from ASTs to UD. Experiments show that there is a lot of similarity behind these two abstractions and our method is used to bootstrap parallel UD treebanks for 31 languages. In the second study, we study the inverse problem i.e. converting UD trees to ASTs. This is motivated with the goal of helping GF-based interlingual translation by using dependency parsers as a robust front end instead of the parser used in GF. \n\nThe second half of this thesis focuses on the topic of data augmentation for parsing – specifically using grammar-based backends for aiding in dependency parsing. We propose a generic method to generate synthetic UD treebanks using interlingua grammars and the methods developed in the first half. Results show that these synthetic treebanks are an alternative to develop parsing models, especially for under-resourced languages without much resources. This study is followed up by another study on out-of-vocabulary words (OOVs) – a more focused problem in parsing. OOVs pose an interesting problem in parser development and the method we present in this paper is a generic simplification that can act as a drop-in replacement for any symbolic parser. Our idea of replacing unknown words with known, similar words results in small but significant improvements in experiments using two parsers and for a range of 7 languages.
Many text corpora exhibit socially problematic biases, which can be propagated or amplified in the models trained on such data. For example, doctor cooccurs more frequently with male pronouns than female pronouns. In this study we (i) propose a metric to measure gender bias; (ii) measure bias in a text corpus and the text generated from a recurrent neural network language model trained on the text corpus; (iii) propose a regularization loss term for the language model that minimizes the projection of encoder-trained embeddings onto an embedding subspace that encodes gender; (iv) finally, evaluate efficacy of our proposed method on reducing gender bias. We find this regularization method to be effective in reducing gender bias up to an optimal weight assigned to the loss term, beyond which the model becomes unstable as the perplexity increases. We replicate this study on three training corpora---Penn Treebank, WikiText-2, and CNN/Daily Mail---resulting in similar conclusions.
The emergence of deep learning as a commanding technique for learning heterogeneous layers of feature representations have consequently substituted traditional machine learning algorithms which are generally poor in analyzing compound sentences. Additionally, convolutional and recurrent neural networks have auspiciously yielded state-of-the-art results in sentiment classification and Natural Language Processing (NLP). In this paper, a deep sentiment representation model through the combination of multiple Convolutional Neural Networks (CNN) kernels with Long Short-Term Memory (LSTM) is proposed for sentiment classification. Our model gains word vector representation using pre-trained Global Vectors for Word Representation (GloVe) embeddings, thereafter used as input to the CNN layer which extracts higher local text representations. Finally, Bidirectional LSTM (biLSTM) generates sentiment classification of sentence representation based on context dependent features. Our combined approach of CNN and biLSTM was experimented using the Stanford Large Movie Review Dataset (IMDB) and Stanford Sentiment Treebank Dataset (SSTB) for binary classification. The evaluation achieves outstanding results in outperforming several existing approaches with 90.4% accuracy on the Stanford Sentiment Treebank dataset and 94.8% accuracy on the Stanford Large Movie Review dataset. These results are achieved with a drastic reduction of model parameters and without a pooling layer in the CNN architecture, helping to retain local and structural information in comparison to other existing deep neural network frameworks.
In sequence learning tasks such as language modelling, Recurrent Neural\nNetworks must learn relationships between input features separated by time.\nState of the art models such as LSTM and Transformer are trained by\nbackpropagation of losses into prior hidden states and inputs held in memory.\nThis allows gradients to flow from present to past and effectively learn with\nperfect hindsight, but at a significant memory cost. In this paper we show that\nit is possible to train high performance recurrent networks using information\nthat is local in time, and thereby achieve a significantly reduced memory\nfootprint. We describe a predictive autoencoder called bRSM featuring recurrent\nconnections, sparse activations, and a boosting rule for improved cell\nutilization. The architecture demonstrates near optimal performance on a\nnon-deterministic (stochastic) partially-observable sequence learning task\nconsisting of high-Markov-order sequences of MNIST digits. We find that this\nmodel learns these sequences faster and more completely than an LSTM, and offer\nseveral possible explanations why the LSTM architecture might struggle with the\npartially observable sequence structure in this task. We also apply our model\nto a next word prediction task on the Penn Treebank (PTB) dataset. We show that\na 'flattened' RSM network, when paired with a modern semantic word embedding\nand the addition of boosting, achieves 103.5 PPL (a 20-point improvement over\nthe best N-gram models), beating ordinary RNNs trained with BPTT and\napproaching the scores of early LSTM implementations. This work provides\nencouraging evidence that strong results on challenging tasks such as language\nmodelling may be possible using less memory intensive, biologically-plausible\ntraining regimes.\n
Throughout many studies which focus on brain laterality as a key component to the outcome of an experiment, or to a participants’ reaction to a stimulus, it can be noted that the different areas of the brain involved in a task response must work together to produce a viable outcome (i.e. lateralized brain processes). While there are laterality components in relation to the cognitive processes of handed and footed responses, it is still largely unknown how the different areas of the brain interpret the emotional stimuli to then affect these outcomes. The purpose of our study is to determine how emotional context affects a simple cognitive task that includes handed and footed responses, and if any observed differences can be traced back to the different systems at work within the brain. Subjects will be tested over two days for handed and footed responses in a cognitive Simon Task. Subjects will be tested with and without emotional context (i.e. a background image of a specific valence and arousal rating), and any resulting differences between non-emotional context and emotional context reaction times will be compared.
Do individual sounds carry meaning? The relationship between sound and meaning in human languages is typicallyassumed to be arbitrary, though recent research provides evidence for the existence of both iconicity and systematicitybetween word forms and their meaning. However, this research has not asked whether individual sounds in a languagecovary in systematic ways with aspects of meaning. In two analyses, we find evidence for more systematicity betweenthe initial phones of words and those words concreteness ratings than one would expect in a truly arbitrary lexicon. Thissuggests that initial phones may act as cues to aspects of word meaning, and raises questions about whether languagelearners detect and exploit these cues.
This talk will explore the role of individual social actors and their communities and networks in the formation, maintenance and dissolution of linguistic norms. It will consider several standardisation episodes in the history of Old and Early Middle English and connect them to other unification processes in the political and cultural history of England at the time. It will be suggested that the suppression of variability on the linguistic level often accompanies, or is a symptom of, a similar suppression on the ideological level. Political, religious, legal, and linguistic processes mingle in various ways in this period (as they do today) to replicate and enhance the social order, to support a reform movement, or to refute dissent. \nThree case studies will offer insights into these processes: 1) shire courts and the ‘standardised’ lexis of the Anglo-Saxon Chronicle in the reign of King Alfred and Edward the Elder; 2) chancery norms and charters of the eleventh century; and 3) religious reform and linguistic focusing in the thirteenth century.
Abstract This chapter is devoted to the presentation of the tools and methods used for the different steps of the semi-automatic syntactic annotation: automatic preprocessing; microsyntactic parsing with the FRMG tool, correction of the parsing with the Arborator tool, agreement analysis, post-validation correction, and development of the final format of the Rhapsodie syntactic treebank. As FRMG is a parser for written French that was not configured to analyze disfluencies and reformulation, we used our manual pile marking to unfold the piles and produce a series of simplified “sentences” with only government relations. Despite having two annotators plus a validator for the corrections, we found a substantial number of errors in the post-validation procedure by using a set of rules to determine the well-formedness of the trees.
Contextualized embeddings, which capture appropriate word meaning depending\non context, have recently been proposed. We evaluate two meth ods for\nprecomputing such embeddings, BERT and Flair, on four Czech text processing\ntasks: part-of-speech (POS) tagging, lemmatization, dependency pars ing and\nnamed entity recognition (NER). The first three tasks, POS tagging,\nlemmatization and dependency parsing, are evaluated on two corpora: the Prague\nDependency Treebank 3.5 and the Universal Dependencies 2.3. The named entity\nrecognition (NER) is evaluated on the Czech Named Entity Corpus 1.1 and 2.0. We\nreport state-of-the-art results for the above mentioned tasks and corpora.\n
We present a recurrent neural network memory that uses sparse coding to\ncreate a combinatoric encoding of sequential inputs. Using several examples, we\nshow that the network can associate distant causes and effects in a discrete\nstochastic process, predict partially-observable higher-order sequences, and\nenable a DQN agent to navigate a maze by giving it memory. The network uses\nonly biologically-plausible, local and immediate credit assignment. Memory\nrequirements are typically one order of magnitude less than existing LSTM, GRU\nand autoregressive feed-forward sequence learning models. The most significant\nlimitation of the memory is generalization to unseen input sequences. We\nexplore this limitation by measuring next-word prediction perplexity on the\nPenn Treebank dataset.\n
In this paper, we discuss constituent ordering generalizations in Japanese. Japanese has SOV as its basic order, but a significant range of argument order variations brought about by ‘scrambling’ is permitted. Although scrambling does not induce much in the way of semantic effects, it is conceivable that marked orders are derived from the unmarked order under some pragmatic or other motivations. The difference in the effect of basic and derived order is not reflected in native speaker’s grammaticality judgments, but we suggest that the intuition about the ordering of arguments may be attested in corpus data. By using the Keyaki treebank (a proper subset of which is NINJAL Parsed Corpus of Modern Japanese (NPCMJ)), it is shown that the naturally-occurring corpus data confirm that marked orderings of arguments are less frequent than their unmarked ordering counterparts. We suggest some possible motivations lying behind the argument order variations.
Classic natural language processing resources such as the Penn Treebank (Marcus et al. 1993) have long been used both as evaluation data for many linguistic tasks and as training data for a variety of off-the-shelf language processing tools. Recent work has highlighted a gender imbalance in the authors of this text data (Garimella et al. 2019) and hypothesized that tools created with such resources will privilege users from particular demographic groups (Hovy and Søgaard 2015). Domain adaptation is typically employed as a strategy in machine learning to adjust models trained and evaluated with data from different genres. However, the present work seeks to evaluate whether domain adaptation to demographic groups such as age or gender may be an effective strategy to ameliorate the effects of biased or outdated training corpora in linguistic preprocessing tasks. We find adaptation to demographic groups to be an effective strategy for improving preprocessing performance across all demographic groups.
Mirror-sensory synesthetes mirror the pain or touch that they observe in other people on their own bodies. This type of synesthesia has been associated with enhanced empathy. We investigated whether the enhanced empathy of people with mirror-sensory synesthesia influences the experience of situations involving touch or pain and whether it affects their prosocial decision making. Mirror-sensory synesthetes (<i>N</i> = 18, all female), verified with a touch-interference paradigm, were compared with a similar number of age-matched control individuals (all female). Participants viewed arousing images depicting pain or touch; we recorded subjective valence and arousal ratings, and physiological responses, hypothesizing more extreme reactions in synesthetes. The subjective impact of positive and negative images was stronger in synesthetes than in control participants; the stronger the reported synesthesia, the more extreme the picture ratings. However, there was no evidence for differential physiological or hormonal responses to arousing pictures. Prosocial decision making was assessed with an economic game assessing altruism, in which participants had to divide money between themselves and a second player. Mirror-sensory synesthetes donated more money than non-synesthetes, showing enhanced prosocial behaviour, and also scored higher on the Interpersonal Reactivity Index as a measure of empathy. Our study demonstrates the subjective impact of mirror-sensory synesthesia and its stimulating influence on prosocial behaviour.This article is part of the discussion meeting issue ‘Bridging senses: new developments in synaesthesia’.
In two experiments, the influence of inducing negative mood on cognitive performance was explored by analyzing physical arm reaching movements as indicators of mind wandering. Mood was induced by viewing a series of six photos per mood condition that were previously established for their emotionally valenced and arousal ratings. A reach tracking device recorded three metrics of arm movement that were expected to reflect instances of mind wandering: initiation latency, movement time, and arm curvature. In the first experiment, 29 participants were randomly assigned into one of two induced-mood groups, negative mood (n = 15) or neutral mood (n = 14). Participants performed a simple Go/No-go task in which arm movements were detected by the reach tracker. The first experiment indicated that the mood inducement was successful but the effect of negative mood on either self-reported mind wandering or variances in arm movement were not significant. Thus, the second experiment prompted the change to a visual-search target-selection task in which variances in initiation latency, movement time, and curvature were expected to be more pronounced. The second experiment consisted of 23 participants who were also randomly assigned to either negative (n = 12) or neutral (n = 11) mood condition. The second experiment revealed that the mood induction was still successful but that there were still no significant effects observed between mood and indicators of mind wandering. Though the results of this study did not reflect initial predictions, it may suggest that low-arousing negative moods in healthy individuals are not associated with increased mind wandering.
Antiaddictive social advertising is a special speech genre of modern communication with specific features determined by the target setting and the chosen strategy. Advertising can be considered as a special functional style, within which separate genres are distinguished, first of all commercial advertising and social advertising. Social advertising, which refers to ethical categories, has become especially popular because society is faced with such problems, the solution of which depends on mass behavior. Advertising has become an integral part of the daily life of a person. Since advertising is mass-replicated in the media, it can enter the consciousness of the addressee, even against his/her will and desire, no wonder advertising is sometimes defined as the “fifth power” after media, whose power is considered the “fourth power”. Therefore, it is so important for advertising to follow the moral principles of society, it is so important to comply with modern ethical and linguistic norms. If commercial advertising is widespread, the antiaddictive one is little known. Many specialists in advertising argue for the need of the strategy of “shock” advertising, a remarkable feature of which is hyperbolization, used to cause fear. In antiaddictive advertising there is often a morphological imperative in the meaning of categorical motivation. It is justified as such advertizing has to be as much as possible appellate, has to be understood unambiguously. Anti-drug social advertising should show positive, motivation to a healthy lifestyle.
Abstract: Temperatures above 20° Celsius have shown to adversely impact human behavior, leading to increased aggression and violence. Climate change will contribute to both the magnitude and severity of this pattern as temperatures continue their rise. Contributions to this field of research have only recently begun to analyze online behavior and language as a proxy for hedonic state, or well-being. From a development perspective this study is relevant since the poor tend to live in some of the warmest regions on earth, and would thus be disproportionately impacted by increased temperatures. We use several sources of data; U.S. based daily statewide temperature data from 2016 through 2017, as well as localized viewer chat data from a live video streaming website. We will sort chatting comments looking for key words (i.e. hate speech, swearing, etc.), and with the use of a word rating system we then assess the overall mood of the chatters contingent on high temperature readings on the precise day of the communications. After controlling for spatiotemporal fixed effects, we find strong evidence that hedonic state decreases above 20°c.
Death is a vague, frightening and abstract concept, mostly considered as a taboo. The current study investigates the ways of conceptualizing death among Persian language speakers in a representative corpus of Persian texts. It seeks answer to the question: What kind of linguistic tools are used by Persian language speakers in order to conceptualize the phenomenon of death and what cultural elements are involved in this regard.To answer the research question, we used the Persian Linguistic Database (PLDB) and a selection of obituaries and epitaphs as data and attempted to identify and extract all the expressions which were directly or indirectly related to the concept of death. Using the conceptual metaphor theory (CMT), by Lakoff and Johnson (1987), and cultural conceptualization by Sharifian (2011), the ways death is conceptualized were identified on the basis of a cognitive-cultural approach.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
Like odor identification, remote odor memory, reflected in familiarity ratings, is impaired in AD (Murphy, Nature Reviews Neurology, 2019). We investigated the relative abilities of standard screening (MMSE), odor identification and remote odor memory to predict transitions from amnestic MCI (aMCI) to AD in a sample from the UCSD ADRC. The sample contained 170 controls, 210 AD, 26 aMCI converters to AD and 42 aMCI non-converters. A receiver operating characteristic (ROC) curve plots the trade-off between sensitivity and specificity. The area under the curve (AUC) indicates how well a marker discriminates patients from controls. Analyses showed higher predictive value for converting from aMCI to AD in ApoE ε4+ carriers for odor familiarity, odor identification and for the combination than for the MMSE. ROC/AUCs for the conversion from aMCI to AD have ranged from.63 -.67 for CSF biomarkers. Odor familiarity and odor identification had similar AUC values; however, combining odor familiarity and odor identification produced an ROC/AUC value of 1.0 in ε4 carriers, appreciably higher than for MMSE alone (.58). Olfactory biomarkers show real promise as early, non-invasive indicators of disease, particularly in samples enriched with ε4 carriers. Although odor identification has been the focus of olfactory biomarker work, the results suggest that other measures of olfactory function have the potential to enhance prediction. Combining odor familiarity and odor identification produced a predictive value of 1.0 in ε4 carriers. The results warrant further investigation into the potential for enhancing drug trials and clinical screening. Supported by NIH grants R01AG004085-26 (CM) and P50AG005131 (UCSD ADRC). We thank the UCSD ADRC and particularly Drs. Douglas Galasko and David Salmon.
Cet article presente la creation d’un treebank journalistique serbe, ParCoJour. Il est compose de 30K tokens et dote de trois couches d’annotation: etiquetage morphosyntaxique, lemmatisation et annotation syntaxique. Une fois construit, ParCoJour a ete utilise dans trois experiences afin d’evaluer l’impact du domaine textuel sur le parsing du serbe en comparant les performances de Talismane, un systeme par apprentissage automatique, sur deux types de corpus, journalistique et litteraire: 1) parsing du corpus journalistique avec un modele entraine sur le corpus journalistique; 2) parsing du corpus journalistique avec un modele entraine sur le corpus litteraire; 3) parsing du corpus litteraire avec un modele entraine sur le corpus journalistique. Les resultats sont compares a ceux ou les deux corpus relevaient du domaine litteraire. Le changement de domaine textuel dans la deuxieme et la troisieme experience entraine une baisse des performances, mais les resultats de parsing restent satisfaisants.
The article discusses the current changing of linguistic norms in English as a lingua franca of global communication nowadays. It aims at both determining the causes of language deviations and analyzing language errors as well as their impact on the effectiveness of the English language communication. Based on the analysis of abundant empirical material, we prove that language innovations are caused by the immanent structural, functional, and pragmatic variability / instability of the English language; they are also associated with cognitive and sociocultural evolution. The research methodology includes: a corpus-based analysis of speech errors; interpretative, context and discourse analyses of the sources of language errors, as well as their distribution, adaptation, habitualization, legitimization, and regulation. We discuss the degree of influence of these processes on native and non-native speakers. Special attention is paid to multilingual interference and the Internet language creation. The findings show that it is impossible to separate language errors from language innovations today. Such conventional governing principles of error normalization as credibility, codification, and approval are still playing an important role while the demographic and geographic principles are losing their significance. The Internet communication often proliferates error normalization processes, which result in evaluating (accepting or rejecting) any innovation according to the principle of “virtual validity”. In conclusion, the English language status as a language of the international communication significantly transforms its norms, rules, and traditions. We think, this will not worsen it, but allow people of different nationalities to communicate in English more effective using their “English variant”, which is the most adapted one to their cognitive, functional and pragmatic needs.
It is common knowledge and prescribed in all normative Portuguese grammars that the verb must agree in number and person with its subject, whether the latter is superposed or postponed to the verb. The lack of agreement between subject and verb (concordância verbal varíavel) is seen by users with a scholastic education as being wrong and linked to the poorest social strata, that is, with a low or zero level of education. However, cases of lack of agreement are not uncommon in informal speech of users with medium or high levels of education. This dichotomy between linguistic norms and orality, and its perception by the Brazilian population (i.e. lack of agreement as a sign of a lower educational and social level) can be verified also in the artistic reproduction, namely in the filmic dialogues of Brazilian national cinema, where verbal agreement variation is used to typify characters with little or no education. This paper will attempt to analyse how this linguistic phenomenon is interpreted by film discourse, i.e. a reproduction of orality. First, the linguistic issue will be presented, that is, the agreement and the lack of it in some registers of Brazilian Portuguese, and a brief presentation of the type of data on which this research was conducted (filmic dialogues from Brazilian films). Then, cases of variable verbal agreement (hereafter CVV) of the first plural person (hence 1PP) will be presented. In this perspective, the analysis of the alternation of use between two pronominal forms in subject function for the 1PP will be considered as a possible cause for the variation of verbal agreement between verb and subject with the 1PP. The approach adopted in this research brings together the variational studies and tools of Corpora Linguistics in an attempt to offer a critical view of the linguistic choices involved in film production.
English Abstract: The article analyses the relationship between etymology of political neologisms, ways of their formation and methods of rendering them from English into Russian. Translation strategy is directed at the recipient of the translated text and should therefore be pragmatic and based on the functional and stylistic norms of the translation language. The author concludes that the most efficient way of word-formation in the English language is compounding (morphological neologisms) and derivation of new meaning for the already existing words (semantic neologisms). The translation methods demonstrating high potential are transliteration, calquing (loan translation) and descriptive translation. Adequate translation is based on the following criteria: conciseness, single interpretation, conformity with the linguistic norms of the translation language. Russian Abstract: В работе затрагиваются вопросы взаимосвязи между этимологией, способами образования политических неологизмов и приемами их передачи с английского языка на русский. Выбор переводческой стратегии ориентирован на получателя текста перевода, поэтому должен обеспечивать прагматические задачи в соответствии с функционально-стилистическими нормами языка перевода. В работе делается вывод о том, что наиболее продуктивным способом образования новых слов является процесс словосложения (морфологические неологизмы) и наделения уже существующих в языке слов новым значением (семантические неологизмы); самыми эффективными приемами перевода неологизмов являются транслитерация, калькирование и описательный перевод. Критериями адекватного перевода политических неологизмов служат краткость, однозначность толкования, соответствие нормам языка перевода.
This contains data files needed for FinMeter. This data is complementary for FinMeter Python library described in: Mika Hämäläinen and Khalid Alnajjar (2019). Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing. In <em>the Proceedings of The 12th International Conference on Natural Language Generation</em>. Sources: The pretrained vectors for Finnish (es - I know) and English (en) are from E. Grave, P. Bojanowski, P. Gupta, A. Joulin, T. Mikolov, <em>Learning Word Vectors for 157 Languages. Creative Commons Attribution-Share-Alike License 3.0</em>. See https://fasttext.cc/docs/en/crawl-vectors.html The word2vec model trained on the Finnish Internet ParseBank is from Kanerva, Jenna; Luotolahti, Juhani; Laippala, Veronika; Ginter, Filip: Syntactic N-gram Collection from a Large-Scale Corpus of Internet Finnish. Proceedings of the Sixth International Conference Baltic HLT. 2014. paper. Creative Commons Attribution-ShareAlike 4.0 International License. See http://bionlp.utu.fi/finnish-internet-parsebank.html The Finnish concreteness data has been automatically translated from Brysbaert, Marc, Amy Beth Warriner, and Victor Kuperman. "Concreteness ratings for 40 thousand generally known English word lemmas." <em>Behavior research methods</em> 46.3 (2014): 904-911. Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. see http://crr.ugent.be/archives/1330
The present study examined the effects of labeling on pain tolerance, sensation, and affect for individuals who are high or low pain catastrophizers, as measured through the Pain Catastrophizing Scale (PCS). Pasticipants completed the PCS and were randomly assigned to 1 of 3 labeling conditions: A maximizing, a minimizing, and a neutral label condition. All participants then took part in a cold-pressor test. The cold-pressor measure of pain tolerance, as well as visual analog scales of sensory and affective ratings of pain, provided the dependent measures. Participants also completed the Anxiety Sensitivity Index (ASI) and the Somatic Amplification Questionnaire (SAQ). Results indicated that high pain catastrophizers have significantly reduced pain tolerance, increased pain sensations, and increased pain unpleasantness compared with low pain catastrophizers. In addition, significant correlations were found between the dependent measures. Main effects for labeling, and interaction effects between pain catastrophizing and labeling, were not supported.
This contains data files needed for FinMeter. This data is complementary for FinMeter Python library described in: Mika Hämäläinen and Khalid Alnajjar (2019). Let's FACE it. Finnish Poetry Generation with Aesthetics and Framing. In <em>the Proceedings of The 12th International Conference on Natural Language Generation</em>. Sources: The pretrained vectors for Finnish (es - I know) and English (en) are from E. Grave, P. Bojanowski, P. Gupta, A. Joulin, T. Mikolov, <em>Learning Word Vectors for 157 Languages. Creative Commons Attribution-Share-Alike License 3.0</em>. See https://fasttext.cc/docs/en/crawl-vectors.html The word2vec model trained on the Finnish Internet ParseBank is from Kanerva, Jenna; Luotolahti, Juhani; Laippala, Veronika; Ginter, Filip: Syntactic N-gram Collection from a Large-Scale Corpus of Internet Finnish. Proceedings of the Sixth International Conference Baltic HLT. 2014. paper. Creative Commons Attribution-ShareAlike 4.0 International License. See http://bionlp.utu.fi/finnish-internet-parsebank.html The Finnish concreteness data has been automatically translated from Brysbaert, Marc, Amy Beth Warriner, and Victor Kuperman. "Concreteness ratings for 40 thousand generally known English word lemmas." <em>Behavior research methods</em> 46.3 (2014): 904-911. Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. see http://crr.ugent.be/archives/1330
The present study explores the images of standard Lithuanian of young people at a gymnasium located in the area of the standard language. The data were obtained from a questionnaire based on the methodological principles of perceptual dialectology. The image of the standard language in the consciousness of the respondents emerges from the analysis of the questionnaire data: the frequency of linguistic codes, the mental maps of the standard language areas, and associations of the standard language. The analysis of the data shows that the gymnasium students tend to distance themselves from the regional linguistic code. The respondents’ disassociation from the local variety and their stronger preference for the code of the standard language is probably related to their sense of language security in the area of the linguistic homeland (including that of the standard language). The mental maps show that the young people associate the standard Lithuanian with the larger or smaller area of central Lithuania, which includes cities (Kaunas, Vilnius), adjacent non-dialect areas (Jonava, Kaišiadorys), and one or two dialect zones; it nearly overlaps the area of the standard language delineated in the second decade of the twenty-first century. Vilnius is the part of this image – probably of its status of the capital city and a significant social, cultural, and urban centre of attraction. The gymnasium students think that speakers of the standard language are city dwellers first and foremost, while the mental connection between the code of standard language and education occurs less often. Such views might have emerged due to the location of the city – hence that of the respondents’ linguistic homeland. Identifying the standard language user as an ordinary person or a Lithuanian could most likely be explained by the fact that the standard language is not only a national language to the young people: it is also an equivalent of their linguistic code. The gymnasium students do not associate the standard language with linguistic norms (the correct use of language). The consistency of the young generation’s attitudes (both those visualised on the maps and verbalised in the questionnaire answers) suggests the high value of the variety spoken in the area they associate with the standard language. The results of the study provide insights into the functioning, vitality, and continuity of the standard language in this area.
Sentiment Analysis is an application of Natural Langue Processing to analyze social media corpora to extract insights of corpora. Sentiment analytical results are the real feedback of the customers, which enables the organizations and companies to take appropriate decision on their products and business policies. Stemming plays in-evitable and vital role in sentiment analysis. Stemming is one of the phase of preprocessing the social media corpora. Today most of the researches uses strong stemmers to identify stem words of social media corpora. The most popular stemming algorithms such as Lancaster and Porter stemming algorithms causes prejudiced the meaning of the words. The over-stemmed words mislead the sentiment classification process. To prevent the over-stemming the Unprejudiced lighter stemming algorithm is proposed to sustain the meaning of the stemmed words. The propose Un-prejudiced algorithm uses lexical database and Parts of speech of Python Natural Language Tool Kit. There are a few stemming algorithm accuracy evaluation methods, in this paper we focused on Paice Error-rate relative to truncation (ERRT) measure to evaluate the accuracy of Lancaster, Porter and Unprejudiced stemming algorithms. The experiments were conducted on 25,758 source words and results were evaluated using Paice stem evaluation method and Sirsat method. The Paice Evaluation ERRT values 0.47209, 0.28703, 0.15502 of Lancaster, Porter, Unprejudiced respectively are proved that the Unprejudiced stemmer is more accurate than Lancaster and Porter. Sirsat’s stem evaluation method Average Words Conflation Factor (AWCF) results 10310.31, 14031.17, 23349.87 of Lancaster, Porter, Unprejudiced respectively are also proved the Unprejudiced stemming algorithm is more accurate than Lancaster and Porter stemming algorithms.
Abstract This chapter discusses the positioning of Belarus in the international context of socioeconomic development based on an assessment of the country's dynamics in world rankings. The country's presence in the recognized world rankings and its holding high positions in them is an obvious advantage for achieving a favorable investment image. Ratings characterize the country's comparative position at the international level in a number of areas: from credit capacity to human capital development. There has been analyzed the position of the Republic of Belarus in several recognized international comparisons, such as Human Development Index, Doing Business, ICT Development Index, Global Innovation Index, Sustainable Development Goals Index, Corruption Perceptions Index, Rule of Law Index, Worldwide Governance Indicators, and others. However, Belarus is not yet participating in the international competitiveness assessment through such popular international ratings as Global Competitiveness Index and Global Entrepreneurship Monitor. The research findings show that the strongest aspects of the socioeconomic development of Belarus are in place due to the high educational level of the human capital development, gender equality, and the implementation of the UN sustainable development goals. The analysis also shows that the weaknesses of institutional environment and public administration do not enable the full implementation of the planned goals of socioeconomic development.
Objects with sharp contours are preferred less than objects with smooth contours, as sharp angles are thought to be an indicator of threat. In two experiments, we probe the link between low level visual features, such as contour curvature, and affective ratings. In Experiment 1, we used artist-traced line drawings of all images from the International Affective Picture System (IAPS) image set. We computationally extracted the contour curvature, length, and orientation statistics of all images, and explored whether these features are predictive of emotional valence scores. Our results replicate previous research, finding a significant negative relationship between high curvature (i.e., angularity) and emotional valence (p = 0.012). Additionally, we find that length is positively related to valence, such that images containing long contours are rated as more positive (p = 0.049). In Experiment 2, we composed new, content-free line drawings of contours with different combinations of length, curvature, and orientation values. Sixty-seven participants were presented with these images on Amazon Mechanical Turk (MTurk) and had to categorize them as positive or negative. A linear mixed effects model revealed that low curvature, long, horizontal contours predicted participants’ positive responses, while short, high curvature contours predicted participants’ negative responses. Taken together, these findings have implications for theories of threat detection such as the Snake Detection Theory, which posits that humans evolved to fear snakes and thus our thalamic nuclei are able to detect snakes rapidly and automatically. It is unlikely, however, that the thalamus has a true representation of “snake”. Our findings suggest a more plausible scenario, whereby visual features associated with threatening stimuli, such as snakes, are quickly detected and passed on to visual cortex for further processing. We have also identified the low-level contour features that are associated with positive valence.
This dissertation investigates the frequency, semantic, and functional characteristics of recurring discontinuous formulaic language in a learner corpus of argumentative and literary essays. Discontinuous sequences of words, or ‘frames’, are recurrent sequences of words that have one or more variable slots. For example, in the * of and it is * to, where the asterisks represent variable slots in the sequences of words. A corpus of English argumentative essays authored by native speakers of English, Japanese, and Spanish is analyzed using modern methods in corpus linguistics to determine which frames are used frequently. Frequent frames are compared between the L1 groups to investigate trends in structure, frequency, and variability.\nThe focus of analysis then narrows to a group of 30 recurring frames comprised of only function words, that is, function word frames such as in the * of, the * of the, and to the * that. The 30 frames are grouped based on structural characteristics, such as noun and preposition-based frames, for further analysis to better understand their semantic and functional characteristics. A lexical database is used to explore the semantic characteristics of fillers of the frames. To determine discourse functions of all instances of each of the 30 target frames, a well-known taxonomy previously applied to continuous sequences of words is adapted and applied to the present context. Discourse functions of the frames are then used as dependent variables in a multinomial logistic regression conducted with four distinct predictor variables: (1) L1, (2) proficiency level, (3) topic of essay, and (4) specific frame. The purpose of the regression is to see which predictor best accounts for discourse function fulfilled by frames from the structural groups.\nFindings from the various analyses first indicate that Japanese learners of English use function word frames at far lower rates than the L1 English and Spanish speakers. Secondly, fillers of the structural groups of function word frames tend to be abstract nouns and the frames largely serve the discourse function of intangible framing of a following noun phrase. In terms of predictor variables, the frames themselves, particularly prepositions, best predict discourse function. The results lend support to the idea that function words, despite carrying little meaning in isolation, are semantically motivated and systematically contribute meaning to larger sequences of words. Pedagogical implications of this study include the teaching of function words from a more phraseological perspective as well as highlighting the connection between frames, fillers, and discourse functions.