Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This thesis studies the connections between parsing friendly representations and interlingua grammars developed for multilingual language generation. Parsing friendly representations refer to dependency tree representations that can be used for robust, accurate and scalable analysis of natural language text. Shared multilingual abstractions are central to both these representations. Universal Dependencies (UD) is a framework to develop cross-lingual representations, using dependency trees for multlingual representations. Similarly, Grammatical Framework (GF) is a framework for interlingual grammars, used to derive abstract syntax trees (ASTs) corresponding to sentences. The first half of this thesis explores the connections between the representations behind these two multilingual abstractions. The first study presents a conversion method from abstract syntax trees (ASTs) to dependency trees and present the mapping between the two abstractions – GF and UD – by applying the conversion from ASTs to UD. Experiments show that there is a lot of similarity behind these two abstractions and our method is used to bootstrap parallel UD treebanks for 31 languages. In the second study, we study the inverse problem i.e. converting UD trees to ASTs. This is motivated with the goal of helping GF-based interlingual translation by using dependency parsers as a robust front end instead of the parser used in GF. \n\nThe second half of this thesis focuses on the topic of data augmentation for parsing – specifically using grammar-based backends for aiding in dependency parsing. We propose a generic method to generate synthetic UD treebanks using interlingua grammars and the methods developed in the first half. Results show that these synthetic treebanks are an alternative to develop parsing models, especially for under-resourced languages without much resources. This study is followed up by another study on out-of-vocabulary words (OOVs) – a more focused problem in parsing. OOVs pose an interesting problem in parser development and the method we present in this paper is a generic simplification that can act as a drop-in replacement for any symbolic parser. Our idea of replacing unknown words with known, similar words results in small but significant improvements in experiments using two parsers and for a range of 7 languages.
This paper suggests annotation guidelines to build a Universal Dependencies (UD) treebank for Korean. We discuss the part-of-speech annotation of Korean specific-categories such as prenouns, numeral classifiers, and (pre)final endings, and propose how to implement UD scheme in Korean regarding selecting a head and assigning dependency relations to dependents. UD prioritizes content words over functional words since the former exhibits less cross-linguistic variations. In a noun phrase, for instance, a core noun is always a head of the entire noun phrase independently of a language. The rest are treated as a dependent: not only a modifier such as an adjective but also a functional category such as an article, numeral quantifier, demonstrative, and so on. However, when it comes to head-less constructions such as coordination or predicate ellipsis, UD firmly advocates the head-initial strategy. The present application of UD to Korean tries to follow UD’s principles as much as possible. Korean is a head-final language, so that headed constructions are analyzed head-finally. In contrast, head-less ones are tagged head-initially. This might disregard language-specific characteristics from a linguistic perspective, but the strategy allows us to build up a set of treebanks in a cross-linguistically consistent way (i.e., the fundamental purpose of UD).
In two experiments, the influence of inducing negative mood on cognitive performance was explored by analyzing physical arm reaching movements as indicators of mind wandering. Mood was induced by viewing a series of six photos per mood condition that were previously established for their emotionally valenced and arousal ratings. A reach tracking device recorded three metrics of arm movement that were expected to reflect instances of mind wandering: initiation latency, movement time, and arm curvature. In the first experiment, 29 participants were randomly assigned into one of two induced-mood groups, negative mood (n = 15) or neutral mood (n = 14). Participants performed a simple Go/No-go task in which arm movements were detected by the reach tracker. The first experiment indicated that the mood inducement was successful but the effect of negative mood on either self-reported mind wandering or variances in arm movement were not significant. Thus, the second experiment prompted the change to a visual-search target-selection task in which variances in initiation latency, movement time, and curvature were expected to be more pronounced. The second experiment consisted of 23 participants who were also randomly assigned to either negative (n = 12) or neutral (n = 11) mood condition. The second experiment revealed that the mood induction was still successful but that there were still no significant effects observed between mood and indicators of mind wandering. Though the results of this study did not reflect initial predictions, it may suggest that low-arousing negative moods in healthy individuals are not associated with increased mind wandering.
Modern machine learning (ML) techniques are transforming many disciplines ranging from transportation to healthcare by uncovering patterns in data, developing autonomous systems that mimic human abilities, and supporting human decision-making. Modern ML techniques, such as deep neural networks, are fueling the rapid developments in artificial intelligence. Engineering design researchers have increasingly used and developed ML techniques to support a wide range of activities from preference modeling to uncertainty quantification in high-dimensional design optimization problems. This special issue brings together fundamental scientific contributions across these areas.The special issue consists of 24 papers spread over two issues of the Journal of Mechanical Design. The papers use various ML techniques, including artificial neural networks, Gaussian processes, reinforcement learning, clustering techniques, and natural language processing. Based on their research objective, the papers can be broadly classified into four groups: (i) ML to support surrogate modeling, design exploration, and optimization, (ii) ML for design synthesis, (iii) ML for extracting human preferences and design strategies, and (iv) comparative studies of ML techniques and research platforms to help design researchers. The papers are summarized in Secs. 1–4. An analysis of the themes covered in the special issue and the potential opportunities for future research in ML for Engineering Design are presented in Sec. 5.In the paper titled Multifidelity Physics-Constrained Neural Network and Its Application in Materials Modeling, Liu and Yang address how to incorporate multifidelity, physics-based constraints into neural network predictions. The paper contributes two key insights. First, the paper extends existing Physics-Constraints Neural Network architectures by imposing a multifidelity constraint scheme wherein an auxiliary network minimizes discrepancies between low and high fidelity models—essentially learning how to correct the low-fidelity one. Second, it proposes an adaptive weighting scheme to control the convergence of individual losses among the different fidelities. They demonstrate the impact of these improvements on several fundamental multiscale material modeling challenges including two-dimensional heat transfer, phase transition, and dendritic growth problems. On these problems, the proposed multifidelity, physics-based constraints decrease the prediction error up to order of magnitude compared with networks without such constraints. This achieves comparable accuracy to that of direct numerical solutions of the underlying equations.Sarkar et al. present a multifidelity modeling and information-theoretic sequential sampling strategy for optimization in their paper titled Multifidelity and Multiscale Bayesian Framework for High-Dimensional Engineering Design and Calibration. The approach is based on modeling of the varied fidelity information sources via Gaussian processes, augmented with efficient active learning strategies that involve sequential selection of optimal points in a multiscale architecture. The strategy is demonstrated using the design optimization of a compressor rotor and calibration of a microstructure prediction model.In the paper titled A Case Study of Deep Reinforcement Learning for Engineering Design: Application to Microfluidic Devices for Flow Sculpting, Lee et al. address how to design micro-fluidic flow sculpting devices by overcoming some of the key weaknesses of evolutionary optimization-based methods, namely, poor sample efficiency and slow optimization convergence. The paper adapts deep reinforcement learning (DRL) techniques to the flow sculpting task and also studies the effectiveness of transfer learning on accelerating the design of target flow shapes. The paper demonstrates that DRL is able to match 90% of the target flow shapes using significantly fewer sculpting pillars than comparable GA models as well as provides a means to interpret the learned model (using Principal Components) that existing approaches to fluidic sculpting do not provide.Lynch et al., in their paper Machine Learning to Aid Tuning of Numerical Parameters in Topology Optimization, present an ML-based meta-learning framework to determine tuning parameters in topology optimization. The parameters are learned from similar optimization problems carried out in the past and adjusted for the problem at hand. This helps in avoiding costly trial-and-error involved in manual parameter tuning.In the paper Data-Driven Design Space Exploration and Exploitation for Design for Additive Manufacturing, Xiong et al. present a data-driven approach for design search and optimization at successive stages in the design process. They use Bayesian network classifier in the embodiment design stage and Gaussian process regression in the detailed design phase. The approach is illustrated in the paper through the design of a customized ankle brace design.Odonkor and Lewis apply data-driven design to the design of operational strategies of complex systems, specifically distributed energy resources. The paper is titled Data-Driven Design of Control Strategies for Distributed Energy Systems. The problem of maximizing arbitrage value is formulated as an optimization problem and solved using reinforcement learning. The approach is demonstrated for shared distributed energy resources in multi-building residential clusters.In Globally Approximate Gaussian Processes for Big Data With Application to Data-Driven Metamaterials Design by Bostanabad et al., a globally approximate Gaussian process (GAGP) is introduced for the purpose of handling large datasets. A GAGP is constructed by pooling several Gaussian processes using identical hyperparameters but built from different subsets of the training data. The predictive capability of GAGPs is shown to be at least as good as state-of-the-art supervised learning methods. It is demonstrated on the unit-cell design of metamaterials through inverse optimization.Liu et al. present a method for the design for crashworthiness involving categorical multimaterial structures in their paper titled Design for Crashworthiness of Categorical Multimaterial Structures Using Cluster Analysis and Bayesian Optimization. Following a topology optimization, the dimensionality of the problem is reduced through clustering followed by a Bayesian optimization to assign a given material to a specific cluster. The approach is applied to the maximization of absorbed energy of an S-rail.Garriga et al. propose a framework to assist the optimization of aircraft systems at the early design stages. The approach in their paper titled A Machine Learning Enabled Multifidelity Platform for the Integrated Design of Aircraft Systems is based on the screening of designs using clustering followed by an identification of the best candidate on a Pareto front. The framework enables the use of models of various fidelities and is demonstrated on a primary flight control system and a landing gear.In Synthesizing Designs With Interpart Dependencies Using Hierarchical Generative Adversarial Networks, Chen and Fuge present a method for synthesizing hierarchical designs with inter-part dependencies using generative models learned from examples. The method constructs multiple generative models using generative adversarial networks (GANs) while satisfying the dependencies through part dependency graphs. The paper lays the foundation for extending the use of generative models from creative individual parts to more realistic engineering systems.The objective in Evolving a Psycho-Physical Distance Metric for Generative Design Exploration of Diverse Shapes by Khan et al. is to incorporate humans’ psychological perceptions about design into the design exploration process. A psycho-physical distance metric is proposed that enables the augmentation of CAD designs based on feedback from users. Results reveal that the proposed method generates more distinct variations of CAD designs compared with a baseline Euclidean distance method.Oh et al. in their paper titled Deep Generative Design: Integration of Topology Optimization and Generative Models present a design framework for creating diverse aesthetic designs that are optimized for engineering performance. The framework integrates topology optimization and generative adversarial networks (GANs) to generate large numbers of design options from limited previous design data. The approach is validated using a 2D wheel design problem.Deshpande and Purwar, in their paper Computational Creativity Via Assisted Variational Synthesis of Mechanisms Using Deep Generative Models, present an approach for variational synthesis of mechanisms and an End-to-End synthesis pipeline that accepts raw, high-level input from users and provides them with distinct concept solutions. The approach is based on learning the probability distribution of linkage parameters and their interdependence to perform tasks such as input conditioning, imputation, and variational synthesis. The approach is a step in the direction of enhancing users’ computational creativity for engineering design.Stump et al. in their paper, Spatial Grammar-Based Recurrent Neural Network for Design Form and Behavior Optimization, present a method for simultaneous optimization of form and behavior through a combination of physics-based models and ML techniques. Specifically, they use character-Recurrent Neural Networks to embody spatial grammars and reinforcement learning to optimize the behavior. The design of a modular multi-hull sailing craft is used as a demonstration problem.Suryadi and Kim utilize machine-learning algorithms for customer choice modeling in their paper titled A Data-Driven Methodology to Construct Customer Choice Sets Using Online Data and Customer Reviews. They present an approach that utilizes publicly available online data and customer reviews from e-commerce websites to construct customer choice sets in the absence of both an actual choice set and customer sociodemographic data. The approach consists of clustering (i) products based on their attributes and (ii) customers based on their reviews, and constructing the choice-sets based on a sampling probability scenario that relies on product and customer clusters. The approach generates choice models with higher predictive ability than randomly constructed choice sets.In their paper Extracting Customer Perceptions of Product Sustainability From Online Reviews, El Dehaibi et al. seek to extract perceived sustainable design features from online reviews. Annotators from Amazon’s Mechanical Turk are used to annotate product reviews and develop a natural language processing model that predicts the positive/negative sentiment of sustainable phrases. The results reveal that the model is more efficient at predicting positive sentiment pertaining to sustainable product features compared with negative sentiments.Raina et al. take a step toward transfer learning from human designers to computational agents in their paper titled Transferring Design Strategies From Human To Computer and Across Design Problems. They present an approach where design strategies are represented using a probabilistic model that provides a general mechanism to transfer strategies from human designers to computational design agents and to generate new designs. The approach is illustrated using a configuration design problem.The goal in Learning to Design From Humans: Imitating Human Designers Through Deep Learning by Raina et al. is to teach computational agents to generate designs without the need for explicit information about objective or performance metrics. A deep learning model is proposed that learns from historical human data and identifies the important regions of a design space. The results reveal that the machine learning agent learns to create designs that are comparable to human-generated ones, despite not having the same explicit feedback that humans do to guide them through the design exploration process.He et al. address the challenge of mining large numbers of design ideas generated from the crowd in their paper titled Mining and Representing the Concept Space of Existing Ideas for Directed Ideation. The authors use natural language processing to extract keywords as elementary concepts and represent the concepts in a way that they can be recombined to generate new ideas.In the paper titled A Data-Driven Approach to Product Usage Context Identification From Online Customer Reviews, Suryadi and Kim use machine learning and natural language processing to identify and cluster usage contexts from a large volume of customer reviews. The methodology also captures sentiments toward a particular usage context in a sentence. The methodology enables designers to effectively use online product reviews by focusing on several specific reviews regarding particular usage contexts and potentially to identify market opportunities for new products that excel in specific usage contexts.Sharpe et al. illuminate differences between Supervised Learning algorithms in terms of how and where different algorithms may apply to different Engineering Design applications. Their paper titled A Comparative Evaluation of Supervised Machine Learning Classification Techniques for Engineering Design Applications does this by comparing four common supervised learning approaches—Support Vector Machines, Random Forests, Gaussian Näive Bayes, and shallow depth Neural Networks—across six example problems that demonstrate different facets or challenges classifiers may face within the engineering design. The results from the work are multifaceted with different algorithms performing better or worse under different conditions and performance measures. However, this leads to the general notion of strong problem dependence for the classifier choice and highlights the importance of understanding appropriate benchmark problems within the engineering design that can shed light on such issues in the future.The availability of data enables not just designers but also design researchers. Rahman et al., in their paper A Computer-Aided Design Based Research Platform for Design Thinking Studies, present a research platform to support data-driven design-thinking and decision-making research. Through the use of fine-grained design action data and unsupervised clustering methods in conjunction with design process models, the authors show how the platform enables data-driven research studies on designers’ sequential decision-making behaviors.In Design Repository Effectiveness for 3D Convolutional Neural Networks: Application to Additive Manufacturing, Williams et al. address the question of whether or not a data repository is useful for training effective ML tools. The authors experimentally test the effects of changes in CAD datasets on the precision and generalizability of trained convolutional neural networks (CNNs) for additive manufacturing applications. The study sheds light on how standardization of design repositories can influence the performance of ML tools.Cunningham et al. study the construction of a performance surrogate based on 3D point cloud representations. In their paper titled An Investigation of Surrogate Models for Efficient Performance-Based Decoding of 3D Point Clouds, a radial basis function (RBF) surrogate is used to link performance and cloud representation mapped onto a latent vector. The proposed RBF-based approach was found to be more efficient and accurate than traditional neural network-based approaches.In the call for proposals for this special issue, the guest editors posed three primary questions: How to effectively use ML for new design applications that are not well-supported by existing ML practice or tools?How to leverage the unique aspects of engineering design in creating new ML approaches?How to share benchmark problems or datasets that can measure ML progress in design?Looking back at the papers collectively within this Special Issue helps shed light on the areas that are receiving significant within the design research and the areas where are opportunities for future design research a strong of using machine learning. techniques have used for data-driven techniques such as Gaussian process regression and neural networks have applications in learning complex between design and performance natural language processing used for mining customer and reinforcement learning is an part of control systems design. of these techniques are in the papers for the special issue Bostanabad et et Suryadi and Xiong et of the deep learning techniques such as convolutional neural networks (CNNs) and generative adversarial networks (GANs) are their way into engineering design research and and and et papers in this special issue address diverse applications including modeling, additive distributed energy systems, and topology optimization, synthesis, mechanism preference modeling, and learning from human of the areas where is potential for ML techniques, but are not well represented in this special issue, modeling human design of market systems, of products with humans or the design for use of data from product usage to using data from of the product or data, or across different aspects of the are opportunities for design challenges in engineering such as and and that from the products and ML can also be used to support engineering design for supporting studies and the generalizability of research terms of understanding and of machine learning methods, the papers in the special issue on the the of multifidelity or multiple data or with the common surrogate modeling and physics-based constraints into ML and ML-based models of human preferences and of the papers in the special issue present surrogate modeling this an active of research for the past two the approaches are using ML techniques. A of papers issues that ML models in multifidelity and et The availability of multifidelity models is in engineering design. In some designers’ are to construct multiple and fidelities of models or of a system as to with it at the appropriate This is not that ML systems are to In this of do multiple or important to to this special issue and is an active of of the approaches on existing with changes that constraints. of the is that many ML systems are not used to many of the problems that need for engineering design. in et al., how the between form and behavior the generated or in and Lewis the between Control and Design these of between and within Engineering and existing ML approaches do not need to for papers the of design or constraints in a constraints by or and hierarchical or et Chen and and constraints and This is important the of many ML systems relies on their to but specific of into the This is where engineering design researchers are well to are opportunities for research in engineering with ML models, such as physics-based models from or via more system models, and studies a problem and as a In the can ML be used with a between design and A is learning from multiple of a or learning among design with multiple data structures and of the papers to this special issue used supervised or reinforcement learning where data or available via used Learning clustering and dimensionality in natural and have from unsupervised learning of using deep learning or and this be a for future work in design. are opportunities for approaches to or for in uncertainty with or data. and calibration of ML models an papers on the of ML for optimization and specifically how to about generative models of two of approach to on the of the the approach Optimization as as by approaches that used or latent methods et al. and et or Optimization as such as that formulated the problem using reinforcement learning or inverse problem et are active and areas of in to for optimization that not in this special issue, such as direct inverse design papers how to best or behavior into an ML This is a are many opportunities for human information and strategies into ML models, and models of and computational or for behavior. This to approaches have used to models into the same be for of models, such as involving human or to ML models of human behavior designers or the design research on in both and and this is within the special However, are constraints that engineering design on of a notion that be to function and to that out as a of approaches to do not such aspects and this a that the Engineering Design can This also for are the fundamental of designers as conditions and does this datasets have ML research and a common for performance. collectively the the among have collectively of of They have in and This special issue set out to to new datasets for engineering design papers in the special issue propose datasets or platforms optimization problems et 2D and 3D shapes et Chen and and manufacturing et while not datasets do use platforms such as to data or these papers within the Special is a wide to be by future work that can useful datasets or for machine learning within engineering design. are some of the areas that papers in the Special Issue and have used ML but for good benchmark datasets datasets for human or including how designers or for more complex than optimization do not have the the model of for Engineering problems such as or that link or use multiple of design representations. datasets a CAD and a by the do not datasets have ML approaches in in the have in for and and datasets for approaches to specific in do not have a good of engineering problems and a given is to This is not unique to Engineering Design. How do or the and that it a realistic benchmark for Computer In some to be is the design on of or the of where and design datasets apply and how their results to datasets and understanding of their or be for future of how to such models within Engineering Design. how and of ML models within this on the of or ML models influence the design of a How does this Design are some of the and the papers in this special issue the of research opportunities in this The papers in the special issue many points from future researchers and may set guest the for their in the and feedback to the Special to and for their help with the for the process of of this special issue, and for the
Abstract Normative and cognitive-linguistic accounts of linguistic meaning are often portrayed and conceived as mutually exclusive alternatives. This dichotomy stems from an insufficient understanding of what the phenomenological accessibility of meaning and usage-basedness of language entail. Namely, the theoretical premises of Cognitive Linguistics actually presuppose socially grounded, normative linguistic meanings. The question remains, what kind of entities normative meanings are like. The present chapter makes a case for construal, linguistic perspective-taking usually analyzed as a conceptual phenomenon, as a normative facet of meaning. Analysis presented here suggests that construal emerges as an inherent property of linguistic expressions via conventionalization of intentionality. This analysis does not only expand the area of linguistic normativity but also points to the integral relation between linguistic norms and intentionality.
This chapter highlights the depth of the relationship between language and emotions. It aims to clarify some important distinctions and terms with respect to the linguistic encoding of emotions. Anthropological research has long confirmed that people in different human groups deal with emotional experience in different ways and different languages also offer very different means to talk about it. The chapter shows that emotional experience involves both some physiological and neurological mechanisms, which – for some of them – may be universal, as well as the cognitive organization of these mechanisms into experience, governed by culturally specific social and linguistic norms. A basic criterion that differentiates between descriptive and expressive linguistic resources is the semiotic status of the linguistic devices in question. Descriptive resources consist mostly of lexical resources, that is words, and some constructions.
Deep Universal Dependencies is a collection of treebanks derived semi-automatically from Universal Dependencies (http://hdl.handle.net/11234/1-2988). It contains additional deep-syntactic and semantic annotations. Version of Deep UD corresponds to the version of UD it is based on. Note however that some UD treebanks have been omitted from Deep UD.
People who experience trauma can develop enduring trauma-related symptoms. In daily life, post-trauma symptoms (e.g., elevated physiological arousal) can be triggered by affectively salient cues in the environment, especially by cues that act as trauma reminders. Trauma exposure is associated with enduring changes in two biological stress systems: the sympathetic nervous system (SNS) and the hypothalamic-pituitary-adrenal (HPA) axis. In women, activity in both systems is additionally modulated by fluctuations in levels of sex hormones (e.g., estradiol), which could influence physiological responses to trauma reminders. Additionally, previous work has linked the sex hormone estradiol with affect, suggesting that menstrual cycle might influence trauma-related symptoms or daily affect more broadly within the context of trauma exposure. However, we do not yet have a clear understanding of how estradiol influences affective experiences post-trauma. We used a multi-method approach to examine the influence of estradiol on daily affective experiences in a non-clinical, trauma-exposed sample of 40 naturally cycling premenopausal women. The first specific goal of this study was to test the hypothesis that low estradiol would be related to trauma symptoms, including an asymmetrical profile of SNS and HPA axis stress reactivity to a naturalistic trauma reminder. Lower estradiol was related to greater number and severity of PTSD symptoms, and participants in low versus high estradiol menstrual cycle phases showed higher SNS and reduced HPA axis reactivity to a trauma reminder. These results suggest that lower estradiol is associated with a less adaptive profile of stress system reactivity and increased PTSD symptom expression. The second specific goal of this study was to test the influence of menstrual cycle phase on daily affect in a subset of 30 participants. We assessed affective experience over the course of a 10-day ecological momentary assessment (EMA) period, which included the early follicular (low estradiol) and late follicular (high estradiol) phases. We selected these menstrual cycle phases to capture a portion of the cycle where estradiol increased, whereas progesterone remained low, allowing us to test the effects of estradiol without the confound of progesterone. Participants reported more frequent aversive affective experiences, defined as negatively valenced, high arousal states, including PTSD symptoms, during the early versus late follicular phase. During the early versus late follicular phase, participants also reported greater negative and positive affect and showed greater variability in affective ratings. These results suggest that lower estradiol menstrual cycle phases are characterized by more frequent aversive affective experiences, greater affective lability and increased PTSD symptom severity. Together, these results have potential implications for clinical assessment, as menstrual cycle phase at the time of assessment could influence diagnosis of PTSD or symptom severity. Additionally, clinicians working with women with PTSD might anticipate greater affective lability and increased symptom severity during low estradiol phases of the menstrual cycle.
Over the last few decades, corpora with comprehensive syntactic annotation, known as treebanks or parsed corpora, have been created in various formats for major languages of the world (e.g., As modes of accessing annotation have become more linguistically sophisticated, so these corpus resources have become more relevant for linguistics in general by providing sources of insight into factors that only become visible through analysis generalized over structures: phenomena in co-occurrence, frequency, constituency, embeddability, scope, agreement, dependency, etc. These insights are spurring new research and refinements in both corpus techniques and theoretical understanding. While much research has concentrated on challenges inherent in the creation as well as correction of annotated corpora (e.g., Examples include linking corpora to external resources like lexical databases, abstracting the contents sufficiently to be of use to non-experts, exploration of crosslinguistic patterns, etc. This special issue consists of five articles focused on applying parsed corpora research in three areas: (I) enrichment and
Modern health worries (MHW) represent individual differences in the perceived threat posed to health and well-being by aspects of modern life. Current evidence suggests that MHW are positively associated with trait negative emotionality, and given that trait negative emotionality is associated with state emotional reactivity to environmental stressors, it is reasonable to expect that persons with elevated MHW would show increased state emotional reactivity to MHW-related stimuli. Consequently, this study aimed to investigate the association of MHW with state emotional reactivity (i.e., valence and arousal) to MHW-related stimuli (i.e., images of air pollution). Combining these stimuli with other stimuli varying in valence and arousal allowed us to examine whether MHW are specifically associated with emotional reactivity induced by MHW-related stimuli. A total of 73 college students viewed 48 images encompassing eight different content areas, including a subset of MHW-related images (air pollution); each image was rated for valence and arousal. Participants also completed measures of MHW, trait negative emotionality (i.e., neuroticism), and demographics. After controlling for neuroticism and gender, results suggest that MHW only predicted valence rating for images of air pollution; conversely, MHW appears to be associated with arousal ratings in response to a variety of stimuli. Implications and limitations are discussed.
This paper presents the submission by the CMU-01 team to the SIGMORPHON 2019\ntask 2 of Morphological Analysis and Lemmatization in Context. This task\nrequires us to produce the lemma and morpho-syntactic description of each token\nin a sequence, for 107 treebanks. We approach this task with a hierarchical\nneural conditional random field (CRF) model which predicts each coarse-grained\nfeature (eg. POS, Case, etc.) independently. However, most treebanks are\nunder-resourced, thus making it challenging to train deep neural models for\nthem. Hence, we propose a multi-lingual transfer training regime where we\ntransfer from multiple related languages that share similar typology.\n
BACKGROUND: People experiencing mental illness require services that provide them with a sense of personal safety, a place where they can experience a reduction to their distress and assistance in managing their feelings. Interventions need to explore therapies that enhance feelings of personal safety and comfort for consumers and within a forensic mental health service, therapies and support that can assist in combating the antecedents to violent offending. The practice of Qigong is reported to have numerous health benefits; however, little has been reported regarding the possible benefits of Qigong for people experiencing severe mental illness and, more specifically, for people experiencing severe mental illness who have serious offending histories such as forensic consumers. This study explores the possibility of using Qigong to reduce personal frustrations that can lead to violence. OBJECTIVES: The object of this study was to explore whether Qigong is an effective intervention on positive affect traits for forensic mental health consumers, and whether other benefits are experienced. METHODS: An exploratory design using quantitative and qualitative approaches was used. Consumers participated in weekly Qigong groups delivered for a 10-wk period. Data were collected using an adapted version of the positive affect rating scale measuring the degree to which people experience different positive emotions. Qualitative measures were added to the scale to obtain a deeper understanding of the consumer experience, with 67 scales completed. CONCLUSIONS: Consumers in a forensic hospital responded positively to participating in Qigong groups. Strategies such as Qigong are interventions that mental health clinicians can use to promote positive feelings of personal relaxation, peacefulness, and safety. Qigong can promote positive affective traits for consumers in forensic hospitals. These positive affective traits can act as protective factors to inpatient aggression and violence. Forensic consumers report that Qigong is easy to learn and helpful for them in managing their frustrations. The findings from this study may add to the paucity of data discussing the use of Qigong with consumers as an effective relaxation intervention and possibly as an intervention in reducing negative affective states by promoting positive affective states, thereby reducing aggression and possible violence occurring within the forensic inpatient environment.
Lectometric approaches measure distances between language varieties (dialects, sociolects, registers etc.) by aggregating over observed differences in the realizations of a set of linguistic variables. In <em>lexical</em> lectometry, a variable consists of the alternative lexical expressions for one concept. In <em>corpus-based</em> lectometry, the observed realizations are culled from stratified corpora. Measuring semantically defined variables in corpora, and aggregating over them, poses specific methodological challenges that have been tackled in a number of studies (Heylen & Ruette 2013; Ruette et al. 2014; Ruette, Ehret & Szmrecsanyi 2016) with different statistical techniques, including Distributional Semantic Models. Yet so far, no general framework for corpus-based lexical lectometry has been formulated that systematically describes the issues and options in each step of the procedure so that it can be straightforwardly applied to new data and new languages, other than English (Ruette, Ehret & Szmrecsanyi 2016), Dutch (Geeraerts, Grondelaers & Speelman 1999) and Portuguese (Soares da Silva 2010). This paper can be characterized as a twofold extension of the previous studies. First, it aims to establish a general framework for lexical lectometry research that considers most if not all options for different steps. Second, we want to go beyond the Indo-European languages by extending the framework on a typologically unrelated language, i.e. Chinese varieties. For the general framework, we propose that a proper lexical lectometry research normally should involve the following steps: (1) compilation of a lectally stratified corpus; (2) sampling concepts as measuring points for lectometry; (3) identification of lexical expressions per concept; (4) disambiguation of lexical expressions in corpus data; (5) calculation of aggregated lexico-lectometric distances; (6) evaluation of measurement reliability and validity. For each step, we further provide possible options and caveats. For instance, step 2 and 3 can rely on existing concept-based lexical databases, like a synonym dictionary, or use corpus-driven keyword extraction and semantic vector space models. Step 4 can either make use of token-level distributional semantics models or rely on simpler n-gram language models. To assess the portability of the general framework, both in practical and linguistic-typological terms, we perform a lexical lectometric analysis for varieties of Chinese based on data from large-scale corpora of Mainland Chinese, Taiwan Chinese and Singapore Chinese.
Introduction: Hoarding behaviour is a common symptom seen in patients with Obsessive Compulsive Disorder (OCD). The phenomenology and prevalence of hoarding in OCD have been under studied in India and the phenomenon is less explored on routine clinical examination. Aim: To study the prevalence and phenomenology of hoarding as a symptom in patients with OCD and tried to elucidate some differences between OCD patients with and without hoarding symptoms. Materials and Methods: A total of 50 patients with OCD and 50 relatives of psychiatric patients were the subjects for the study. The OCD group was administered the Yale Brown Obsessive Compulsive Scale (YBOCS), the Hoarding Rating Scale and the Clutter Image Rating Scale. The 50 cases of OCD were further divided on the presence and absence of hoarding as a symptom into 2 groups and the scores on the scales used were statistically analysed using descriptive statistics like frequency and percentages, chi-square test and unpaired t-test. Results: The mean duration of illness was 8.01±5.17 years and the mean age of onset of the illness was 27.28±7.11 years for all patients with OCD. OCD patients with hoarding had a shorter total duration of illness than those without hoarding. Newspapers and scrap were hoarded the most with sentimental reasons along with importance of goods were cited as reasons for the behaviour. The two groups showed significant differences on compulsive sub scale of the YBOCS and no differences were noted in the other scales used. Conclusion: Patients having OCD with hoarding as one of the symptoms may differ from those not having hoarding. However larger studies across diverse groups are needed to corroborate these findings.
Affect fluctuates in a moment-to-moment fashion, and it reflects the continuous relationship between the individual and the environment. Despite substantial research, there remain important open questions regarding how the continuous stream of sensory input is dynamically represented in experienced affect. Here, approaching affect as a temporally dependent process, we show that momentary affect is shaped by a combination of changes in recent stimuli (i.e. visually presented images for the current studies) and previously experienced affect. We also found that this temporally dependent relationship is influenced by context uncertainty. Participants, in each trial, viewed sequentially presented images and subsequently reported their affective experience, which was modeled based on images’ normative affect ratings and participants’ previously reported affect. Study 1 showed that self-reported valence and arousal in a given trial is partly shaped by the affective impact of the given images and previously experienced affect. In Study 2, we manipulated context uncertainty by controlling occurrence probabilities for normatively pleasant and unpleasant images in separate trials. Increasing context uncertainty (i.e. random occurrence of pleasant and unpleasant images) is associated with increased negative affect when the overall effect context is controlled. In addition, the relative contribution of the most recent image to momentary affect increased with increasing context uncertainty. Taken together, these findings provide clear behavioral evidence that affective experience fluctuates in a temporally dependent and continuous fashion based on recent changes in input variables and previous internal state, and that these fluctuations are sensitive to the affective context and its certainty.
The Plain Meaning Rule is often assailed on the grounds that it is unprincipled—that it substitutes for careful analysis an interpreter’s ad hoc and impressionistic intuition about the meaning of legal texts. But what if judges and lawyers had the means to test their intuitions about plain meaning systematically? Then initial linguistic impressions about the meaning of a legal text might be viewed as hypotheses to be tested, rather than determinative criteria upon which to base important decisions. There exists very little legal scholarship on corpus linguistics—the study of language function and use through large, electronic linguistic databases called corpora—and the role that corpus methods might play in legal interpretation. This omission becomes more and more striking as scholars and jurists (and even the United States Supreme Court) have found themselves persuaded by corpus-based arguments. This Article argues that the plain or ordinary meaning of a given term in a given context is an empirical matter that may be quantified through corpus-based methods. These methods, when applied to questions of legal ambiguity, present significant advantages over existing empirical approaches to plain meaning and over the prevailing intuition-based interpretive approach of many courts. Because large, sophisticated linguistic corpora are widely available and easy to use, and because corpus methods offer a more principled and systematic alternative to the impressionistic interpretation of legal texts, corpus linguistics may one day revolutionize the process of legal interpretation.
Nonsuicidal self-injury (NSSI; e.g., cutting or burning the skin without suicidal intent) is a dangerous and increasingly prevalent health-risk behavior. Despite advances in NSSI research over the past decade, many aspects of NSSI remain poorly understood. In particular, there are few strong predictors of NSSI, it is unclear how positive attitudes toward NSSI develop, and there are no empirically supported treatments for NSSI. In the present study, I addressed these topics with a multi-method, experimental, and longitudinal approach. For Aim 1 of the study, I examined baseline differences between NSSI (n = 58) and control (n = 86) adult participants on NSSI-themed versions of five measures that cover different aspects of attitudes: the implicit association test (IAT); the affect misattribution procedure (AMP); explicit affective ratings; startle eyeblink reactivity; and startle postauricular reactivity. Compared to the control group, the NSSI group displayed significantly more positive attitudes on all five measures. Moreover, AMP scores and explicit ratings prospectively predicted self-cutting frequency over the ensuing six months. For Aim 2, I employed pain offset relief conditioning in an attempt to induce more positive implicit attitudes toward NSSI in the control group. This conditioning significantly diminished startle eyeblink reactivity in the context of NSSI images, but did not significantly affect any other measures. For Aim 3, I tested the ability of aversive conditioning in the NSSI group to reverse positive implicit attitudes toward NSSI and to reduce NSSI behaviors over the subsequent six months. Aversive conditioning normalized startle eyeblink and postauricular reactivity, but did not significantly affect any other measures. Results also provided preliminary support for the hypothesis that aversive conditioning prospectively reduces self-cutting. In conjunction with my other recent studies (Franklin et al., 2010; 2011, 2012, 2013), these findings have prompted a new theoretical framework called the Benefits and Barriers model of NSSI.
The practice of assessing brand management in construction in Ukraine is in a passive stage, but due to the entry into the Ukrainian market of foreign companies for which regular evaluation of their brand — the need for survival in a competitive environment, Ukrainian companies are beginning to pay more attention to the creation and formation of their competitive trading because of its high business image / rating. A modern toolkit based on appropriate approaches is used to form organizational and economic foundations. The article analyzes modern scientific approaches to assessing the economic potential of an enterprise, identifies the main trends and factors that affect the assessment of construction enterprises. The article is devoted to the study of theoretical and methodological assessments of branding of construction enterprises. The role and importance of innovation in ensuring the efficient operation of modern enterprises is emphasized. It is found that construction, especially innovative, is of great social importance and has a significant economic effect. The importance of determining the potential of innovative development of construction in general and construction enterprises in particular is substantiated. The purpose of this article is to systematically investigate the interpretation of the potential of innovative development of construction enterprises. The theoretical basis of the research is the scientific works of foreign and domestic scientists on the problems of identifying the essence of innovative development potential. A systematic study of the general characteristics of the construction company brand was conducted and the priority directions for choosing the development of the economic potential of the enterprises were determined. The necessity to understand the potential of the enterprise in the unity of all its elements, which are subject to the achievement of the overall goals of the enterprise, is substantiated. The weight of the component of the brand in the potential of the construction industry enterprises is substantiated. Some aspects of the development of theoretical and methodological approaches to the estimation of the intellectual capital of construction enterprises are formulated. Existing theoretical and methodological approaches to the brand assessment of construction enterprises are analyzed.
This study focuses on a comprehensive analysis and manual re-annotation of the Turkish IMST-UD Treebank, which was automatically converted from the IMST Treebank (Sulubacak et al., 2016b). In accordance with the Universal Dependencies' guidelines and the necessities of Turkish grammar, the existing treebank was revised. The current study presents the revisions that were made alongside the motivations behind the major changes. Moreover, it reports the parsing results of a transition-based dependency parser and a graph-based dependency parser obtained over the previous and updated versions of the treebank. In light of these results, we have observed that the re-annotation of the Turkish IMST-UD treebank improves performance with regards to dependency parsing.
The relationship between words in a sentence often tell us more about the underlying semantic content of a document than its actual words individually. Natural language understanding has seen an increasing effort in the formation of techniques that try to produce non-trivial features, in the last few years, especially after robust word embeddings models became prominent, when they proved themselves able to capture and represent semantic relationships from massive amounts of data. These new dense vector representations indeed leverage the baseline in natural language processing, but they still fall short in dealing with intrinsic issues in linguistics, such as polysemy and homonymy. Systems that make use of natural language at its core, can be affected by a weak semantic representation of human language, resulting in inaccurate outcomes based on poor decisions. In this subject, word sense disambiguation and lexical chains have been exploring alternatives to alleviate several problems in linguistics, such as semantic representation, definitions, differentiation, polysemy, and homonymy. However, little effort is seen in combining recent advances in token embeddings (e.g. words, documents) with word sense disambiguation and lexical chains. To collaborate in building a bridge between these areas, this work proposes a collection of algorithms to extract semantic features from large corpora as its main contributions, named MSSA, MSSA-D, MSSA-NR, FLLC II, and FXLC II. The MSSA techniques focus on disambiguating and annotating each word by its specific sense, considering the semantic effects of its context. The lexical chains group derive the semantic relations between consecutive words in a document in a dynamic and pre-defined manner. These original techniques' target is to uncover the implicit semantic links between words using their lexical structure, incorporating multi-sense embeddings, word sense disambiguation, lexical chains, and lexical databases. A few natural language problems are selected to validate the contributions of this work, in which our techniques outperform state-of-the-art systems. All the proposed algorithms can be used separately as independent components or combined in one single system to improve the semantic representation of words, sentences, and documents. Additionally, they can also work in a recurrent form, refining even more their results.
يعد القرآن الكريم من مصادر المعرفة، وقد تولدت منه فروع واسعة؛ إذ نُزل القرآن الكريم باللغة العربية، ولا يوجد خيار آخر لإتقان المعرفة الواردة فيه إلا من خلال تعلم اللغة العربية. تهدف هذه الدراسة إلى بيان مفهوم المدونة العربية القرآنية ومكوناتها، والكشف عن علاقة تعلم اللغة العربية بالقرآن الكريم، وبيان كيفية تعليم وتعلم القواعد العربية الأساسية عبر المدونة العربية القرآنية، وستتبع الدراسة المنهج الوصفي والتحليلي. إن وجود العلاقة بين اللغة العربية والقرآن الكريم، يدفع الطلبة المتخصصين في اللغة العربية أن يربطوا اللغة العربية بالقرآن؛ لذلك نرى أن المدونة العربية القرآنية تساعدهم على فهم القواعد القرآنية بطريقة مثيرة للاهتمام. في نظرة شاملة يمكن أن نستنتج أن المدونة العربية القرآنية هي واحدة من أهم الأدوات الحسابية التي تم إنتاجها في خدمة اللغة العربية؛ حيث توفر للمتعلمين ما يحتاجون إليه في مجال اللغة واللغويات والدراسات الحاسوبية، كما تمهد الطريق للباحثين لدراسة الهياكل المورفولوجية والنحوية من خلال دراسات الحوسبة العميقة للقرآن.
 الكلمات المفتاحية: المصرف القرآني، نموذج حاسوبي، المدونة العربية القرآنية، المعجم القرآني.
 Abstract 
 The Holy Quran is a source of knowledge and it has generated wide branches of knowledge. The Holy Quran was revealed in Arabic. Hence, there is no other option to master its knowledge except by learning the Arabic language. This study aims at explaining the concept of the Arabic Quranic Corpus and its components, revealing the relationship between learning the Arabic language and the Holy Quran, and showing how to teach and learn basic Arabic grammar through the Quranic Arabic Corpus. The study will follow the descriptive and analytical approach. The existence of the relationship between the Arabic language and the Holy Quran prompts Arabic learners to associate Arabic with the Qur'an. Therefore, we see that the Quranic Arabic Corpus helps them to understand Quranic rules in an interesting way. In a comprehensive view, we can conclude that the Arabic Quranic Corpus is one of the most important web-based medium produced to serve the Arabic language. It provides learners with what they need in the field of language, linguistics and computer studies, and paves the way for researchers to study morphological and grammatical structures through technology with detail description of grammars.
 Keywords: Quranic Treebank, Computational Model, Arabic Quranic Corpus, Qur’anic Dictionary.
Paper is dedicated to the testing of the concept of literacy, based on the questionnaire, carried out in the school year of 2018/2019 among the students of two secondary vocational schools in Vrsac, Belgrade and Grammar School in Vrsac (200 respondents). The primary hypothesis of the research was that detection and detailed study of high school students conceptosphere on literacy identify the fields to improve the teaching of Serbian as a mother tongue in secondary schools and the aim of work that, based on the collected and then processed data in analytical, cognitive and descriptive method, is to (a) isolate the dominant concepts of (non)literacy, (b) look at the tendencies of spreading and shaping the notion of literacy induced by the needs of modern life, and also that, in order to improve linguistic culture in all domains and all educational levels -(c) point to the possibility of improving the teaching of the Serbian language as a mother tongue. According to results of the survey secondary school students experience literacy in the 21 st century as a complex concept; from the one who is literate expecting linguistic knowledge, what are the basic, traditionally accepted parameters, and recognize illiteracy as the lack of ability to apply knowledge in the field of language. They also demonstrated that it is necessary to improve the efficiency of teaching approaches designed to improve functional literacy in a variety of communicative situations; increase the number of hours and exercises in the field of spelling, or nurture and acquire more comprehensive and knowledge in use and skills of different forms of literacy needed for managing in 21 st century; more attention should be paid to including relevant language handbooks in teaching; more explicit, on frequent and more familiar examples to students, point to the advantages of knowing and respecting the linguistic norm, paving the way for a better linguistic culture and enrichment of the mother tongue.
Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Speech-to-Text (STT) and Text-to-Speech (TTS) systems, the question remains how to best leverage state-of-the-art language models (which capture rich semantic features, but are trained on only written text) on inputs with ASR errors. In this paper, we present Telephonetic, a data augmentation framework that helps robustify language model features to ASR corrupted inputs. To capture phonetic alterations, we employ a character-level language model trained using probabilistic masking. Phonetic augmentations are generated in two stages: a TTS encoder (Tacotron 2, WaveGlow) and a STT decoder (DeepSpeech). Similarly, semantic perturbations are produced by sampling from nearby words in an embedding space, which is computed using the BERT language model. Words are selected for augmentation according to a hierarchical grammar sampling strategy. Telephonetic is evaluated on the Penn Treebank (PTB) corpus, and demonstrates its effectiveness as a bootstrapping technique for transferring neural language models to the speech domain. Notably, our language model achieves a test perplexity of 37.49 on PTB, which to our knowledge is state-of-the-art among models trained only on PTB.
Lexical Markup Framework (LMF) or ISO 24613 [1] is a de jure standard that\nprovides a framework for modelling and encoding lexical information in\nretrodigitised print dictionaries and NLP lexical databases. An in-depth review\nis currently underway within the standardisation subcommittee,\nISO-TC37/SC4/WG4, to find a more modular, flexible and durable follow up to the\noriginal LMF standard published in 2008. In this paper we will present some of\nthe major improvements which have so far been implemented in the new version of\nLMF.\n
The intelligent information processing of the standard Zhuang language spoken mainly in Southern China is presently in its infancy, and lacks a well-defined language corpus and automatic part-of-speech tagging methods. Therefore, this study proposes an adversarial part-of-speech tagging method based on reinforcement learning, which solves the problems associated with a lack of a language corpus, time-consuming laborious manual marking, and the low performance of machine marking. Firstly, we construct a markup dictionary based on the grammatical characteristics of standard Zhuang and the Penn Chinese Treebank. Secondly, a dependency syntax analysis is applied for constructing the semantic information feature vectors of sentences, and long short-term memory is adopted as the policy network architecture to enhance available information using recurrent memory, and a conditional random field is employed as the discriminant network to perform label inference with global normalization. Finally, we use reinforcement learning as the model framework, target parts of speech as the feedback of the environment, and then obtain the optimal policy through adversarial learning. The results show that the combination of reinforcement learning and adversarial network alleviates the dependence of the model on the training corpus to some extent, and can quickly and effectively expand the scale of the annotation dictionary for the Zhuang language, thereby obtaining better labeling results.
According to World Intellectual Property Organization (2017) report, over 3 million patents exist in the patent database, but only certain numbers have commercial potential. Generally, to assess the commercial potential of patent, it consumes time and requires various expertise. Currently, several models have been developed to address this matter, which to assess using questionnaire tool for portfolios by human. So that occurs bias any limitation exists that our research will address by artificial intelligence. Hence, this research applies a Natural Language Programming to assess for commercial potential of patent, consisting of five steps - (i) Morphological analysis based on the Lexical database, (ii) Syntactic analysis of sentence to check syntax sentence patterns, (iii) Sematic analysis to interpret the meaning of words derived from the previous step, (iv) Discourse integration from context of domain together with the main sentence providing more accurate sentence analysis, and (v) Pragmatic analysis to ensure the correct meaning of interpretation. Then, the obtained data is used to determine criterion factors and formulate the model for assessing commercial potential of patent using Natural Language programming. This finding should deliver an alternative effective patent assessment system, which addresses some current deficiency in patent's assessment for commercial potential.
Antiaddictive social advertising is a special speech genre of modern communication with specific features determined by the target setting and the chosen strategy. Advertising can be considered as a special functional style, within which separate genres are distinguished, first of all commercial advertising and social advertising. Social advertising, which refers to ethical categories, has become especially popular because society is faced with such problems, the solution of which depends on mass behavior. Advertising has become an integral part of the daily life of a person. Since advertising is mass-replicated in the media, it can enter the consciousness of the addressee, even against his/her will and desire, no wonder advertising is sometimes defined as the “fifth power” after media, whose power is considered the “fourth power”. Therefore, it is so important for advertising to follow the moral principles of society, it is so important to comply with modern ethical and linguistic norms. If commercial advertising is widespread, the antiaddictive one is little known. Many specialists in advertising argue for the need of the strategy of “shock” advertising, a remarkable feature of which is hyperbolization, used to cause fear. In antiaddictive advertising there is often a morphological imperative in the meaning of categorical motivation. It is justified as such advertizing has to be as much as possible appellate, has to be understood unambiguously. Anti-drug social advertising should show positive, motivation to a healthy lifestyle.
Perception of emotional valence and emotional memory performance vary across the menstrual cycle. However, the consequences of altered ovarian hormone levels due to the intake of hormonal contraceptives on these emotional and cognitive processes remain to be established. In the present study, which included 2169 healthy young females, we show that hormonal contraceptives (HC) users rated emotional pictures as more emotional than HC-non-users and outperformed non-users in terms of better memory recall of emotional pictures. The observed association between HC-status and memory performance was partially mediated by the perception of emotional picture valence, indicating that increased valence ratings of emotional pictures in HC-users led to their better emotional memory performance. These findings extend the knowledge about the relation of HC-intake with emotional valence perception and emotional memory performance. Further, the findings might stimulate further research investigating the interrelation of enhanced memory for emotional events and the increased risk for anxiety-related psychiatric disorders in women.
Recent research on discourse relations has found that they are cued not only by discourse markers (DMs) but also by other textual signals and that signaling information is indicative of genres. While several corpora exist with discourse relation signaling information such as the Penn Discourse Treebank (PDTB, Prasad et al. 2008) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada 2018), they both annotate the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. 1993), which is limited to the news domain. Thus, this paper adapts the signal identification and anchoring scheme (Liu and Zeldes, 2019) to three more genres, examines the distribution of signaling devices across relations and genres, and provides a taxonomy of indicative signals found in this dataset.
espanolResumen en castellano: Esta tesis trata de la morfologia verbal del ingles antiguo para identificar y lematizar los verbos debiles de esta lengua en un corpus al que se accede a traves de una base de datos lexica. La lematizacion es una de las tareas mas importantes a la hora de construir un diccionario. Sin embargo, es una de las tareas pendientes en el campo de la linguistica historica debido a que no existen corpora exhaustivos y lematizados de esta lengua. El enfoque de esta tesis doctoral esta en la lematizacion de las tres clases de verbos debiles del ingles antiguo, aunque las areas de la Lexicografia y la Linguistica de Corpus son tambien relevantes para esta investigacion. Las fuentes principales de esta investigacion son las formas flexivas que estan atestiguadas en el Dictionary of Old English Corpus (DOEC) y que estan disponibles en el lematizador Norna, las fuentes lexicograficas que existen publicadas sobre esta lengua, principalmente el Dictionary of Old English (DOE), y otras fuentes textuales como el York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) y una indexacion de fuentes secundarias del ingles antiguo. El objetivo principal supone la identificacion de las flexiones de los verbos debiles y de su lematizacion con uno de los lemas propuestos en las listas de referencia. Conseguir este objetivo implica manejar las fuentes disponibles en ingles antiguo para poder lematizar y validar los resultados del analisis y el diseno de un metodo que combine busquedas automaticas en la base de datos lexica Nerthus y la revision manual de los resultados. La metodologia incluye cuatro pasos sucesivos con diversas tareas en cada paso. El primero de estos pasos tiene como objetivo la lematizacion de las formas canonicas de los verbos debiles lanzando cadenas de busquedas especificas para cada clase de verbos debiles en el lematizador Norna, donde esta disponible un indice de tipos del DOEC, la fuente de informacion mas fiable de la que se dispone en ingles antiguo. Despues, los resultados se validan con el DOE y se anaden las formas no-canonicas de los verbos debiles entre las letras A y H. El tercer paso tiene como objetivo identificar las formas no-canonicas de las terminaciones flexivas y de las vocales de los radicales que aparecen con mas frecuencia en los verbos debiles para generar patrones de lematizacion. La busqueda de estos patrones y de la lista de prefijos no-canonicos que esta disponible en Norna culmina en la lematizacion de las formas flexivas no transparentes de los verbos debiles. La validacion de los resultados de las letras I a la Y supone el ultimo paso de la metodologia, donde se comparan los datos obtenidos con el analisis sintactico del YCOE y con los datos que se obtienen de una base de datos de indexacion de las fuentes secundarias del ingles antiguo. Los problemas que surgen a lo largo del proceso de lematizacion tienen que ver principalmente con las peculiaridades del ingles antiguo y las limitaciones de la lematizacion de tipos que esta investigacion sigue. La discusion de los resultados del analisis concluye esta tesis. Las principales aportaciones de esta tesis son las listas de lemas y sus formas flexivas, especialmente las de los verbos entre las letras I y la Y ya que no estan disponibles todavia, y el metodo que se ha disenado para identificar estas formas, incluyendo los patrones de lematizacion generados para lematizar las formas con terminaciones no comunes y vocales no canonicas en el radical. EnglishThis thesis deals with the verbal morphology of the Old English language in order to identify and lemmatise weak verbs in a corpus accessed through a lexical database. Lemmatisation is a pending task in the field of historical linguistics given the lack of comprehensive and lemmatised corpora in this language. The focus of this doctoral dissertation is on the lemmatisation of the three classes of weak verbs, although the linguistic fields of Lexicography and Corpus Linguistics are also relevant to this research. The main aim involves the identification of the canonical and non-canonical realisations of the Old English weak verbs and their lemmatisation with a lemma from a reference list of weak verbs. Achieving this goal involves, firstly, the use of the available sources of the Old English language in order to lemmatise and validate the results and, secondly, the design of a semi-automatic research methodology that combines automatic searches in the lexical database Nerthus and the manual revision of the results in order to achieve this task. The sources for this investigation are the inflectional forms that are attested in the Dictionary of Old English Corpus (DOEC) which are available in the lemmatiser Norna, the lexicographical sources published on the Old English language, mainly the Dictionary of Old English (DOE), and other textual sources such as the York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) and an index of secondary sources of Old English. The methodology comprises four successive steps and several tasks within each step. The first step aims at the lemmatisation of the transparent forms of weak verbs with the search of specific query strings for each subclass of weak verbs in the lemmatiser Norna, where an index type of the DOEC, the most reliable source of information regarding the Old English language, is available. Then, the second step validates the results with the DOE and adds to the analysis the non-canonical attestations for the weak verbs from the letter A-H. Thirdly, the identification of the most recurrent non-canonical inflectional endings and stem vowels attested in weak verbs gives rise to lemmatisation patterns. The search of these sets of correspondences and the list of non-canonical prefixes that is available in Norna results in the lemmatisation of the non-canonical inflections of weak verbs. The validation of the results from the letter I-Y concludes the research methodology with the syntactic parsing provided by the YCOE and the data retrieved from the index of secondary sources of Old English Freya. The issues that arise throughout the lemmatisation process mainly concern the idiosyncrasy of the Old English language writing system and the limitations of the lemmatisation by type that this investigation follows. The quantitative and qualitative discussion of the results of the analysis concludes this thesis. The main contributions of this thesis are the lists of weak lemmas and their lemmatised inflectional forms, specially those of the verbs I-Y which are not available yet and the designed research methodology to identify these forms, including the sets of lemmatisation patterns of the non-canonical inflectional endings and stem vowels of weak verbs.
Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider the correlation between parameters and therefore can lead to unstable and slow convergence in training. Second-order optimization methods provide a possible solution to this issue. However these methods are either computationally heavy or do not have competitive performance. In this paper, a novel optimization method - stochastic natural gradient based on minimum variance assumption (SNGM) is proposed for training RNNLMs. It allows the natural gradient method to operate at a comparable training efficiency to the SGD method. By modifying the gradient according to the local curvature of the KL-divergence between current and updated probabilistic distributions, the proposed SNGM approach is shown to outperform both the SGD and limited memory BFGS methods across three tasks: Penn Treebank, Switchboard conversational speech recognition and AMI meeting room transcription in terms of both perplexity and word error rate.
Previous studies have shown that hoarding behavior usually starts at a subclinical level in early adolescence and gradually worsens; however, a limited number of studies have examined the prevalence of hoarding behavior and its association with developmental disorders in young adults. The aims of this study were to estimate the prevalence of hoarding behavior and to identify correlations between hoarding behavior and developmental disorder traits in university students. The study participants included 801 university students (616 men, 185 women) who completed questionnaires (ASRS: Adult ADHD Self-Report Scale version 1.1, AQ16: Autism-Spectrum Quotient with 16 items, and CIR: Clutter Image Rating). Among 801 participants, 27 (3.4%) exceeded the CIR cut-off score. Moreover, the participants with hoarding behavior had a significantly higher percentage of ADHD traits compared to participants without hoarding behavior (HB(+) vs HB(−), 40.7% vs 21.7%). In addition, 7.4% of HB(+) participants had autism spectrum disorder (ASD) traits, compared to 4.1% of HB(−) participants. A correlation analysis revealed that the CIR composite score had a stronger correlation with the ASRS inattentive score than with the hyperactivity/impulsivity score (CIR composite vs ASRS IA, r = 0.283; CIR composite vs ASRS H/I, r = 0.147). The results showed a high prevalence of ADHD traits in the university students with hoarding behavior. Moreover, we found that the hoarding behavior was more strongly correlated with inattentive symptoms rather than with hyperactivity/impulsivity symptoms. Our results support the concept of a common pathophysiology behind hoarding behavior and ADHD in young adults.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
The emergence of China as a global economic power in the 21st Century has brought about surging needs for cross-lingual and cross-cultural mediation, typically performed by translators. Advances in Artificial Intelligence and Language Engineering have been bolstered by Machine learning and suitable Big Data cultivation. They have helped to meet some of the translator's needs, though the technical specialists have not kept pace with the practical and expanding requirements in language mediation. One major technical and linguistic hurdle involves words outside the vocabulary of the translator or the lexical database he/she consults, especially Multi-Word Expressions (Compound Words) in technical subjects. A further problem lies in the multiplicity of renditions of a term in the target language. This paper discusses a proactive approach following the successful extraction and application of sizable bilingual Multi-Word Expressions (Compound Words) for language mediation in technical subjects, which do not fall within the expertise of typical translators, who have inadequate appreciation of the range of new technical tools available to help him/her. Our approach draws on the personal reflections of translators and teachers of translation and is based on the prior R&D efforts relating to 300,000 comparable Chinese-English patents. The subsequent protocol we have developed aims to be proactive in meeting four identified practical challenges in technical translation (e.g. patents). It has broader economic implication in the Age of Big Data (Tsou et al, 2015) and Trade War, as the workload, if not, the challenges, increasingly cannot be met by currently available front-line translators. We shall demonstrate how new tools can be harnessed to spearhead the application of language technology not only in language mediation but also in the “teaching” and “learning” of translation. It shows how a better appreciation of their needs may enhance the contributions of the technical specialists, and thus enhance the resultant synergetic benefits. © 2019 Incoma Ltd. All rights reserved.
Counterconditioning (CC) is a form of retroactive interference that inhibits expression of learned behavior. But similar to extinction, CC can be a fairly weak and impermanent form of interference, and the original behavior is prone to relapse. Research on CC is limited, especially in humans, but prior studies suggest it is more effective than extinction at modifying some behaviors (e.g., preference or valence ratings) than others (e.g., physiological arousal). Here, we used a within-subjects design to compare the effects of aversive-to-appetitive CC versus standard extinction on two separate tests of long-term memory in human adults: implicit physiological arousal and explicit episodic memory. Participants underwent Pavlovian fear conditioning to two semantic categories (animals, tools) paired with an electric shock. Conditioned stimuli (i.e., category exemplars) from one category were then extinguished, while stimuli from the other category were paired with a positive outcome. Participants returned 24-h later for a test of skin conductance responses (SCR) to the conditioned exemplars, as well as a surprise recognition memory test for stimuli encoded the previous day. Results showed reduced SCRs at a test for unique stimuli from a category that had undergone CC, relative to stimuli from a category that had undergone standard extinction. Additionally, participants selectively remembered more stimuli encoded during CC than extinction. These results provide new evidence that aversive-to-appetitive CC, as compared to extinction, strengthens memory for items directly associated with a positive outcome, which may provide stronger retrieval competition against a fear memory at test to help diminish fear relapse.
We introduce a language-agnostic evolutionary technique for automatically extracting chunks from dependency treebanks. We evaluate these chunks on a number of morphosyntactic tasks, namely POS 1 tagging, morphological feature tagging, and dependency parsing. We test the utility of these chunks in a host of different ways. We first learn chunking as one task in a shared multitask framework together with POS and morphological feature tagging. The predictions from this network are then used as input to augment sequence-labelling dependency parsing. Finally, we investigate the impact chunks have on dependency parsing in a multi-task framework. Our results from these analyses show that these chunks improve performance at different levels of syntactic abstraction on English UD treebanks and a small, diverse subset of non-English UD treebanks.
It is widely accepted that the human cognitive system organizes perceptual input into complex hierarchical descriptions which can be represented by tree structures. Tree structures have been used to describe linguistic, musical and visual perception. In this paper, we will investigate whether there exists an underlying model that governs perceptual organization in general. Our key idea is that the cognitive system strives for the simplest structure (the “simplicity principle”), but in doing so it is biased by the likelihood of previous experiences (the “likelihood principle”). We will present a model which combines these two principles by balancing the notion of most likely tree with the notion of shortest derivation. Experiments with linguistic and musical benchmarks (Penn Treebank and Essen Folksong Collection) show that such a combination outperforms models that are based on either simplicity or likelihood alone.
This chapter engages with research in lingua franca scenarios, defining this term and arguing for the value of adopting a linguistic ethnographic approach in this area. The development of the field of English as a Lingua Franca (ELF) is outlined. Historical antecedents for lingua franca studies in interactional sociolinguistics are identified, including work on intercultural miscommunications, the study of English as an international language and the identification of pragmatic strategies in ELF scenarios such as the let-it-pass procedure. The chapter considers critical debates, including the need to challenge the norms of standard language pedagogies, and the importance of maintaining a critical perspective on the current global dominance of English. It provides a review of current areas of focus in this area, including studies of English as lingua franca in higher education in an increasingly internationalised university system, and in workplaces. The range of methods drawn on in ethnographic studies of lingua franca scenarios is described, and implications for practice are identified in relation to the teaching of English and language policy and planning. Directions for future research identified include the emergence of social and linguistic norms in interaction, and developing the focus on lingua francas other than English.
Abstract: Temperatures above 20° Celsius have shown to adversely impact human behavior, leading to increased aggression and violence. Climate change will contribute to both the magnitude and severity of this pattern as temperatures continue their rise. Contributions to this field of research have only recently begun to analyze online behavior and language as a proxy for hedonic state, or well-being. From a development perspective this study is relevant since the poor tend to live in some of the warmest regions on earth, and would thus be disproportionately impacted by increased temperatures. We use several sources of data; U.S. based daily statewide temperature data from 2016 through 2017, as well as localized viewer chat data from a live video streaming website. We will sort chatting comments looking for key words (i.e. hate speech, swearing, etc.), and with the use of a word rating system we then assess the overall mood of the chatters contingent on high temperature readings on the precise day of the communications. After controlling for spatiotemporal fixed effects, we find strong evidence that hedonic state decreases above 20°c.
Discourse relation classification has proven to be a hard task, with rather low performance on several corpora that notably differ on the relation set they use. We propose to decompose the task into smaller, mostly binary tasks corresponding to various primitive concepts encoded into the discourse relation definitions. More precisely, we translate the discourse relations into a set of values for attributes based on distinctions used in the mappings between discourse frameworks proposed by This arguably allows for a more robust representation of discourse relations, and enables us to address usually ignored aspects of discourse relation prediction, namely multiple labels and underspecified annotations. We study experimentally which of the conceptual primitives are harder to learn from the Penn Discourse Treebank English corpus, and propose a correspondence to predict the original labels, with preliminary empirical comparisons with a direct model.
In the paper the analysis of the perspective of the technology of intelligent monitoring systems is conducted, as intellectualization is the main direction of development of modern technologies, and the property of intellectuality should be inherent in all the latest information management systems. Various strategies for intellectualization of monitoring are aimed at implementing intellectual information support for decision-makers using monitoring tools. Such support can be realized by building fuzzy linguistic databases/knowledge together with fuzzy inference subsystems, and information for decision making can be displayed on the automated workplace of the decision maker.
Implicit causal relation recognition aims to identify the causal relation between a pair of arguments. It is a challenging task due to the lack of conjunctions and the shortage of labeled data. In order to improve the identification performance, we come up with an approach to expand the training dataset. On the basis of the hypothesis that there inherently exists causal relations in WHY-type Question-Answer (QA) pairs, we utilize WHY-type QA pairs for the training set expansion. In practice, we first collect WHY-type QA pairs from the Knowledge Bases (KBs) of the reading comprehension tasks, and then convert them into narrative argument pairs by Question-Statement Conversion (QSC). In order to alleviate redundancy, we use active learning (AL) to select informative samples from the synthetic argument pairs. The sampled synthetic argument pairs are added to the Penn Discourse Treebank (PDTB), and the expanded PDTB is used to retrain the neural network-based classifiers. Experiments show that our method yields a performance gain of 2.42% F 1-score when AL is used, and 1.61% without using.
The aim of the present study was to examine whether offspring at high and low familial risk for depression differ in the immediate and more lasting behavioural and physiological effects of hedonically-based mood repair. Participants (9- to 22-year olds) included never-depressed offspring at high familial depression risk (high-risk, n = 64), offspring with similar familial background and personal depression histories (high-risk/DEP, n = 25), and never-depressed offspring at low familial risk (controls, n = 62). Offspring provided affect ratings at baseline, after sad mood induction, immediately following hedonically-based mood repair, and at subsequent, post-repair epochs. Physiological reactivity, indexed via respiratory sinus arrhythmia (RSA), was assessed during the protocol. Following mood induction and mood repair, high- and low-risk (control) offspring reported comparable changes in levels of sadness and RSA. However, sadness increased among high-risk offspring following the post-repair epoch, whereas low-risk offspring maintained mood repair benefits. High-risk/DEP offspring also reported higher levels of sadness following the post-repair epoch than did low-risk offspring. Change in RSA did not differ across the three offspring groups. Self-ratings confirm that one source of difficulty associated with depression risk is diminished ability to maintain hedonically-based mood repair gains, which were not apparent at the physiological level.
We present a strategy to automate the extraction of semantic relations from texts. Both machine learning and rule-based techniques are investigated and the impact of different linguistic knowledge is analyzed for the various approaches. To implement the extraction system RExtractor, several natural language processing tools have been improved: from sentence splitting and tokenization modules to dependency syntax parsers. Furthermore, we created the Czech Legal Text Treebank with several layers of linguistic annotation, which is used to train and test each stage of the proposed system. As a result of the performed work, new Semantic Web resources and tools are available for automatic processing of texts.
A workshop on open resources for the original languages of the Bible in Copenhagen in March 2018 was the start of a new Copenhagen Alliance for Open Biblical Resources. The point of departure for the workshop was the need for programs and applications like Paratext and Bible Online Learner to have access to high-quality and reliable open data in order to assist Bible translators, teachers and students of Biblical Hebrew and New Testament Greek. The publication of contributions presents papers on methods for annotation, resources tracing patristic quotations and data for detached constructions in Biblical Hebrew. Reports cover tasks and data for Bible translation and research, treebanks, and applications like STEPBible and Bible Online Learner.