Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Do individual sounds carry meaning? The relationship between sound and meaning in human languages is typicallyassumed to be arbitrary, though recent research provides evidence for the existence of both iconicity and systematicitybetween word forms and their meaning. However, this research has not asked whether individual sounds in a languagecovary in systematic ways with aspects of meaning. In two analyses, we find evidence for more systematicity betweenthe initial phones of words and those words concreteness ratings than one would expect in a truly arbitrary lexicon. Thissuggests that initial phones may act as cues to aspects of word meaning, and raises questions about whether languagelearners detect and exploit these cues.
Highly regularized LSTMs achieve impressive results on several benchmark datasets in language modeling. We propose a new regularization method based on decoding the last token in the context using the predicted distribution of the next token. This biases the model towards retaining more contextual information, in turn improving its ability to predict the next token. With negligible overhead in the number of parameters and training time, our Past Decode Regularization (PDR) method improves perplexity on the Penn Treebank dataset by up to 1.8 points and by up to 2.3 points on the WikiText-2 dataset, over strong regularized baselines using a single softmax. With a mixture-of-softmax model, we show gains of up to 1.0 perplexity points on these datasets. In addition, our method achieves 1.169 bits-per-character on the Penn Treebank Character dataset for character level language modeling. Each of these results constitute improvements over models without PDR in their respective settings.
Throughout many studies which focus on brain laterality as a key component to the outcome of an experiment, or to a participants’ reaction to a stimulus, it can be noted that the different areas of the brain involved in a task response must work together to produce a viable outcome (i.e. lateralized brain processes). While there are laterality components in relation to the cognitive processes of handed and footed responses, it is still largely unknown how the different areas of the brain interpret the emotional stimuli to then affect these outcomes. The purpose of our study is to determine how emotional context affects a simple cognitive task that includes handed and footed responses, and if any observed differences can be traced back to the different systems at work within the brain. Subjects will be tested over two days for handed and footed responses in a cognitive Simon Task. Subjects will be tested with and without emotional context (i.e. a background image of a specific valence and arousal rating), and any resulting differences between non-emotional context and emotional context reaction times will be compared.
In sequence learning tasks such as language modelling, Recurrent Neural\nNetworks must learn relationships between input features separated by time.\nState of the art models such as LSTM and Transformer are trained by\nbackpropagation of losses into prior hidden states and inputs held in memory.\nThis allows gradients to flow from present to past and effectively learn with\nperfect hindsight, but at a significant memory cost. In this paper we show that\nit is possible to train high performance recurrent networks using information\nthat is local in time, and thereby achieve a significantly reduced memory\nfootprint. We describe a predictive autoencoder called bRSM featuring recurrent\nconnections, sparse activations, and a boosting rule for improved cell\nutilization. The architecture demonstrates near optimal performance on a\nnon-deterministic (stochastic) partially-observable sequence learning task\nconsisting of high-Markov-order sequences of MNIST digits. We find that this\nmodel learns these sequences faster and more completely than an LSTM, and offer\nseveral possible explanations why the LSTM architecture might struggle with the\npartially observable sequence structure in this task. We also apply our model\nto a next word prediction task on the Penn Treebank (PTB) dataset. We show that\na 'flattened' RSM network, when paired with a modern semantic word embedding\nand the addition of boosting, achieves 103.5 PPL (a 20-point improvement over\nthe best N-gram models), beating ordinary RNNs trained with BPTT and\napproaching the scores of early LSTM implementations. This work provides\nencouraging evidence that strong results on challenging tasks such as language\nmodelling may be possible using less memory intensive, biologically-plausible\ntraining regimes.\n
The emergence of deep learning as a commanding technique for learning heterogeneous layers of feature representations have consequently substituted traditional machine learning algorithms which are generally poor in analyzing compound sentences. Additionally, convolutional and recurrent neural networks have auspiciously yielded state-of-the-art results in sentiment classification and Natural Language Processing (NLP). In this paper, a deep sentiment representation model through the combination of multiple Convolutional Neural Networks (CNN) kernels with Long Short-Term Memory (LSTM) is proposed for sentiment classification. Our model gains word vector representation using pre-trained Global Vectors for Word Representation (GloVe) embeddings, thereafter used as input to the CNN layer which extracts higher local text representations. Finally, Bidirectional LSTM (biLSTM) generates sentiment classification of sentence representation based on context dependent features. Our combined approach of CNN and biLSTM was experimented using the Stanford Large Movie Review Dataset (IMDB) and Stanford Sentiment Treebank Dataset (SSTB) for binary classification. The evaluation achieves outstanding results in outperforming several existing approaches with 90.4% accuracy on the Stanford Sentiment Treebank dataset and 94.8% accuracy on the Stanford Large Movie Review dataset. These results are achieved with a drastic reduction of model parameters and without a pooling layer in the CNN architecture, helping to retain local and structural information in comparison to other existing deep neural network frameworks.
This paper presents the submission by the CMU-01 team to the SIGMORPHON 2019\ntask 2 of Morphological Analysis and Lemmatization in Context. This task\nrequires us to produce the lemma and morpho-syntactic description of each token\nin a sequence, for 107 treebanks. We approach this task with a hierarchical\nneural conditional random field (CRF) model which predicts each coarse-grained\nfeature (eg. POS, Case, etc.) independently. However, most treebanks are\nunder-resourced, thus making it challenging to train deep neural models for\nthem. Hence, we propose a multi-lingual transfer training regime where we\ntransfer from multiple related languages that share similar typology.\n
BACKGROUND: People experiencing mental illness require services that provide them with a sense of personal safety, a place where they can experience a reduction to their distress and assistance in managing their feelings. Interventions need to explore therapies that enhance feelings of personal safety and comfort for consumers and within a forensic mental health service, therapies and support that can assist in combating the antecedents to violent offending. The practice of Qigong is reported to have numerous health benefits; however, little has been reported regarding the possible benefits of Qigong for people experiencing severe mental illness and, more specifically, for people experiencing severe mental illness who have serious offending histories such as forensic consumers. This study explores the possibility of using Qigong to reduce personal frustrations that can lead to violence. OBJECTIVES: The object of this study was to explore whether Qigong is an effective intervention on positive affect traits for forensic mental health consumers, and whether other benefits are experienced. METHODS: An exploratory design using quantitative and qualitative approaches was used. Consumers participated in weekly Qigong groups delivered for a 10-wk period. Data were collected using an adapted version of the positive affect rating scale measuring the degree to which people experience different positive emotions. Qualitative measures were added to the scale to obtain a deeper understanding of the consumer experience, with 67 scales completed. CONCLUSIONS: Consumers in a forensic hospital responded positively to participating in Qigong groups. Strategies such as Qigong are interventions that mental health clinicians can use to promote positive feelings of personal relaxation, peacefulness, and safety. Qigong can promote positive affective traits for consumers in forensic hospitals. These positive affective traits can act as protective factors to inpatient aggression and violence. Forensic consumers report that Qigong is easy to learn and helpful for them in managing their frustrations. The findings from this study may add to the paucity of data discussing the use of Qigong with consumers as an effective relaxation intervention and possibly as an intervention in reducing negative affective states by promoting positive affective states, thereby reducing aggression and possible violence occurring within the forensic inpatient environment.
Lectometric approaches measure distances between language varieties (dialects, sociolects, registers etc.) by aggregating over observed differences in the realizations of a set of linguistic variables. In <em>lexical</em> lectometry, a variable consists of the alternative lexical expressions for one concept. In <em>corpus-based</em> lectometry, the observed realizations are culled from stratified corpora. Measuring semantically defined variables in corpora, and aggregating over them, poses specific methodological challenges that have been tackled in a number of studies (Heylen & Ruette 2013; Ruette et al. 2014; Ruette, Ehret & Szmrecsanyi 2016) with different statistical techniques, including Distributional Semantic Models. Yet so far, no general framework for corpus-based lexical lectometry has been formulated that systematically describes the issues and options in each step of the procedure so that it can be straightforwardly applied to new data and new languages, other than English (Ruette, Ehret & Szmrecsanyi 2016), Dutch (Geeraerts, Grondelaers & Speelman 1999) and Portuguese (Soares da Silva 2010). This paper can be characterized as a twofold extension of the previous studies. First, it aims to establish a general framework for lexical lectometry research that considers most if not all options for different steps. Second, we want to go beyond the Indo-European languages by extending the framework on a typologically unrelated language, i.e. Chinese varieties. For the general framework, we propose that a proper lexical lectometry research normally should involve the following steps: (1) compilation of a lectally stratified corpus; (2) sampling concepts as measuring points for lectometry; (3) identification of lexical expressions per concept; (4) disambiguation of lexical expressions in corpus data; (5) calculation of aggregated lexico-lectometric distances; (6) evaluation of measurement reliability and validity. For each step, we further provide possible options and caveats. For instance, step 2 and 3 can rely on existing concept-based lexical databases, like a synonym dictionary, or use corpus-driven keyword extraction and semantic vector space models. Step 4 can either make use of token-level distributional semantics models or rely on simpler n-gram language models. To assess the portability of the general framework, both in practical and linguistic-typological terms, we perform a lexical lectometric analysis for varieties of Chinese based on data from large-scale corpora of Mainland Chinese, Taiwan Chinese and Singapore Chinese.
Dependency grammar induction is the task of learning dependency syntax without annotated training data. Traditional graph-based models with global inference achieve state-ofthe-art results on this task but they require O(n3) run time. Transition-based models enable faster inference with O(n) time complexity, but their performance still lags behind. In this work, we propose a neural transition-based parser for dependency grammar induction, whose inference procedure utilizes rich neural features with O(n) time complexity. We train the parser with an integration of variational inference, posterior regularization and variance reduction techniques. The resulting framework outperforms previous unsupervised transition-based dependency parsers and achieves performance comparable to graph-based models, both on the English Penn Treebank and on the Universal Dependency Treebank. In an empirical comparison, we show that our approach substantially increases parsing speed over graphbased models.
Introduction: Hoarding behaviour is a common symptom seen in patients with Obsessive Compulsive Disorder (OCD). The phenomenology and prevalence of hoarding in OCD have been under studied in India and the phenomenon is less explored on routine clinical examination. Aim: To study the prevalence and phenomenology of hoarding as a symptom in patients with OCD and tried to elucidate some differences between OCD patients with and without hoarding symptoms. Materials and Methods: A total of 50 patients with OCD and 50 relatives of psychiatric patients were the subjects for the study. The OCD group was administered the Yale Brown Obsessive Compulsive Scale (YBOCS), the Hoarding Rating Scale and the Clutter Image Rating Scale. The 50 cases of OCD were further divided on the presence and absence of hoarding as a symptom into 2 groups and the scores on the scales used were statistically analysed using descriptive statistics like frequency and percentages, chi-square test and unpaired t-test. Results: The mean duration of illness was 8.01±5.17 years and the mean age of onset of the illness was 27.28±7.11 years for all patients with OCD. OCD patients with hoarding had a shorter total duration of illness than those without hoarding. Newspapers and scrap were hoarded the most with sentimental reasons along with importance of goods were cited as reasons for the behaviour. The two groups showed significant differences on compulsive sub scale of the YBOCS and no differences were noted in the other scales used. Conclusion: Patients having OCD with hoarding as one of the symptoms may differ from those not having hoarding. However larger studies across diverse groups are needed to corroborate these findings.
Affect fluctuates in a moment-to-moment fashion, and it reflects the continuous relationship between the individual and the environment. Despite substantial research, there remain important open questions regarding how the continuous stream of sensory input is dynamically represented in experienced affect. Here, approaching affect as a temporally dependent process, we show that momentary affect is shaped by a combination of changes in recent stimuli (i.e. visually presented images for the current studies) and previously experienced affect. We also found that this temporally dependent relationship is influenced by context uncertainty. Participants, in each trial, viewed sequentially presented images and subsequently reported their affective experience, which was modeled based on images’ normative affect ratings and participants’ previously reported affect. Study 1 showed that self-reported valence and arousal in a given trial is partly shaped by the affective impact of the given images and previously experienced affect. In Study 2, we manipulated context uncertainty by controlling occurrence probabilities for normatively pleasant and unpleasant images in separate trials. Increasing context uncertainty (i.e. random occurrence of pleasant and unpleasant images) is associated with increased negative affect when the overall effect context is controlled. In addition, the relative contribution of the most recent image to momentary affect increased with increasing context uncertainty. Taken together, these findings provide clear behavioral evidence that affective experience fluctuates in a temporally dependent and continuous fashion based on recent changes in input variables and previous internal state, and that these fluctuations are sensitive to the affective context and its certainty.
The Plain Meaning Rule is often assailed on the grounds that it is unprincipled—that it substitutes for careful analysis an interpreter’s ad hoc and impressionistic intuition about the meaning of legal texts. But what if judges and lawyers had the means to test their intuitions about plain meaning systematically? Then initial linguistic impressions about the meaning of a legal text might be viewed as hypotheses to be tested, rather than determinative criteria upon which to base important decisions. There exists very little legal scholarship on corpus linguistics—the study of language function and use through large, electronic linguistic databases called corpora—and the role that corpus methods might play in legal interpretation. This omission becomes more and more striking as scholars and jurists (and even the United States Supreme Court) have found themselves persuaded by corpus-based arguments. This Article argues that the plain or ordinary meaning of a given term in a given context is an empirical matter that may be quantified through corpus-based methods. These methods, when applied to questions of legal ambiguity, present significant advantages over existing empirical approaches to plain meaning and over the prevailing intuition-based interpretive approach of many courts. Because large, sophisticated linguistic corpora are widely available and easy to use, and because corpus methods offer a more principled and systematic alternative to the impressionistic interpretation of legal texts, corpus linguistics may one day revolutionize the process of legal interpretation.
Nonsuicidal self-injury (NSSI; e.g., cutting or burning the skin without suicidal intent) is a dangerous and increasingly prevalent health-risk behavior. Despite advances in NSSI research over the past decade, many aspects of NSSI remain poorly understood. In particular, there are few strong predictors of NSSI, it is unclear how positive attitudes toward NSSI develop, and there are no empirically supported treatments for NSSI. In the present study, I addressed these topics with a multi-method, experimental, and longitudinal approach. For Aim 1 of the study, I examined baseline differences between NSSI (n = 58) and control (n = 86) adult participants on NSSI-themed versions of five measures that cover different aspects of attitudes: the implicit association test (IAT); the affect misattribution procedure (AMP); explicit affective ratings; startle eyeblink reactivity; and startle postauricular reactivity. Compared to the control group, the NSSI group displayed significantly more positive attitudes on all five measures. Moreover, AMP scores and explicit ratings prospectively predicted self-cutting frequency over the ensuing six months. For Aim 2, I employed pain offset relief conditioning in an attempt to induce more positive implicit attitudes toward NSSI in the control group. This conditioning significantly diminished startle eyeblink reactivity in the context of NSSI images, but did not significantly affect any other measures. For Aim 3, I tested the ability of aversive conditioning in the NSSI group to reverse positive implicit attitudes toward NSSI and to reduce NSSI behaviors over the subsequent six months. Aversive conditioning normalized startle eyeblink and postauricular reactivity, but did not significantly affect any other measures. Results also provided preliminary support for the hypothesis that aversive conditioning prospectively reduces self-cutting. In conjunction with my other recent studies (Franklin et al., 2010; 2011, 2012, 2013), these findings have prompted a new theoretical framework called the Benefits and Barriers model of NSSI.
This article focuses on the empirical experience and conclusions, resulting from the creation of language research and acquisition tools for Livonian – one of the smallest languages in Europe. A cluster was created for Livonian containing three interconnected databases, each with distinct types of data – lexical, morphological, and a corpus. The lexical database contains the lemmas and their data, the morphological database stores morphological forms, while all textual material, including the dictionary examples, is in the corpus. When indexing the corpus, every word refers to a lemma in the lexical database and its morphological information (new lemmas are added prior to indexation), ensuring consistency of the language data, and from each database the full data set of the other databases can be accessed. The function of each cluster is to extract the maximum amount of information from limited data sources. While technologies designed for languages with a large number of speakers focus on using quantitative methods and automation to extract qualitative information from a large and constantly expanding amount of linguistic data, the main function of technologies designed for small languages is to extract the same type of information from a limited and largely static data set. This article also examines a string of problems faced when working with a small amount of resources (inadequate language data, insufficient personnel, lack of rules for automating processes, etc.) and methods for resolving these problems in the case of Livonian.
The practice of assessing brand management in construction in Ukraine is in a passive stage, but due to the entry into the Ukrainian market of foreign companies for which regular evaluation of their brand — the need for survival in a competitive environment, Ukrainian companies are beginning to pay more attention to the creation and formation of their competitive trading because of its high business image / rating. A modern toolkit based on appropriate approaches is used to form organizational and economic foundations. The article analyzes modern scientific approaches to assessing the economic potential of an enterprise, identifies the main trends and factors that affect the assessment of construction enterprises. The article is devoted to the study of theoretical and methodological assessments of branding of construction enterprises. The role and importance of innovation in ensuring the efficient operation of modern enterprises is emphasized. It is found that construction, especially innovative, is of great social importance and has a significant economic effect. The importance of determining the potential of innovative development of construction in general and construction enterprises in particular is substantiated. The purpose of this article is to systematically investigate the interpretation of the potential of innovative development of construction enterprises. The theoretical basis of the research is the scientific works of foreign and domestic scientists on the problems of identifying the essence of innovative development potential. A systematic study of the general characteristics of the construction company brand was conducted and the priority directions for choosing the development of the economic potential of the enterprises were determined. The necessity to understand the potential of the enterprise in the unity of all its elements, which are subject to the achievement of the overall goals of the enterprise, is substantiated. The weight of the component of the brand in the potential of the construction industry enterprises is substantiated. Some aspects of the development of theoretical and methodological approaches to the estimation of the intellectual capital of construction enterprises are formulated. Existing theoretical and methodological approaches to the brand assessment of construction enterprises are analyzed.
The relationship between words in a sentence often tell us more about the underlying semantic content of a document than its actual words individually. Natural language understanding has seen an increasing effort in the formation of techniques that try to produce non-trivial features, in the last few years, especially after robust word embeddings models became prominent, when they proved themselves able to capture and represent semantic relationships from massive amounts of data. These new dense vector representations indeed leverage the baseline in natural language processing, but they still fall short in dealing with intrinsic issues in linguistics, such as polysemy and homonymy. Systems that make use of natural language at its core, can be affected by a weak semantic representation of human language, resulting in inaccurate outcomes based on poor decisions. In this subject, word sense disambiguation and lexical chains have been exploring alternatives to alleviate several problems in linguistics, such as semantic representation, definitions, differentiation, polysemy, and homonymy. However, little effort is seen in combining recent advances in token embeddings (e.g. words, documents) with word sense disambiguation and lexical chains. To collaborate in building a bridge between these areas, this work proposes a collection of algorithms to extract semantic features from large corpora as its main contributions, named MSSA, MSSA-D, MSSA-NR, FLLC II, and FXLC II. The MSSA techniques focus on disambiguating and annotating each word by its specific sense, considering the semantic effects of its context. The lexical chains group derive the semantic relations between consecutive words in a document in a dynamic and pre-defined manner. These original techniques' target is to uncover the implicit semantic links between words using their lexical structure, incorporating multi-sense embeddings, word sense disambiguation, lexical chains, and lexical databases. A few natural language problems are selected to validate the contributions of this work, in which our techniques outperform state-of-the-art systems. All the proposed algorithms can be used separately as independent components or combined in one single system to improve the semantic representation of words, sentences, and documents. Additionally, they can also work in a recurrent form, refining even more their results.
يعد القرآن الكريم من مصادر المعرفة، وقد تولدت منه فروع واسعة؛ إذ نُزل القرآن الكريم باللغة العربية، ولا يوجد خيار آخر لإتقان المعرفة الواردة فيه إلا من خلال تعلم اللغة العربية. تهدف هذه الدراسة إلى بيان مفهوم المدونة العربية القرآنية ومكوناتها، والكشف عن علاقة تعلم اللغة العربية بالقرآن الكريم، وبيان كيفية تعليم وتعلم القواعد العربية الأساسية عبر المدونة العربية القرآنية، وستتبع الدراسة المنهج الوصفي والتحليلي. إن وجود العلاقة بين اللغة العربية والقرآن الكريم، يدفع الطلبة المتخصصين في اللغة العربية أن يربطوا اللغة العربية بالقرآن؛ لذلك نرى أن المدونة العربية القرآنية تساعدهم على فهم القواعد القرآنية بطريقة مثيرة للاهتمام. في نظرة شاملة يمكن أن نستنتج أن المدونة العربية القرآنية هي واحدة من أهم الأدوات الحسابية التي تم إنتاجها في خدمة اللغة العربية؛ حيث توفر للمتعلمين ما يحتاجون إليه في مجال اللغة واللغويات والدراسات الحاسوبية، كما تمهد الطريق للباحثين لدراسة الهياكل المورفولوجية والنحوية من خلال دراسات الحوسبة العميقة للقرآن.
 الكلمات المفتاحية: المصرف القرآني، نموذج حاسوبي، المدونة العربية القرآنية، المعجم القرآني.
 Abstract 
 The Holy Quran is a source of knowledge and it has generated wide branches of knowledge. The Holy Quran was revealed in Arabic. Hence, there is no other option to master its knowledge except by learning the Arabic language. This study aims at explaining the concept of the Arabic Quranic Corpus and its components, revealing the relationship between learning the Arabic language and the Holy Quran, and showing how to teach and learn basic Arabic grammar through the Quranic Arabic Corpus. The study will follow the descriptive and analytical approach. The existence of the relationship between the Arabic language and the Holy Quran prompts Arabic learners to associate Arabic with the Qur'an. Therefore, we see that the Quranic Arabic Corpus helps them to understand Quranic rules in an interesting way. In a comprehensive view, we can conclude that the Arabic Quranic Corpus is one of the most important web-based medium produced to serve the Arabic language. It provides learners with what they need in the field of language, linguistics and computer studies, and paves the way for researchers to study morphological and grammatical structures through technology with detail description of grammars.
 Keywords: Quranic Treebank, Computational Model, Arabic Quranic Corpus, Qur’anic Dictionary.
Paper is dedicated to the testing of the concept of literacy, based on the questionnaire, carried out in the school year of 2018/2019 among the students of two secondary vocational schools in Vrsac, Belgrade and Grammar School in Vrsac (200 respondents). The primary hypothesis of the research was that detection and detailed study of high school students conceptosphere on literacy identify the fields to improve the teaching of Serbian as a mother tongue in secondary schools and the aim of work that, based on the collected and then processed data in analytical, cognitive and descriptive method, is to (a) isolate the dominant concepts of (non)literacy, (b) look at the tendencies of spreading and shaping the notion of literacy induced by the needs of modern life, and also that, in order to improve linguistic culture in all domains and all educational levels -(c) point to the possibility of improving the teaching of the Serbian language as a mother tongue. According to results of the survey secondary school students experience literacy in the 21 st century as a complex concept; from the one who is literate expecting linguistic knowledge, what are the basic, traditionally accepted parameters, and recognize illiteracy as the lack of ability to apply knowledge in the field of language. They also demonstrated that it is necessary to improve the efficiency of teaching approaches designed to improve functional literacy in a variety of communicative situations; increase the number of hours and exercises in the field of spelling, or nurture and acquire more comprehensive and knowledge in use and skills of different forms of literacy needed for managing in 21 st century; more attention should be paid to including relevant language handbooks in teaching; more explicit, on frequent and more familiar examples to students, point to the advantages of knowing and respecting the linguistic norm, paving the way for a better linguistic culture and enrichment of the mother tongue.
Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Speech-to-Text (STT) and Text-to-Speech (TTS) systems, the question remains how to best leverage state-of-the-art language models (which capture rich semantic features, but are trained on only written text) on inputs with ASR errors. In this paper, we present Telephonetic, a data augmentation framework that helps robustify language model features to ASR corrupted inputs. To capture phonetic alterations, we employ a character-level language model trained using probabilistic masking. Phonetic augmentations are generated in two stages: a TTS encoder (Tacotron 2, WaveGlow) and a STT decoder (DeepSpeech). Similarly, semantic perturbations are produced by sampling from nearby words in an embedding space, which is computed using the BERT language model. Words are selected for augmentation according to a hierarchical grammar sampling strategy. Telephonetic is evaluated on the Penn Treebank (PTB) corpus, and demonstrates its effectiveness as a bootstrapping technique for transferring neural language models to the speech domain. Notably, our language model achieves a test perplexity of 37.49 on PTB, which to our knowledge is state-of-the-art among models trained only on PTB.
Lexical Markup Framework (LMF) or ISO 24613 [1] is a de jure standard that\nprovides a framework for modelling and encoding lexical information in\nretrodigitised print dictionaries and NLP lexical databases. An in-depth review\nis currently underway within the standardisation subcommittee,\nISO-TC37/SC4/WG4, to find a more modular, flexible and durable follow up to the\noriginal LMF standard published in 2008. In this paper we will present some of\nthe major improvements which have so far been implemented in the new version of\nLMF.\n
The intelligent information processing of the standard Zhuang language spoken mainly in Southern China is presently in its infancy, and lacks a well-defined language corpus and automatic part-of-speech tagging methods. Therefore, this study proposes an adversarial part-of-speech tagging method based on reinforcement learning, which solves the problems associated with a lack of a language corpus, time-consuming laborious manual marking, and the low performance of machine marking. Firstly, we construct a markup dictionary based on the grammatical characteristics of standard Zhuang and the Penn Chinese Treebank. Secondly, a dependency syntax analysis is applied for constructing the semantic information feature vectors of sentences, and long short-term memory is adopted as the policy network architecture to enhance available information using recurrent memory, and a conditional random field is employed as the discriminant network to perform label inference with global normalization. Finally, we use reinforcement learning as the model framework, target parts of speech as the feedback of the environment, and then obtain the optimal policy through adversarial learning. The results show that the combination of reinforcement learning and adversarial network alleviates the dependence of the model on the training corpus to some extent, and can quickly and effectively expand the scale of the annotation dictionary for the Zhuang language, thereby obtaining better labeling results.
According to World Intellectual Property Organization (2017) report, over 3 million patents exist in the patent database, but only certain numbers have commercial potential. Generally, to assess the commercial potential of patent, it consumes time and requires various expertise. Currently, several models have been developed to address this matter, which to assess using questionnaire tool for portfolios by human. So that occurs bias any limitation exists that our research will address by artificial intelligence. Hence, this research applies a Natural Language Programming to assess for commercial potential of patent, consisting of five steps - (i) Morphological analysis based on the Lexical database, (ii) Syntactic analysis of sentence to check syntax sentence patterns, (iii) Sematic analysis to interpret the meaning of words derived from the previous step, (iv) Discourse integration from context of domain together with the main sentence providing more accurate sentence analysis, and (v) Pragmatic analysis to ensure the correct meaning of interpretation. Then, the obtained data is used to determine criterion factors and formulate the model for assessing commercial potential of patent using Natural Language programming. This finding should deliver an alternative effective patent assessment system, which addresses some current deficiency in patent's assessment for commercial potential.
Recent research on discourse relations has found that they are cued not only by discourse markers (DMs) but also by other textual signals and that signaling information is indicative of genres. While several corpora exist with discourse relation signaling information such as the Penn Discourse Treebank (PDTB, Prasad et al. 2008) and the Rhetorical Structure Theory Signalling Corpus (RST-SC, Das and Taboada 2018), they both annotate the Wall Street Journal (WSJ) section of the Penn Treebank (PTB, Marcus et al. 1993), which is limited to the news domain. Thus, this paper adapts the signal identification and anchoring scheme (Liu and Zeldes, 2019) to three more genres, examines the distribution of signaling devices across relations and genres, and provides a taxonomy of indicative signals found in this dataset.
espanolResumen en castellano: Esta tesis trata de la morfologia verbal del ingles antiguo para identificar y lematizar los verbos debiles de esta lengua en un corpus al que se accede a traves de una base de datos lexica. La lematizacion es una de las tareas mas importantes a la hora de construir un diccionario. Sin embargo, es una de las tareas pendientes en el campo de la linguistica historica debido a que no existen corpora exhaustivos y lematizados de esta lengua. El enfoque de esta tesis doctoral esta en la lematizacion de las tres clases de verbos debiles del ingles antiguo, aunque las areas de la Lexicografia y la Linguistica de Corpus son tambien relevantes para esta investigacion. Las fuentes principales de esta investigacion son las formas flexivas que estan atestiguadas en el Dictionary of Old English Corpus (DOEC) y que estan disponibles en el lematizador Norna, las fuentes lexicograficas que existen publicadas sobre esta lengua, principalmente el Dictionary of Old English (DOE), y otras fuentes textuales como el York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) y una indexacion de fuentes secundarias del ingles antiguo. El objetivo principal supone la identificacion de las flexiones de los verbos debiles y de su lematizacion con uno de los lemas propuestos en las listas de referencia. Conseguir este objetivo implica manejar las fuentes disponibles en ingles antiguo para poder lematizar y validar los resultados del analisis y el diseno de un metodo que combine busquedas automaticas en la base de datos lexica Nerthus y la revision manual de los resultados. La metodologia incluye cuatro pasos sucesivos con diversas tareas en cada paso. El primero de estos pasos tiene como objetivo la lematizacion de las formas canonicas de los verbos debiles lanzando cadenas de busquedas especificas para cada clase de verbos debiles en el lematizador Norna, donde esta disponible un indice de tipos del DOEC, la fuente de informacion mas fiable de la que se dispone en ingles antiguo. Despues, los resultados se validan con el DOE y se anaden las formas no-canonicas de los verbos debiles entre las letras A y H. El tercer paso tiene como objetivo identificar las formas no-canonicas de las terminaciones flexivas y de las vocales de los radicales que aparecen con mas frecuencia en los verbos debiles para generar patrones de lematizacion. La busqueda de estos patrones y de la lista de prefijos no-canonicos que esta disponible en Norna culmina en la lematizacion de las formas flexivas no transparentes de los verbos debiles. La validacion de los resultados de las letras I a la Y supone el ultimo paso de la metodologia, donde se comparan los datos obtenidos con el analisis sintactico del YCOE y con los datos que se obtienen de una base de datos de indexacion de las fuentes secundarias del ingles antiguo. Los problemas que surgen a lo largo del proceso de lematizacion tienen que ver principalmente con las peculiaridades del ingles antiguo y las limitaciones de la lematizacion de tipos que esta investigacion sigue. La discusion de los resultados del analisis concluye esta tesis. Las principales aportaciones de esta tesis son las listas de lemas y sus formas flexivas, especialmente las de los verbos entre las letras I y la Y ya que no estan disponibles todavia, y el metodo que se ha disenado para identificar estas formas, incluyendo los patrones de lematizacion generados para lematizar las formas con terminaciones no comunes y vocales no canonicas en el radical. EnglishThis thesis deals with the verbal morphology of the Old English language in order to identify and lemmatise weak verbs in a corpus accessed through a lexical database. Lemmatisation is a pending task in the field of historical linguistics given the lack of comprehensive and lemmatised corpora in this language. The focus of this doctoral dissertation is on the lemmatisation of the three classes of weak verbs, although the linguistic fields of Lexicography and Corpus Linguistics are also relevant to this research. The main aim involves the identification of the canonical and non-canonical realisations of the Old English weak verbs and their lemmatisation with a lemma from a reference list of weak verbs. Achieving this goal involves, firstly, the use of the available sources of the Old English language in order to lemmatise and validate the results and, secondly, the design of a semi-automatic research methodology that combines automatic searches in the lexical database Nerthus and the manual revision of the results in order to achieve this task. The sources for this investigation are the inflectional forms that are attested in the Dictionary of Old English Corpus (DOEC) which are available in the lemmatiser Norna, the lexicographical sources published on the Old English language, mainly the Dictionary of Old English (DOE), and other textual sources such as the York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) and an index of secondary sources of Old English. The methodology comprises four successive steps and several tasks within each step. The first step aims at the lemmatisation of the transparent forms of weak verbs with the search of specific query strings for each subclass of weak verbs in the lemmatiser Norna, where an index type of the DOEC, the most reliable source of information regarding the Old English language, is available. Then, the second step validates the results with the DOE and adds to the analysis the non-canonical attestations for the weak verbs from the letter A-H. Thirdly, the identification of the most recurrent non-canonical inflectional endings and stem vowels attested in weak verbs gives rise to lemmatisation patterns. The search of these sets of correspondences and the list of non-canonical prefixes that is available in Norna results in the lemmatisation of the non-canonical inflections of weak verbs. The validation of the results from the letter I-Y concludes the research methodology with the syntactic parsing provided by the YCOE and the data retrieved from the index of secondary sources of Old English Freya. The issues that arise throughout the lemmatisation process mainly concern the idiosyncrasy of the Old English language writing system and the limitations of the lemmatisation by type that this investigation follows. The quantitative and qualitative discussion of the results of the analysis concludes this thesis. The main contributions of this thesis are the lists of weak lemmas and their lemmatised inflectional forms, specially those of the verbs I-Y which are not available yet and the designed research methodology to identify these forms, including the sets of lemmatisation patterns of the non-canonical inflectional endings and stem vowels of weak verbs.
Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider the correlation between parameters and therefore can lead to unstable and slow convergence in training. Second-order optimization methods provide a possible solution to this issue. However these methods are either computationally heavy or do not have competitive performance. In this paper, a novel optimization method - stochastic natural gradient based on minimum variance assumption (SNGM) is proposed for training RNNLMs. It allows the natural gradient method to operate at a comparable training efficiency to the SGD method. By modifying the gradient according to the local curvature of the KL-divergence between current and updated probabilistic distributions, the proposed SNGM approach is shown to outperform both the SGD and limited memory BFGS methods across three tasks: Penn Treebank, Switchboard conversational speech recognition and AMI meeting room transcription in terms of both perplexity and word error rate.
Previous studies have shown that hoarding behavior usually starts at a subclinical level in early adolescence and gradually worsens; however, a limited number of studies have examined the prevalence of hoarding behavior and its association with developmental disorders in young adults. The aims of this study were to estimate the prevalence of hoarding behavior and to identify correlations between hoarding behavior and developmental disorder traits in university students. The study participants included 801 university students (616 men, 185 women) who completed questionnaires (ASRS: Adult ADHD Self-Report Scale version 1.1, AQ16: Autism-Spectrum Quotient with 16 items, and CIR: Clutter Image Rating). Among 801 participants, 27 (3.4%) exceeded the CIR cut-off score. Moreover, the participants with hoarding behavior had a significantly higher percentage of ADHD traits compared to participants without hoarding behavior (HB(+) vs HB(−), 40.7% vs 21.7%). In addition, 7.4% of HB(+) participants had autism spectrum disorder (ASD) traits, compared to 4.1% of HB(−) participants. A correlation analysis revealed that the CIR composite score had a stronger correlation with the ASRS inattentive score than with the hyperactivity/impulsivity score (CIR composite vs ASRS IA, r = 0.283; CIR composite vs ASRS H/I, r = 0.147). The results showed a high prevalence of ADHD traits in the university students with hoarding behavior. Moreover, we found that the hoarding behavior was more strongly correlated with inattentive symptoms rather than with hyperactivity/impulsivity symptoms. Our results support the concept of a common pathophysiology behind hoarding behavior and ADHD in young adults.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
The emergence of China as a global economic power in the 21st Century has brought about surging needs for cross-lingual and cross-cultural mediation, typically performed by translators. Advances in Artificial Intelligence and Language Engineering have been bolstered by Machine learning and suitable Big Data cultivation. They have helped to meet some of the translator's needs, though the technical specialists have not kept pace with the practical and expanding requirements in language mediation. One major technical and linguistic hurdle involves words outside the vocabulary of the translator or the lexical database he/she consults, especially Multi-Word Expressions (Compound Words) in technical subjects. A further problem lies in the multiplicity of renditions of a term in the target language. This paper discusses a proactive approach following the successful extraction and application of sizable bilingual Multi-Word Expressions (Compound Words) for language mediation in technical subjects, which do not fall within the expertise of typical translators, who have inadequate appreciation of the range of new technical tools available to help him/her. Our approach draws on the personal reflections of translators and teachers of translation and is based on the prior R&D efforts relating to 300,000 comparable Chinese-English patents. The subsequent protocol we have developed aims to be proactive in meeting four identified practical challenges in technical translation (e.g. patents). It has broader economic implication in the Age of Big Data (Tsou et al, 2015) and Trade War, as the workload, if not, the challenges, increasingly cannot be met by currently available front-line translators. We shall demonstrate how new tools can be harnessed to spearhead the application of language technology not only in language mediation but also in the “teaching” and “learning” of translation. It shows how a better appreciation of their needs may enhance the contributions of the technical specialists, and thus enhance the resultant synergetic benefits. © 2019 Incoma Ltd. All rights reserved.
A workshop on open resources for the original languages of the Bible in Copenhagen in March 2018 was the start of a new Copenhagen Alliance for Open Biblical Resources. The point of departure for the workshop was the need for programs and applications like Paratext and Bible Online Learner to have access to high-quality and reliable open data in order to assist Bible translators, teachers and students of Biblical Hebrew and New Testament Greek. The publication of contributions presents papers on methods for annotation, resources tracing patristic quotations and data for detached constructions in Biblical Hebrew. Reports cover tasks and data for Bible translation and research, treebanks, and applications like STEPBible and Bible Online Learner.
Abstract This chapter analyzes the interface between prosodic features and situational variables in the Rhapsodie treebank. First, we provide general quantitative information. Second, we present a preliminary set of statistical analyses performed on six specific prosodic variables, that provides two types of information. While Principal Component Analysis is useful to identify major trends in the data, inferential analysis is the first step in modeling the effect of the different modalities of each situational variable on the prosodic patterns, before considering predictive modeling. Finally, by focusing on the IPA unit, we present a third statistical method based on the computation of specificity indexes.
In this paper, we extend recent approaches to Lexicalized Tree Adjoining Grammar (LTAG) parsing that combine supertagging with dependency parsing. In other words, we assign supertags (= unanchored elementary trees) to lexical items and we compute substitution/adjunction arcs between them. Kasai et al. (2017, 2018) jointly predict these structures with a neural graph-based parser. Predicting 1-best supertags and dependency arcs (as in Kasai et al. (2017, 2018)) however leads only to partial parsing due to incompatibilities between elementary trees and derivation trees. We therefore extend the approach described in Kasai et al. (2017, 2018) to n-best supertags and k-best dependency arcs and combine it with a subsequent A*-parsing step that extends the TAG parser from Waszczuk (2017). We show that this architecture allows for efficient full TAG parsing while being sufficiently accurate. We test our architecture on an LTAG extracted from the French Treebank (FTB).
We recast dependency parsing as a sequence labeling problem, exploring several encodings of dependency trees as labels. While dependency parsing by means of sequence labeling had been attempted in existing work, results suggested that the technique was impractical. We show instead that with a conventional BiLSTM-based model it is possible to obtain fast and accurate parsers. These parsers are conceptually simple, not needing traditional parsing algorithms or auxiliary structures. However, experiments on the PTB and a sample of UD treebanks show that they provide a good speed-accuracy tradeoff, with results competitive with more complex approaches.
This two-phased, sequential mixed-methods study investigates how raters are influenced by different rating scales on a college-level English as a second language (ESL) writing placement test. In Phase I, nine certified raters rated 152 essays using a holistic, profile-based scale; in Phase II, they rated 200 essays using a binary, analytic scale developed based on the holistic scale and 100 essays using both rating scales. Ratings were examined both quantitatively through Rasch modeling and qualitatively via think-aloud protocols and semi-structured interviews. Findings from Phase I revealed that, despite satisfactory internal consistency, the raters demonstrated relatively low rater agreement and individual differences in their use of the holistic scale. Findings from Phase II showed that the binary, analytic scale led to much improvement in rater consensus and rater consistency. Another finding from Phase II suggests that the binary, analytic scale helped the raters deconstruct the holistic scale, reducing their cognitive burden. This study represents a creative use of a binary, analytic scale to guide raters through a holistic rating scale. Implications regarding how a rating scale affects rating behavior and performance are discussed.
In recent years, neural networks have been widely used for language modeling in different tasks of natural language processing. Results show that long short-term memory (LSTM) neural networks are appropriate for language modeling due to their ability to process long sequences. Furthermore, many studies are shown that extra information improve language models (LMs) performance. In this research, we propose parallel structures for incorporating part-of-speech tags into language modeling task using both the unidirectional and bidirectional type of LSTMs. Words and part-of-speech tags are given to the network as parallel inputs. In this way, to concatenate these two paths, two different structures are proposed according to the type of network used in the parallel part. We analyze the efficiency on Penn Treebank (PTB) dataset using perplexity measure. These two proposed structures show improvements in comparison to the baseline models. Not only does the bidirectional LSTM method gain the lowest perplexity, but it also has the lowest training parameters among our proposed methods. The perplexity of proposed structures has reduced 1.5% and %13 for unidirectional and bidirectional LSTMs, respectively.
WordNet is a lexical database for languages, the difference between WordNet and dictionaries in general is that WordNet focuses on the synonyms. The main unit of WordNet is synonym set (synset), synset is a set of one or more words that have the same meaning and certainly can be replaced in certain contexts. Synset is a very important element in implementing WordNet. In this paper, an analysis of the synonym extraction process is carried out by using commutative approach, the data test obtained from the Oxford Paperback Thesaurus by taking 51 word entries. Commutative method has similar characters with synonym set, synonym set can replace each other in certain contexts. The data test extraction process is carried out until the performance measurement evaluation process using F1Score. The system generates synonym sets that matched with the manual extraction, the result of F1Score between the program and Princeton synonym sets are worth 10%.
Different from the current syntax parsing based on deep learning, we present a novel Chinese parsing method, which is based on Sliding Match of Semantic String (SMOSS). (1) Training stage: In a treebank, headwords of tree nodes are represented by semantic codes given in the Synonym Dictionary (Tongyici Cilin). N-gram semantic templates are extracted from every layer of a syntax tree by means of sliding window to establish one N-gram semantic template library. (2) Parsing stage: Words of a sentence, including headwords of chunks, are represented by the semantic codes from Tongyici Cilin. With the sliding window method, N-gram semantic code strings are extracted to match with the templates in the N-gram semantic template library; subsequently, the mapping information of the matched templates is employed to guide the chunking of semantic code strings. The Chinese syntax parsing is completed through continuous matching and chunking. On the same training scale, N-gram semantic template can create favorable conditions for flexible matching and improve the syntax parsing performance. With train and test sets from the Tsinghua Chinese Treebank (TCT), the results are F1-score 99.71% (closed test) and F1-score 70.43% (open test), respectively.
Abstract We study the distribution of the nominal and copular construction of predicate nominals in a subset of authors from the Ancient Greek Dependency Treebank ( AGDT ). We concentrate on the texts of the historians Herodotus, Thucydides (both 5th century BCE) and Polybius (2nd century BCE). The data comprise a sample of 440 sentences (Hdt = 175, Thuc = 91, Pol = 174). We analyze the impact of four features that have been discussed in the literature and can be observed in the annotation of AGDT: (1) order of constituents, (2) part of speech of the subjects, (3) type of clause and (4) length of the clause. Furthermore, we test how the predictive power of these factors varies in time from Herodotus and Thucydides to Polybius with the help of a logistic-regression model. The analysis shows that, contrary to a simplistic opinion, the nominal construction does not drop into irrelevance in Hellenistic Greek. Moreover, an analysis of the distributions in the authors highlights a remarkable continuity in the usage patterns. Further work is needed to improve the predictive power of our logistic-regression model and to integrate more data in view of a more comprehensive quantitative diachronic study.
We investigate the relationship between the notion of nuclearity as proposed in Rhetorical Structure Theory (RST) and the signalling of coherence relations. RST relations are categorized as either mononuclear (comprising a nucleus and a satellite span) or multinuclear (comprising two or more nuclei spans). We examine how mononuclear relations (e.g., Antithesis, Condition) and multinuclear relations (e.g., Contrast, List) are indicated by relational signals, more particularly by discourse markers (e.g., because, however, if, therefore). We conduct a corpus study, examining the distribution of either type of relations in the RST Discourse Treebank (Carlson et al., 2002) and the distribution of discourse markers for those relations in the RST Signalling Corpus (Das et al., 2015). Our results show that discourse markers are used more often to signal multinuclear relations than mononuclear relations. The findings also suggest a complex relationship between the relation types and syntactic categories of discourse markers (subordinating and coordinating conjunctions).
Genital pain is a social experience that needs to be studied as a dyadic interaction between partners. The present study relied on a sample of 42 heterosexual couples to examine the level of congruence between both partners' ratings of pain and sexual arousal in response to experimentally induced vaginal pressure that served as a simulation of vaginal sensations during penetration. We also inferred the men's ability to estimate their partner's level of pain and sexual arousal. Because the relationship has shown to influence pain estimations, we considered the moderating role of perceived partner responsiveness and relationship satisfaction. We found higher disagreement in pain ratings when vaginal pressure was induced in the context of a sexual film compared to a neutral film, with men overestimating the level of pain in women. Also sexual arousal ratings diverged between partners, with men underestimating their partners' level of sexual arousal during the induction of vaginal pressure, regardless of whether they were watching a sexual or neutral film. Importantly, the level of congruence between actual and estimated ratings of pain and sexual arousal depended on how relationally satisfied men and women were and how validated and supported women felt by their male partner. These results make an important contribution to the growing literature on the social determinants of sexual pain experiences.
PURPOSE: Corneal confocal microscopy (CCM) is an imaging method to detect loss of nerve fibers in the cornea. The impact of image quality on the CCM parameters has not been investigated. We developed a quality index (QI) with 3 stages for CCM images and compared the influence of the image quality on the quantification of corneal nerve parameters using 2 modes of analysis in healthy volunteers and patients with known peripheral neuropathy. METHODS: Images of 75 participants were a posteriori analyzed, including 25 each in 3 image quality groups (QI 1-QI 3). Corneal nerve fiber length (CNFL) was analyzed using automated and semiautomated software, and corneal nerve fiber density and corneal nerve branch density were quantified using automated image analysis. Three masked raters assessed CCM image quality (QI) independently and categorized images into groups QI 1-QI 3. In addition, statistical analysis was used to compare interrater reliability. Analysis of variance was used for analysis between the groups. Interrater reliability analysis between the image ratings was performed by calculating Fleiss' kappa and its 95% confidence interval. RESULTS: CNFL, corneal nerve fiber density, and corneal nerve branch density increased significantly with QI (P < 0.001, all post hoc tests P < 0.05). CNFL was higher using semiautomated compared with automated nerve analysis, independent of QI. Fleiss kappa coefficient for interrater reliability of QI was 0.72. CONCLUSIONS: The quantification of corneal nerve parameters depends on image quality, and poorer quality images are associated with lower values for corneal nerve parameters. We propose the QI as a tool to reduce variability in quantification of corneal nerve parameters.
We inspect the multi-head self-attention in Transformer NMT encoders for three source languages, looking for patterns that could have a syntactic interpretation. In many of the attention heads, we frequently find sequences of consecutive states attending to the same position, which resemble syntactic phrases. We propose a transparent deterministic method of quantifying the amount of syntactic information present in the self-attentions, based on automatically building and evaluating phrasestructure trees from the phrase-like sequences. We compare the resulting trees to existing constituency treebanks, both manually and by computing precision and recall.
Despite advances in dependency parsing, languages with small treebanks still present challenges.