Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The Plain Meaning Rule is often assailed on the grounds that it is unprincipled—that it substitutes for careful analysis an interpreter’s ad hoc and impressionistic intuition about the meaning of legal texts. But what if judges and lawyers had the means to test their intuitions about plain meaning systematically? Then initial linguistic impressions about the meaning of a legal text might be viewed as hypotheses to be tested, rather than determinative criteria upon which to base important decisions. There exists very little legal scholarship on corpus linguistics—the study of language function and use through large, electronic linguistic databases called corpora—and the role that corpus methods might play in legal interpretation. This omission becomes more and more striking as scholars and jurists (and even the United States Supreme Court) have found themselves persuaded by corpus-based arguments. This Article argues that the plain or ordinary meaning of a given term in a given context is an empirical matter that may be quantified through corpus-based methods. These methods, when applied to questions of legal ambiguity, present significant advantages over existing empirical approaches to plain meaning and over the prevailing intuition-based interpretive approach of many courts. Because large, sophisticated linguistic corpora are widely available and easy to use, and because corpus methods offer a more principled and systematic alternative to the impressionistic interpretation of legal texts, corpus linguistics may one day revolutionize the process of legal interpretation.
Nonsuicidal self-injury (NSSI; e.g., cutting or burning the skin without suicidal intent) is a dangerous and increasingly prevalent health-risk behavior. Despite advances in NSSI research over the past decade, many aspects of NSSI remain poorly understood. In particular, there are few strong predictors of NSSI, it is unclear how positive attitudes toward NSSI develop, and there are no empirically supported treatments for NSSI. In the present study, I addressed these topics with a multi-method, experimental, and longitudinal approach. For Aim 1 of the study, I examined baseline differences between NSSI (n = 58) and control (n = 86) adult participants on NSSI-themed versions of five measures that cover different aspects of attitudes: the implicit association test (IAT); the affect misattribution procedure (AMP); explicit affective ratings; startle eyeblink reactivity; and startle postauricular reactivity. Compared to the control group, the NSSI group displayed significantly more positive attitudes on all five measures. Moreover, AMP scores and explicit ratings prospectively predicted self-cutting frequency over the ensuing six months. For Aim 2, I employed pain offset relief conditioning in an attempt to induce more positive implicit attitudes toward NSSI in the control group. This conditioning significantly diminished startle eyeblink reactivity in the context of NSSI images, but did not significantly affect any other measures. For Aim 3, I tested the ability of aversive conditioning in the NSSI group to reverse positive implicit attitudes toward NSSI and to reduce NSSI behaviors over the subsequent six months. Aversive conditioning normalized startle eyeblink and postauricular reactivity, but did not significantly affect any other measures. Results also provided preliminary support for the hypothesis that aversive conditioning prospectively reduces self-cutting. In conjunction with my other recent studies (Franklin et al., 2010; 2011, 2012, 2013), these findings have prompted a new theoretical framework called the Benefits and Barriers model of NSSI.
The practice of assessing brand management in construction in Ukraine is in a passive stage, but due to the entry into the Ukrainian market of foreign companies for which regular evaluation of their brand — the need for survival in a competitive environment, Ukrainian companies are beginning to pay more attention to the creation and formation of their competitive trading because of its high business image / rating. A modern toolkit based on appropriate approaches is used to form organizational and economic foundations. The article analyzes modern scientific approaches to assessing the economic potential of an enterprise, identifies the main trends and factors that affect the assessment of construction enterprises. The article is devoted to the study of theoretical and methodological assessments of branding of construction enterprises. The role and importance of innovation in ensuring the efficient operation of modern enterprises is emphasized. It is found that construction, especially innovative, is of great social importance and has a significant economic effect. The importance of determining the potential of innovative development of construction in general and construction enterprises in particular is substantiated. The purpose of this article is to systematically investigate the interpretation of the potential of innovative development of construction enterprises. The theoretical basis of the research is the scientific works of foreign and domestic scientists on the problems of identifying the essence of innovative development potential. A systematic study of the general characteristics of the construction company brand was conducted and the priority directions for choosing the development of the economic potential of the enterprises were determined. The necessity to understand the potential of the enterprise in the unity of all its elements, which are subject to the achievement of the overall goals of the enterprise, is substantiated. The weight of the component of the brand in the potential of the construction industry enterprises is substantiated. Some aspects of the development of theoretical and methodological approaches to the estimation of the intellectual capital of construction enterprises are formulated. Existing theoretical and methodological approaches to the brand assessment of construction enterprises are analyzed.
The relationship between words in a sentence often tell us more about the underlying semantic content of a document than its actual words individually. Natural language understanding has seen an increasing effort in the formation of techniques that try to produce non-trivial features, in the last few years, especially after robust word embeddings models became prominent, when they proved themselves able to capture and represent semantic relationships from massive amounts of data. These new dense vector representations indeed leverage the baseline in natural language processing, but they still fall short in dealing with intrinsic issues in linguistics, such as polysemy and homonymy. Systems that make use of natural language at its core, can be affected by a weak semantic representation of human language, resulting in inaccurate outcomes based on poor decisions. In this subject, word sense disambiguation and lexical chains have been exploring alternatives to alleviate several problems in linguistics, such as semantic representation, definitions, differentiation, polysemy, and homonymy. However, little effort is seen in combining recent advances in token embeddings (e.g. words, documents) with word sense disambiguation and lexical chains. To collaborate in building a bridge between these areas, this work proposes a collection of algorithms to extract semantic features from large corpora as its main contributions, named MSSA, MSSA-D, MSSA-NR, FLLC II, and FXLC II. The MSSA techniques focus on disambiguating and annotating each word by its specific sense, considering the semantic effects of its context. The lexical chains group derive the semantic relations between consecutive words in a document in a dynamic and pre-defined manner. These original techniques' target is to uncover the implicit semantic links between words using their lexical structure, incorporating multi-sense embeddings, word sense disambiguation, lexical chains, and lexical databases. A few natural language problems are selected to validate the contributions of this work, in which our techniques outperform state-of-the-art systems. All the proposed algorithms can be used separately as independent components or combined in one single system to improve the semantic representation of words, sentences, and documents. Additionally, they can also work in a recurrent form, refining even more their results.
يعد القرآن الكريم من مصادر المعرفة، وقد تولدت منه فروع واسعة؛ إذ نُزل القرآن الكريم باللغة العربية، ولا يوجد خيار آخر لإتقان المعرفة الواردة فيه إلا من خلال تعلم اللغة العربية. تهدف هذه الدراسة إلى بيان مفهوم المدونة العربية القرآنية ومكوناتها، والكشف عن علاقة تعلم اللغة العربية بالقرآن الكريم، وبيان كيفية تعليم وتعلم القواعد العربية الأساسية عبر المدونة العربية القرآنية، وستتبع الدراسة المنهج الوصفي والتحليلي. إن وجود العلاقة بين اللغة العربية والقرآن الكريم، يدفع الطلبة المتخصصين في اللغة العربية أن يربطوا اللغة العربية بالقرآن؛ لذلك نرى أن المدونة العربية القرآنية تساعدهم على فهم القواعد القرآنية بطريقة مثيرة للاهتمام. في نظرة شاملة يمكن أن نستنتج أن المدونة العربية القرآنية هي واحدة من أهم الأدوات الحسابية التي تم إنتاجها في خدمة اللغة العربية؛ حيث توفر للمتعلمين ما يحتاجون إليه في مجال اللغة واللغويات والدراسات الحاسوبية، كما تمهد الطريق للباحثين لدراسة الهياكل المورفولوجية والنحوية من خلال دراسات الحوسبة العميقة للقرآن.
 الكلمات المفتاحية: المصرف القرآني، نموذج حاسوبي، المدونة العربية القرآنية، المعجم القرآني.
 Abstract 
 The Holy Quran is a source of knowledge and it has generated wide branches of knowledge. The Holy Quran was revealed in Arabic. Hence, there is no other option to master its knowledge except by learning the Arabic language. This study aims at explaining the concept of the Arabic Quranic Corpus and its components, revealing the relationship between learning the Arabic language and the Holy Quran, and showing how to teach and learn basic Arabic grammar through the Quranic Arabic Corpus. The study will follow the descriptive and analytical approach. The existence of the relationship between the Arabic language and the Holy Quran prompts Arabic learners to associate Arabic with the Qur'an. Therefore, we see that the Quranic Arabic Corpus helps them to understand Quranic rules in an interesting way. In a comprehensive view, we can conclude that the Arabic Quranic Corpus is one of the most important web-based medium produced to serve the Arabic language. It provides learners with what they need in the field of language, linguistics and computer studies, and paves the way for researchers to study morphological and grammatical structures through technology with detail description of grammars.
 Keywords: Quranic Treebank, Computational Model, Arabic Quranic Corpus, Qur’anic Dictionary.
Paper is dedicated to the testing of the concept of literacy, based on the questionnaire, carried out in the school year of 2018/2019 among the students of two secondary vocational schools in Vrsac, Belgrade and Grammar School in Vrsac (200 respondents). The primary hypothesis of the research was that detection and detailed study of high school students conceptosphere on literacy identify the fields to improve the teaching of Serbian as a mother tongue in secondary schools and the aim of work that, based on the collected and then processed data in analytical, cognitive and descriptive method, is to (a) isolate the dominant concepts of (non)literacy, (b) look at the tendencies of spreading and shaping the notion of literacy induced by the needs of modern life, and also that, in order to improve linguistic culture in all domains and all educational levels -(c) point to the possibility of improving the teaching of the Serbian language as a mother tongue. According to results of the survey secondary school students experience literacy in the 21 st century as a complex concept; from the one who is literate expecting linguistic knowledge, what are the basic, traditionally accepted parameters, and recognize illiteracy as the lack of ability to apply knowledge in the field of language. They also demonstrated that it is necessary to improve the efficiency of teaching approaches designed to improve functional literacy in a variety of communicative situations; increase the number of hours and exercises in the field of spelling, or nurture and acquire more comprehensive and knowledge in use and skills of different forms of literacy needed for managing in 21 st century; more attention should be paid to including relevant language handbooks in teaching; more explicit, on frequent and more familiar examples to students, point to the advantages of knowing and respecting the linguistic norm, paving the way for a better linguistic culture and enrichment of the mother tongue.
Speech processing systems rely on robust feature extraction to handle phonetic and semantic variations found in natural language. While techniques exist for desensitizing features to common noise patterns produced by Speech-to-Text (STT) and Text-to-Speech (TTS) systems, the question remains how to best leverage state-of-the-art language models (which capture rich semantic features, but are trained on only written text) on inputs with ASR errors. In this paper, we present Telephonetic, a data augmentation framework that helps robustify language model features to ASR corrupted inputs. To capture phonetic alterations, we employ a character-level language model trained using probabilistic masking. Phonetic augmentations are generated in two stages: a TTS encoder (Tacotron 2, WaveGlow) and a STT decoder (DeepSpeech). Similarly, semantic perturbations are produced by sampling from nearby words in an embedding space, which is computed using the BERT language model. Words are selected for augmentation according to a hierarchical grammar sampling strategy. Telephonetic is evaluated on the Penn Treebank (PTB) corpus, and demonstrates its effectiveness as a bootstrapping technique for transferring neural language models to the speech domain. Notably, our language model achieves a test perplexity of 37.49 on PTB, which to our knowledge is state-of-the-art among models trained only on PTB.
Lexical Markup Framework (LMF) or ISO 24613 [1] is a de jure standard that\nprovides a framework for modelling and encoding lexical information in\nretrodigitised print dictionaries and NLP lexical databases. An in-depth review\nis currently underway within the standardisation subcommittee,\nISO-TC37/SC4/WG4, to find a more modular, flexible and durable follow up to the\noriginal LMF standard published in 2008. In this paper we will present some of\nthe major improvements which have so far been implemented in the new version of\nLMF.\n
Meishan Zhang, Yue Zhang, Guohong Fu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
The intelligent information processing of the standard Zhuang language spoken mainly in Southern China is presently in its infancy, and lacks a well-defined language corpus and automatic part-of-speech tagging methods. Therefore, this study proposes an adversarial part-of-speech tagging method based on reinforcement learning, which solves the problems associated with a lack of a language corpus, time-consuming laborious manual marking, and the low performance of machine marking. Firstly, we construct a markup dictionary based on the grammatical characteristics of standard Zhuang and the Penn Chinese Treebank. Secondly, a dependency syntax analysis is applied for constructing the semantic information feature vectors of sentences, and long short-term memory is adopted as the policy network architecture to enhance available information using recurrent memory, and a conditional random field is employed as the discriminant network to perform label inference with global normalization. Finally, we use reinforcement learning as the model framework, target parts of speech as the feedback of the environment, and then obtain the optimal policy through adversarial learning. The results show that the combination of reinforcement learning and adversarial network alleviates the dependence of the model on the training corpus to some extent, and can quickly and effectively expand the scale of the annotation dictionary for the Zhuang language, thereby obtaining better labeling results.
According to World Intellectual Property Organization (2017) report, over 3 million patents exist in the patent database, but only certain numbers have commercial potential. Generally, to assess the commercial potential of patent, it consumes time and requires various expertise. Currently, several models have been developed to address this matter, which to assess using questionnaire tool for portfolios by human. So that occurs bias any limitation exists that our research will address by artificial intelligence. Hence, this research applies a Natural Language Programming to assess for commercial potential of patent, consisting of five steps - (i) Morphological analysis based on the Lexical database, (ii) Syntactic analysis of sentence to check syntax sentence patterns, (iii) Sematic analysis to interpret the meaning of words derived from the previous step, (iv) Discourse integration from context of domain together with the main sentence providing more accurate sentence analysis, and (v) Pragmatic analysis to ensure the correct meaning of interpretation. Then, the obtained data is used to determine criterion factors and formulate the model for assessing commercial potential of patent using Natural Language programming. This finding should deliver an alternative effective patent assessment system, which addresses some current deficiency in patent's assessment for commercial potential.
espanolResumen en castellano: Esta tesis trata de la morfologia verbal del ingles antiguo para identificar y lematizar los verbos debiles de esta lengua en un corpus al que se accede a traves de una base de datos lexica. La lematizacion es una de las tareas mas importantes a la hora de construir un diccionario. Sin embargo, es una de las tareas pendientes en el campo de la linguistica historica debido a que no existen corpora exhaustivos y lematizados de esta lengua. El enfoque de esta tesis doctoral esta en la lematizacion de las tres clases de verbos debiles del ingles antiguo, aunque las areas de la Lexicografia y la Linguistica de Corpus son tambien relevantes para esta investigacion. Las fuentes principales de esta investigacion son las formas flexivas que estan atestiguadas en el Dictionary of Old English Corpus (DOEC) y que estan disponibles en el lematizador Norna, las fuentes lexicograficas que existen publicadas sobre esta lengua, principalmente el Dictionary of Old English (DOE), y otras fuentes textuales como el York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) y una indexacion de fuentes secundarias del ingles antiguo. El objetivo principal supone la identificacion de las flexiones de los verbos debiles y de su lematizacion con uno de los lemas propuestos en las listas de referencia. Conseguir este objetivo implica manejar las fuentes disponibles en ingles antiguo para poder lematizar y validar los resultados del analisis y el diseno de un metodo que combine busquedas automaticas en la base de datos lexica Nerthus y la revision manual de los resultados. La metodologia incluye cuatro pasos sucesivos con diversas tareas en cada paso. El primero de estos pasos tiene como objetivo la lematizacion de las formas canonicas de los verbos debiles lanzando cadenas de busquedas especificas para cada clase de verbos debiles en el lematizador Norna, donde esta disponible un indice de tipos del DOEC, la fuente de informacion mas fiable de la que se dispone en ingles antiguo. Despues, los resultados se validan con el DOE y se anaden las formas no-canonicas de los verbos debiles entre las letras A y H. El tercer paso tiene como objetivo identificar las formas no-canonicas de las terminaciones flexivas y de las vocales de los radicales que aparecen con mas frecuencia en los verbos debiles para generar patrones de lematizacion. La busqueda de estos patrones y de la lista de prefijos no-canonicos que esta disponible en Norna culmina en la lematizacion de las formas flexivas no transparentes de los verbos debiles. La validacion de los resultados de las letras I a la Y supone el ultimo paso de la metodologia, donde se comparan los datos obtenidos con el analisis sintactico del YCOE y con los datos que se obtienen de una base de datos de indexacion de las fuentes secundarias del ingles antiguo. Los problemas que surgen a lo largo del proceso de lematizacion tienen que ver principalmente con las peculiaridades del ingles antiguo y las limitaciones de la lematizacion de tipos que esta investigacion sigue. La discusion de los resultados del analisis concluye esta tesis. Las principales aportaciones de esta tesis son las listas de lemas y sus formas flexivas, especialmente las de los verbos entre las letras I y la Y ya que no estan disponibles todavia, y el metodo que se ha disenado para identificar estas formas, incluyendo los patrones de lematizacion generados para lematizar las formas con terminaciones no comunes y vocales no canonicas en el radical. EnglishThis thesis deals with the verbal morphology of the Old English language in order to identify and lemmatise weak verbs in a corpus accessed through a lexical database. Lemmatisation is a pending task in the field of historical linguistics given the lack of comprehensive and lemmatised corpora in this language. The focus of this doctoral dissertation is on the lemmatisation of the three classes of weak verbs, although the linguistic fields of Lexicography and Corpus Linguistics are also relevant to this research. The main aim involves the identification of the canonical and non-canonical realisations of the Old English weak verbs and their lemmatisation with a lemma from a reference list of weak verbs. Achieving this goal involves, firstly, the use of the available sources of the Old English language in order to lemmatise and validate the results and, secondly, the design of a semi-automatic research methodology that combines automatic searches in the lexical database Nerthus and the manual revision of the results in order to achieve this task. The sources for this investigation are the inflectional forms that are attested in the Dictionary of Old English Corpus (DOEC) which are available in the lemmatiser Norna, the lexicographical sources published on the Old English language, mainly the Dictionary of Old English (DOE), and other textual sources such as the York-Toronto-Helsinki Parsed Corpus of Old English (YCOE) and an index of secondary sources of Old English. The methodology comprises four successive steps and several tasks within each step. The first step aims at the lemmatisation of the transparent forms of weak verbs with the search of specific query strings for each subclass of weak verbs in the lemmatiser Norna, where an index type of the DOEC, the most reliable source of information regarding the Old English language, is available. Then, the second step validates the results with the DOE and adds to the analysis the non-canonical attestations for the weak verbs from the letter A-H. Thirdly, the identification of the most recurrent non-canonical inflectional endings and stem vowels attested in weak verbs gives rise to lemmatisation patterns. The search of these sets of correspondences and the list of non-canonical prefixes that is available in Norna results in the lemmatisation of the non-canonical inflections of weak verbs. The validation of the results from the letter I-Y concludes the research methodology with the syntactic parsing provided by the YCOE and the data retrieved from the index of secondary sources of Old English Freya. The issues that arise throughout the lemmatisation process mainly concern the idiosyncrasy of the Old English language writing system and the limitations of the lemmatisation by type that this investigation follows. The quantitative and qualitative discussion of the results of the analysis concludes this thesis. The main contributions of this thesis are the lists of weak lemmas and their lemmatised inflectional forms, specially those of the verbs I-Y which are not available yet and the designed research methodology to identify these forms, including the sets of lemmatisation patterns of the non-canonical inflectional endings and stem vowels of weak verbs.
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple LS approach that makes use of the Bidirectional Encoder Representations from Transformers (BERT) which can consider both the given sentence and the complex word during generating candidate substitutions for the complex word. Specifically, we mask the complex word of the original sentence for feeding into the BERT to predict the masked token. The predicted results will be used as candidate substitutions. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement compared with these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 12 Accuracy points on three well-known benchmarks.
Named Entity Recognition (NER) for Myanmar Language is essential to Myanmar natural language processing research work. In this work, NER for Myanmar language is treated as a sequence tagging problem and the effectiveness of deep neural networks on NER for Myanmar language has been investigated. Experiments are performed by applying deep neural network architectures on syllable level Myanmar contexts. Very first manually annotated NER corpus for Myanmar language is also constructed and proposed. In developing our in-house NER corpus, sentences from online news website and also sentences supported from ALT-Parallel-Corpus are also used. This ALT corpus is one part of the Asian Language Treebank (ALT) project under ASEAN IVO. This paper contributes the first evaluation of neural network models on NER task for Myanmar language. The experimental results show that those neural sequence models can produce promising results compared to the baseline CRF model. Among those neural architectures, bidirectional LSTM network added CRF layer above gives the highest F-score value. This work also aims to discover the effectiveness of neural network approaches to Myanmar textual processing as well as to promote further researches on this understudied language.
Walter de Bibbesworth is known primarily for his Tretiz, a rhyming vocabulary of French written at some point between the years 1230 and 1270. According to a prologue transmitted in some manuscripts, the text was written at the request of Dyonise de Munchensi, wife of the powerful lord Warin de Munchensi. Another work attributed to Bibbesworth, a tençon or debate poem written with the Earl of Lincoln, Henry de Lacy, probably dates from the time of the 1270 Crusade. In it, the young lord Henry asks the older Walter to advise on a dilemma: should he honour his vow to go to the Holy Land for the love of Christ, or stay home for the love of his lady? Walter attempts to convince him to place love of the divine before secular motives. Alongside these two works, Bibbesworth is credited with at least one other poetic composition, ‘Amours m'ount si enchaunté’, which revels in the same sense of wordplay and interest in homophony that we see on display in the Tretiz. The two shorter poems have been almost entirely ignored by scholars, a neglect which the present article aims to correct. In the following pages, I argue that reading the shorter poems alongside the Tretiz can elucidate the grammar of poetry in the lyric materials and the poetry of grammar in the treatise. Grammar, the heart of the medieval educational curriculum, exceeded the bounds of strictly linguistic study. The disciplines of grammar (what language is) and rhetoric (what language can do) often overlapped in practice, especially in the description of figurative or expressive language; from Antiquity onwards, ‘the principles of grammar were seen as the key to understanding poetic form’. Grammatical study also contained an inalienable moral aspect; as Paul Gehl comments, ‘the very idea of a linguistic norm was charged with moral meaning’, and several of the texts used for study owed their selection to didactic content as much as linguistic interest. To reflect on Latin language, within the context of grammatical education, thus simultaneously compelled reflection on questions of morality and poetic form, and recent scholarship has very fruitfully explored how the educational experiences of medieval authors informed and shaped the literary works they subsequently produced.
Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider the correlation between parameters and therefore can lead to unstable and slow convergence in training. Second-order optimization methods provide a possible solution to this issue. However these methods are either computationally heavy or do not have competitive performance. In this paper, a novel optimization method - stochastic natural gradient based on minimum variance assumption (SNGM) is proposed for training RNNLMs. It allows the natural gradient method to operate at a comparable training efficiency to the SGD method. By modifying the gradient according to the local curvature of the KL-divergence between current and updated probabilistic distributions, the proposed SNGM approach is shown to outperform both the SGD and limited memory BFGS methods across three tasks: Penn Treebank, Switchboard conversational speech recognition and AMI meeting room transcription in terms of both perplexity and word error rate.
Previous studies have shown that hoarding behavior usually starts at a subclinical level in early adolescence and gradually worsens; however, a limited number of studies have examined the prevalence of hoarding behavior and its association with developmental disorders in young adults. The aims of this study were to estimate the prevalence of hoarding behavior and to identify correlations between hoarding behavior and developmental disorder traits in university students. The study participants included 801 university students (616 men, 185 women) who completed questionnaires (ASRS: Adult ADHD Self-Report Scale version 1.1, AQ16: Autism-Spectrum Quotient with 16 items, and CIR: Clutter Image Rating). Among 801 participants, 27 (3.4%) exceeded the CIR cut-off score. Moreover, the participants with hoarding behavior had a significantly higher percentage of ADHD traits compared to participants without hoarding behavior (HB(+) vs HB(−), 40.7% vs 21.7%). In addition, 7.4% of HB(+) participants had autism spectrum disorder (ASD) traits, compared to 4.1% of HB(−) participants. A correlation analysis revealed that the CIR composite score had a stronger correlation with the ASRS inattentive score than with the hyperactivity/impulsivity score (CIR composite vs ASRS IA, r = 0.283; CIR composite vs ASRS H/I, r = 0.147). The results showed a high prevalence of ADHD traits in the university students with hoarding behavior. Moreover, we found that the hoarding behavior was more strongly correlated with inattentive symptoms rather than with hyperactivity/impulsivity symptoms. Our results support the concept of a common pathophysiology behind hoarding behavior and ADHD in young adults.
Creating an opinionated lexicon is an important step towards a reliable social media analysis system. In this article we are proposing an approach and describing an experiment to build an Arabic polarised lexical database from analysing online implicitly and explicitly rated customer reviews. These reviews are written in modern standard Arabic and Palestinian/Jordanian dialect. Therefore, the produced lexicon contains casual slangs and dialectic entries used by the online community, which is useful for sentiment analysis of informal social media micro-blogs. We have extracted 28,000 entries from processing 15,100 reviews and by expanding the initial lexicon through Google translate. We calculated an implicit rating for every review driven by its text to address the problem of ambiguous opinions of certain online posts, where the text of the review does not match the given rating (the explicit rating). Each entry was given a polarity tag and a confidence score. High confidence scores have increased the precision of the polarisation process. Explicit rating has increased the coverage and confidence of polarity.
The emergence of China as a global economic power in the 21st Century has brought about surging needs for cross-lingual and cross-cultural mediation, typically performed by translators. Advances in Artificial Intelligence and Language Engineering have been bolstered by Machine learning and suitable Big Data cultivation. They have helped to meet some of the translator's needs, though the technical specialists have not kept pace with the practical and expanding requirements in language mediation. One major technical and linguistic hurdle involves words outside the vocabulary of the translator or the lexical database he/she consults, especially Multi-Word Expressions (Compound Words) in technical subjects. A further problem lies in the multiplicity of renditions of a term in the target language. This paper discusses a proactive approach following the successful extraction and application of sizable bilingual Multi-Word Expressions (Compound Words) for language mediation in technical subjects, which do not fall within the expertise of typical translators, who have inadequate appreciation of the range of new technical tools available to help him/her. Our approach draws on the personal reflections of translators and teachers of translation and is based on the prior R&D efforts relating to 300,000 comparable Chinese-English patents. The subsequent protocol we have developed aims to be proactive in meeting four identified practical challenges in technical translation (e.g. patents). It has broader economic implication in the Age of Big Data (Tsou et al, 2015) and Trade War, as the workload, if not, the challenges, increasingly cannot be met by currently available front-line translators. We shall demonstrate how new tools can be harnessed to spearhead the application of language technology not only in language mediation but also in the “teaching” and “learning” of translation. It shows how a better appreciation of their needs may enhance the contributions of the technical specialists, and thus enhance the resultant synergetic benefits. © 2019 Incoma Ltd. All rights reserved.
In recent years, dependency parsing is a fascinating research topic and has a lot of applications in natural language processing. In this paper, we present an effective approach to improve dependency parsing by utilizing supertag features. We performed experiments with the transition-based dependency parsing approach because it can take advantage of rich features. Empirical evaluation on Vietnamese Dependency Treebank showed that, we achieved an improvement of 18.92% in labeled attachment score with gold supertags and an improvement of 3.57% with automatic supertags.
We introduce a language-agnostic evolutionary technique for automatically extracting chunks from dependency treebanks. We evaluate these chunks on a number of morphosyntactic tasks, namely POS 1 tagging, morphological feature tagging, and dependency parsing. We test the utility of these chunks in a host of different ways. We first learn chunking as one task in a shared multitask framework together with POS and morphological feature tagging. The predictions from this network are then used as input to augment sequence-labelling dependency parsing. Finally, we investigate the impact chunks have on dependency parsing in a multi-task framework. Our results from these analyses show that these chunks improve performance at different levels of syntactic abstraction on English UD treebanks and a small, diverse subset of non-English UD treebanks.
The recent availability of large on-line parsed corpora makes it possible to test theories of psycholinguistic complexity by comparing the frequency distributions of closely related constructions. In this paper, we use this technique to test the psycholinguistic theory proposed by Gibson et al. (1993), which includes two independently motivated attachment principles: Recency Preference and Predicate Proximity. In order to test this theory, we examined two general classes of attachment ambiguities from the parsed Wall Street Journal corpus from the Penn Treebank: 1) ambiguities which involve three prospective noun phrase attachment sites; and 2) ambiguities which involve three prospective verb phrase attachment sites. Given three prospective noun phrase (NP) sites in English, the theory most naturally predicts a complexity ordering of NP3 (easiest, most recent), NP1, NP2, but a ranking of VP3, VP2, VP1 for verb phrase attachments. Our corpus analyses support both of these predictions.
Recurrent Neural Networks (RNNs) have dominated language modeling because of their superior performance over traditional N-gram based models. In many applications, a large Recurrent Neural Network language model (RNNLM) or an ensemble of several RNNLMs is used. These models have large memory footprints and require heavy computation. In this paper, we examine the effect of applying knowledge distillation in reducing the model size for RNNLMs. In addition, we propose a trust regularization method to improve the knowledge distillation training for RNNLMs. Using knowledge distillation with trust regularization, we reduce the parameter size to a third of that of the previously published best model while maintaining the state-of-the-art perplexity result on Penn Treebank data. In a speech recognition N-bestrescoring task, we reduce the RNNLM model size to 18.5% of the baseline system, with no degradation in word error rate(WER) performance on Wall Street Journal data set.
Il progetto LiLa: Linking Latin intende creare una Knowledge Base di risorse linguistiche (corpora, lessici digitali e strumenti di trattamento automatico del linguaggio) per lo studio del latino secondo il modello dei Linked Open Data. Questo articolo descrive gli obiettivi e le motivazioni del progetto, soffermandosi in particolare sulla centralità del lemma come forma che consente di creare la rete di informazioni linguistiche. L’architettura di LiLa s’incentra dunque sul nodo cardine rappresentato dal lemma e sulle sue proprietà morfologiche. L’articolo discute le strategie impiegate per creare una banca dati di lemmi latini che possa supportare tale modello, e i primi esperimenti volti a connettere risorse testuali (treebank del latino) alla raccolta dei lemmi.
Automatically generating descriptive captions for images is a well-researched area in computer vision. However, existing evaluation approaches focus on measuring the similarity between two sentences disregarding fine-grained semantics of the captions. In our setting of images depicting persons interacting with branded products, the subject, predicate, object and the name of the branded product are important evaluation criteria of the generated captions. Generating image captions with these constraints is a new challenge, which we tackle in this work. By simultaneously predicting integer-valued ratings that describe attributes of the human-product interaction, we optimize a deep neural network architecture in a multi-task learning setting, which considerably improves the caption quality. Furthermore, we introduce a novel metric that allows us to assess whether the generated captions meet our requirements (i.e., subject, predicate, object, and product name) and describe a series of experiments on caption quality and how to address annotator disagreements for the image ratings with an approach called soft targets. We also show that our novel clause-focused metrics are also applicable to other image captioning datasets, such as the popular MSCOCO dataset.
It is widely accepted that the human cognitive system organizes perceptual input into complex hierarchical descriptions which can be represented by tree structures. Tree structures have been used to describe linguistic, musical and visual perception. In this paper, we will investigate whether there exists an underlying model that governs perceptual organization in general. Our key idea is that the cognitive system strives for the simplest structure (the “simplicity principle”), but in doing so it is biased by the likelihood of previous experiences (the “likelihood principle”). We will present a model which combines these two principles by balancing the notion of most likely tree with the notion of shortest derivation. Experiments with linguistic and musical benchmarks (Penn Treebank and Essen Folksong Collection) show that such a combination outperforms models that are based on either simplicity or likelihood alone.
This chapter engages with research in lingua franca scenarios, defining this term and arguing for the value of adopting a linguistic ethnographic approach in this area. The development of the field of English as a Lingua Franca (ELF) is outlined. Historical antecedents for lingua franca studies in interactional sociolinguistics are identified, including work on intercultural miscommunications, the study of English as an international language and the identification of pragmatic strategies in ELF scenarios such as the let-it-pass procedure. The chapter considers critical debates, including the need to challenge the norms of standard language pedagogies, and the importance of maintaining a critical perspective on the current global dominance of English. It provides a review of current areas of focus in this area, including studies of English as lingua franca in higher education in an increasingly internationalised university system, and in workplaces. The range of methods drawn on in ethnographic studies of lingua franca scenarios is described, and implications for practice are identified in relation to the teaching of English and language policy and planning. Directions for future research identified include the emergence of social and linguistic norms in interaction, and developing the focus on lingua francas other than English.
Discourse relation classification has proven to be a hard task, with rather low performance on several corpora that notably differ on the relation set they use. We propose to decompose the task into smaller, mostly binary tasks corresponding to various primitive concepts encoded into the discourse relation definitions. More precisely, we translate the discourse relations into a set of values for attributes based on distinctions used in the mappings between discourse frameworks proposed by This arguably allows for a more robust representation of discourse relations, and enables us to address usually ignored aspects of discourse relation prediction, namely multiple labels and underspecified annotations. We study experimentally which of the conceptual primitives are harder to learn from the Penn Discourse Treebank English corpus, and propose a correspondence to predict the original labels, with preliminary empirical comparisons with a direct model.
When using computer-aided translation systems in a typical, professional translation workflow, there are several stages at which there is room for improvement. The SCATE (Smart Computer-Aided Translation Environment) project investigated several of these aspects, both from a human-computer interaction point of view, as well as from a purely technological side. This paper describes the SCATE research with respect to improved fuzzy matching, parallel treebanks, the integration of translation memories with machine translation, quality estimation, terminology extraction from comparable texts, the use of speech recognition in the translation process, and human computer interaction and interface design for the professional translation environment. For each of these topics, we describe the experiments we performed and the conclusions drawn, providing an overview of the highlights of the entire SCATE project.
In the paper the analysis of the perspective of the technology of intelligent monitoring systems is conducted, as intellectualization is the main direction of development of modern technologies, and the property of intellectuality should be inherent in all the latest information management systems. Various strategies for intellectualization of monitoring are aimed at implementing intellectual information support for decision-makers using monitoring tools. Such support can be realized by building fuzzy linguistic databases/knowledge together with fuzzy inference subsystems, and information for decision making can be displayed on the automated workplace of the decision maker.
Implicit causal relation recognition aims to identify the causal relation between a pair of arguments. It is a challenging task due to the lack of conjunctions and the shortage of labeled data. In order to improve the identification performance, we come up with an approach to expand the training dataset. On the basis of the hypothesis that there inherently exists causal relations in WHY-type Question-Answer (QA) pairs, we utilize WHY-type QA pairs for the training set expansion. In practice, we first collect WHY-type QA pairs from the Knowledge Bases (KBs) of the reading comprehension tasks, and then convert them into narrative argument pairs by Question-Statement Conversion (QSC). In order to alleviate redundancy, we use active learning (AL) to select informative samples from the synthetic argument pairs. The sampled synthetic argument pairs are added to the Penn Discourse Treebank (PDTB), and the expanded PDTB is used to retrain the neural network-based classifiers. Experiments show that our method yields a performance gain of 2.42% F 1-score when AL is used, and 1.61% without using.
The aim of the present study was to examine whether offspring at high and low familial risk for depression differ in the immediate and more lasting behavioural and physiological effects of hedonically-based mood repair. Participants (9- to 22-year olds) included never-depressed offspring at high familial depression risk (high-risk, n = 64), offspring with similar familial background and personal depression histories (high-risk/DEP, n = 25), and never-depressed offspring at low familial risk (controls, n = 62). Offspring provided affect ratings at baseline, after sad mood induction, immediately following hedonically-based mood repair, and at subsequent, post-repair epochs. Physiological reactivity, indexed via respiratory sinus arrhythmia (RSA), was assessed during the protocol. Following mood induction and mood repair, high- and low-risk (control) offspring reported comparable changes in levels of sadness and RSA. However, sadness increased among high-risk offspring following the post-repair epoch, whereas low-risk offspring maintained mood repair benefits. High-risk/DEP offspring also reported higher levels of sadness following the post-repair epoch than did low-risk offspring. Change in RSA did not differ across the three offspring groups. Self-ratings confirm that one source of difficulty associated with depression risk is diminished ability to maintain hedonically-based mood repair gains, which were not apparent at the physiological level.
We present a strategy to automate the extraction of semantic relations from texts. Both machine learning and rule-based techniques are investigated and the impact of different linguistic knowledge is analyzed for the various approaches. To implement the extraction system RExtractor, several natural language processing tools have been improved: from sentence splitting and tokenization modules to dependency syntax parsers. Furthermore, we created the Czech Legal Text Treebank with several layers of linguistic annotation, which is used to train and test each stage of the proposed system. As a result of the performed work, new Semantic Web resources and tools are available for automatic processing of texts.
A workshop on open resources for the original languages of the Bible in Copenhagen in March 2018 was the start of a new Copenhagen Alliance for Open Biblical Resources. The point of departure for the workshop was the need for programs and applications like Paratext and Bible Online Learner to have access to high-quality and reliable open data in order to assist Bible translators, teachers and students of Biblical Hebrew and New Testament Greek. The publication of contributions presents papers on methods for annotation, resources tracing patristic quotations and data for detached constructions in Biblical Hebrew. Reports cover tasks and data for Bible translation and research, treebanks, and applications like STEPBible and Bible Online Learner.
Abstract This chapter analyzes the interface between prosodic features and situational variables in the Rhapsodie treebank. First, we provide general quantitative information. Second, we present a preliminary set of statistical analyses performed on six specific prosodic variables, that provides two types of information. While Principal Component Analysis is useful to identify major trends in the data, inferential analysis is the first step in modeling the effect of the different modalities of each situational variable on the prosodic patterns, before considering predictive modeling. Finally, by focusing on the IPA unit, we present a third statistical method based on the computation of specificity indexes.
In this paper, we extend recent approaches to Lexicalized Tree Adjoining Grammar (LTAG) parsing that combine supertagging with dependency parsing. In other words, we assign supertags (= unanchored elementary trees) to lexical items and we compute substitution/adjunction arcs between them. Kasai et al. (2017, 2018) jointly predict these structures with a neural graph-based parser. Predicting 1-best supertags and dependency arcs (as in Kasai et al. (2017, 2018)) however leads only to partial parsing due to incompatibilities between elementary trees and derivation trees. We therefore extend the approach described in Kasai et al. (2017, 2018) to n-best supertags and k-best dependency arcs and combine it with a subsequent A*-parsing step that extends the TAG parser from Waszczuk (2017). We show that this architecture allows for efficient full TAG parsing while being sufficiently accurate. We test our architecture on an LTAG extracted from the French Treebank (FTB).
We recast dependency parsing as a sequence labeling problem, exploring several encodings of dependency trees as labels. While dependency parsing by means of sequence labeling had been attempted in existing work, results suggested that the technique was impractical. We show instead that with a conventional BiLSTM-based model it is possible to obtain fast and accurate parsers. These parsers are conceptually simple, not needing traditional parsing algorithms or auxiliary structures. However, experiments on the PTB and a sample of UD treebanks show that they provide a good speed-accuracy tradeoff, with results competitive with more complex approaches.
In this paper we present a work which aims to test the most advanced, state-of-the-art syntactic dependency parsers based on deep neural networks (DNN) on Italian. We made a large set of experiments by using two Italian treebanks containing different text types downloaded from the Universal Dependencies project and propose a new solution based on ensemble systems. We implemented the proposed ensemble solutions by testing different techniques described in literature, obtaining very good parsing results, well above the state of the art for Italian.
Lexical Markup Framework (LMF) or ISO 24613 [1] is a de jure standard that provides a framework for modelling and encoding lexical information in retrodigitised print dictionaries and NLP lexical databases. An in-depth review is currently underway within the standardisation subcommittee, ISO-TC37/SC4/WG4, to find a more modular, flexible and durable follow up to the original LMF standard published in 2008. In this paper we will present some of the major improvements which have so far been implemented in the new version of LMF.
Objective. The article presents the results of the empirical study of the impact of the Internet using experience on the process of the Internet texts understanding.
 Materials & Methods. Different theoretical methods and techniques were used for this purpose: deductive and inductive methods, analysis and synthesis, generalization, systematization. Empirical methods were used for this purpose: experiment (semantic and receptive), method of semantic and pragmatic interpretations, content analysis, subjective scaling procedure. Mathematical methods were used: primary statistics, checks on the normal nature of the data distribution, statistical output, taking into account statistical indicators of fashion and the scope of variation. As well as some interpretive methods that are based on specific principles of systemic, activity, cognitive and organizational approaches.
 Results. The author notes that Internet texts understanding is significantly different from the understanding of oral or written texts, since the Internet text is a pragmatically integral electronic document that constructing of conditionally completed text blocks in the form of «windows», that are opened in separate tabs of the browser, the order that depends on hyperlinks and user behavior. The peculiarities of the Internet texts include enhanced dialogue, divisibility, external informativity, reduced connectivity and comprehension, pragmatic and mostly formal integrity, conditional completeness, complicated structural, as well as hybrid and high degree of permeability, multimedia, presentation, inclination to speech game and collective authorship, saturation with neologisms, emoticons and abbreviations, fragmentation, non-compliance with linguistic norms, and the functioning of a special language etiquette.
 Conclusions. As a result of empirical research involving 716 respondents from different regions of Ukraine, it was determined that experience has the most significant effect on the understanding process at the reception stage, guiding users' activity and the accuracy of their expectations. Experience contributes to the accuracy of predicting the content of Internet texts by 15,0%. At the stage of interpretation, the adequacy and completeness of Internet texts interpretation with the accumulation of experience is improved by almost 10,0%. However, even experienced users were able to correctly interpret only a quarter of the dominant, while random – only a sixth part. Even less important is the experience and Internet activity at the stage of emotional identification, nor the assessment of comprehension, nor the coherence of emotional attitude is almost independent of the Internet using experience. With the accumulation of experience, users evaluate Internet texts more homogeneously, they are easier to realize their own attitude to the Internet texts, it becomes more consistent, however, they underestimate the complexity of Internet texts more than half of the cases, they are also inclined to share texts on the approval and critical like inexperienced readers.
Syntactic structure of sentences obtained from Constituency Parsing is fundamental information in many Natural Language Processing tasks. However, due to the lack of available resources and the complex linguistic features of Vietnamese, the research into Constituency Parsing has not received enough attention in this language. To the best of our knowledge, the study presented in this paper is one of the first investigations to explore this task in Vietnamese. In this work, we present a Spanbased approach which focuses on representing spans through the use of contextualized pre-trained embeddings to obtain optimal parse trees for Vietnamese sentences. The conducted experiments indicate that our system achieved promising results on the VLSP Vietnamese Treebank dataset by significantly outperforming existing methods. The results of this study support the view that encoding context information into the representation of words is effective in improving the parsing performance of Vietnamese. Consequently, this idea can be generalized to apply to other tasks such as Dependency Parsing or other low-resource languages.
Ellipsis is very common in language. It’s necessary for natural language processing to restore the elided elements in a sentence. However, there’s only a few corpora annotating the ellipsis, which draws back the automatic detection and recovery of the ellipsis. This paper introduces the annotation of ellipsis in Chinese sentences, using a novel graph-based representation Abstract Meaning Representation (AMR), which has a good mechanism to restore the elided elements manually. We annotate 5,000 sentences selected from Chinese TreeBank (CTB). We find that 54.98% of sentences have ellipses. 92% of the ellipses are restored by copying the antecedents’ concepts. and 12.9% of them are the new added concepts. In addition, we find that the elided element is a word or phrase in most cases, but sometimes only the head of a phrase or parts of a phrase, which is rather hard for the automatic recovery of ellipsis.
Hate-inducing language, which has become a recurrent decimal in Nigerian socio-political discourse, is not unconnected to the deep-seated boundaries existing amongst different ethnic groups in Nigeria. Linguistic studies on hate language in Nigeria have largely utilised pragmatic and critical discourse analytical tools in identifying the illocutions and ideologies involved but hardly paid attention to the metalinguistic forms deployed in hate speeches. Therefore, the present study, aside adding to the research line of the Natural Semantic Metalanguage (NSM)—which has unduly focused on language typology, explores the metalinguistic evaluators that index hate speech in Nigeria, and relate them to specific pragmatic strategies through which hate speech producers’ intentions are communicated. To achieve this, three full manuscripts of hate speech made by three groups (i.e. Arewa Youth Consultative Forum, Youths of Oduduwa Republic, and Biafra Nation Youth League) from three (northern, western, and eastern, respectively) regions of Nigeria are purposively sampled from Google directories and Radio Biafra archives, subjected to descriptive and quantitative analysis, with insights from the NSM theory and aspects of pragmatic acts. Two categories of metalinguistic evaluators were identified, positive (GOOD) and negative (BAD) evaluators; and these are associated with three pragmatic strategies; namely, blunt condemnation, unshielded exposition, and appeal to emotion. While the condemning and exposing strategies largely utilise negative evaluators in initiating hate on target groups, the emotion-drawing strategy largely employs positive evaluators in boosting the image of the hate-speech producing group in the eyes of the audience. With these findings, the study takes existing scholarship on violence-inducing language a step forward, especially in providing a pragmatic explanation to the proliferation of hate crimes in Nigeria. It also offers a holistic linguistic database and critical meta-language for the teaching of hate-related language and crime, especially in second-language situations.
Combinatory Categorial Grammars provide a transparent interface between surface syntax and underlying semantic representation. Discourse Representation Theory allows the handling of meaning across sentence boundaries. Based on the foundations of these two theories along with the work of Johan Bos on the Boxer framework for English language, we propose an approach to the task of semantic parsing with Discourse Representation Structure for the French language. By giving an example of discourse analysis on French sentences and experimenting on 4,525 sentences taken from the French Treebank corpus, we demonstrate and evaluate the outcomes of our framework.
This two-phased, sequential mixed-methods study investigates how raters are influenced by different rating scales on a college-level English as a second language (ESL) writing placement test. In Phase I, nine certified raters rated 152 essays using a holistic, profile-based scale; in Phase II, they rated 200 essays using a binary, analytic scale developed based on the holistic scale and 100 essays using both rating scales. Ratings were examined both quantitatively through Rasch modeling and qualitatively via think-aloud protocols and semi-structured interviews. Findings from Phase I revealed that, despite satisfactory internal consistency, the raters demonstrated relatively low rater agreement and individual differences in their use of the holistic scale. Findings from Phase II showed that the binary, analytic scale led to much improvement in rater consensus and rater consistency. Another finding from Phase II suggests that the binary, analytic scale helped the raters deconstruct the holistic scale, reducing their cognitive burden. This study represents a creative use of a binary, analytic scale to guide raters through a holistic rating scale. Implications regarding how a rating scale affects rating behavior and performance are discussed.
In recent years, neural networks have been widely used for language modeling in different tasks of natural language processing. Results show that long short-term memory (LSTM) neural networks are appropriate for language modeling due to their ability to process long sequences. Furthermore, many studies are shown that extra information improve language models (LMs) performance. In this research, we propose parallel structures for incorporating part-of-speech tags into language modeling task using both the unidirectional and bidirectional type of LSTMs. Words and part-of-speech tags are given to the network as parallel inputs. In this way, to concatenate these two paths, two different structures are proposed according to the type of network used in the parallel part. We analyze the efficiency on Penn Treebank (PTB) dataset using perplexity measure. These two proposed structures show improvements in comparison to the baseline models. Not only does the bidirectional LSTM method gain the lowest perplexity, but it also has the lowest training parameters among our proposed methods. The perplexity of proposed structures has reduced 1.5% and %13 for unidirectional and bidirectional LSTMs, respectively.
WordNet is a lexical database for languages, the difference between WordNet and dictionaries in general is that WordNet focuses on the synonyms. The main unit of WordNet is synonym set (synset), synset is a set of one or more words that have the same meaning and certainly can be replaced in certain contexts. Synset is a very important element in implementing WordNet. In this paper, an analysis of the synonym extraction process is carried out by using commutative approach, the data test obtained from the Oxford Paperback Thesaurus by taking 51 word entries. Commutative method has similar characters with synonym set, synonym set can replace each other in certain contexts. The data test extraction process is carried out until the performance measurement evaluation process using F1Score. The system generates synonym sets that matched with the manual extraction, the result of F1Score between the program and Princeton synonym sets are worth 10%.