Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Abstract Two-thirds of adults in the United Kingdom currently suffer from overweight or obesity, making it one of the biggest contributors to health problems. Within the framework of the incentive sensitisation theory, it has been hypothesised that overweight people experience heightened reward anticipation when encountering cues that signal food, such as pictures and smells of food, but that they experience less reward from consuming food compared to normal-weight people. There is, however, little evidence for this prediction. Few studies test both anticipation and consumption in the same study, and even fewer with electroencephalography (EEG). This study sought to address this gap in the literature by measuring scalp activity when overweight and normal-weight people encountered cues signalling the imminent arrival of pleasant and neutral taste stimuli, and when they received these stimuli. The behavioural data showed that there was a smaller difference in valence ratings between the pleasant and neutral taste in the overweight than normal-weight group, in accordance with our hypothesis. However, contrary to our hypothesis, the groups did not differ in their electrophysiological response to taste stimuli. Instead, there was a reduction in N1 amplitude to both taste and picture cues in overweight relative to normal-weight participants. This suggests that reduced attention to cues may be a crucial factor in risk of overweight.
In this thesis, we propose novel approaches for supervised RST-style discourse parsing, as well as the methods for utilizing those discourse structures for the benefit of natural language generation. We demonstrate a significant improvement in discourse parsing accuracy on RST-DT and Instr-DT treebanks by incorporating silver-standard supervision. Furthermore, in line with theoretical and empirical connections between the discourse parsing and coreference resolution tasks, we find the evidence of improvement of discourse parsing accuracy on RST-DT when our proposed discourse parsing system is provided with coreference supervision from a coreference resolver trained on OntoNotes corpus. Finally, in extending our work to natural language generation, we demonstrate that our novel content structuring system utilizing silver-standard discourse structures outperforms text-only systems on our proposed task of elementary discourse unit ordering, a significantly more difficult version of sentence ordering task.
This article discusses issues such as the importance of language learning in the education system of Uzbekistan, the relationship between language and society, nation, the status of the Uzbek language as a state language. In addition, the article highlights the basics of communicative and literary literacy of the document and some problems of the Uzbek legal language (translation, thesaurus). At present, when science is developing, new tasks appear for the education system. In particular, the requirement of the time is not only knowledge of the specialty, but also knowledge of the state language, full mastery of the norms of the modern Uzbek literary language. Consequently, a modern lawyer should not only be a professional with a deep knowledge of the laws and rules of public relations, taking into account the specifics of a renewed society in order to establish stability and law and order in society, but also possess linguistic norms. Ensuring the rule of law in society is one of the most important principles of a democratic state, for which laws must be fair by their nature and understandable for people, that is, the language of legal documents created must be detailed and free. Consequently, language and law are closely related concepts. One of the most pressing issues that need to be resolved today is the compilation of an industry language thesaurus, regulation of legal terminology.
Despite the increasing number of U.S. born Latinos, placing heritage and native speakers in the Spanish curriculum is still a challenge (MacGregor-Mendoza & Moreno, 2020). The present article (a) addresses the unique needs of heritage speakers in the Spanish curriculum; (b) problematizes traditional grammar-based placement exams; and (c) describes a multiple-choice placement exam (free upon request) designed and used at Georgia State University (GSU), a major urban university in the Southeastern U.S. Taking a sociolinguistic approach to the dialectical nature of Spanish, the GSU Spanish Language Program Coordinator developed a placement test based on what students—heritage, native, and non-native—do when asked to perform language tasks. The placement test design is outlined using distinctions of linguistic norms, both local/ regional and general. Reference is made to the ways in which diverse types of Spanish speakers align linguistically with general Spanish. This essay responds to the call for language standardization studies that recognize diglossia within a single named language by examining the role of heteroglossia to challenge monolingual language standardization ideologies (McLelland, 2021). Pedagogical implications for identifying and placing K-16 learners in a meaningful Spanish for Heritage Speakers classroom are discussed.
There is empirical evidence that expected yet not current affect predicts decisions.However, common research designs in affective decision-making show consistent methodological problems (e.g., conceptualization of different emotion concepts; measuring only emotional valence, but not arousal).We developed a gambling task that systematically varied learning experience, average feedback balance and feedback consistency.In Experiment 1 we studied whether predecisional current affect or expected affect predict recurrent gambling responses.Furthermore, we exploratively examined how affective information is represented on a neuronal level in Experiment 2. Expected and current valence and arousal ratings as well as Blood Oxygen Level Dependent (BOLD) responses were analyzed using a within-subject design.We used a generalized mixed effect model to predict gambling responses with the different affect variables.Results suggest a guiding function of expected valence for decisions.In the anticipation period, we found activity in brain areas previously associated with valencegeneral processing (e.g., anterior cingulate cortex, nucleus accumbens, thalamus) mostly independent of contextual factors.These findings are discussed in the context of the idea of a valence-general affective work-space, a goal-directed account of emotions, and the hypothesis that current affect might be used to form expectations of future outcomes.In conclusion, expected valence seems to be the best predictor of recurrent decisions in gambling tasks.
RST-style discourse parsing plays a vital role in many NLP tasks, revealing\nthe underlying semantic/pragmatic structure of potentially complex and diverse\ndocuments. Despite its importance, one of the most prevailing limitations in\nmodern day discourse parsing is the lack of large-scale datasets. To overcome\nthe data sparsity issue, distantly supervised approaches from tasks like\nsentiment analysis and summarization have been recently proposed. Here, we\nextend this line of research by exploiting distant supervision from topic\nsegmentation, which can arguably provide a strong and oftentimes complementary\nsignal for high-level discourse structures. Experiments on two human-annotated\ndiscourse treebanks confirm that our proposal generates accurate tree\nstructures on sentence and paragraph level, consistently outperforming previous\ndistantly supervised models on the sentence-to-document task and occasionally\nreaching even higher scores on the sentence-to-paragraph level.\n
Dependency parsing is a crucial step towards deep language understanding and, therefore, widely demanded by numerous Natural Language Processing applications. In particular, left-to-right and top-down transition-based algorithms that rely on Pointer Networks are among the most accurate approaches for performing dependency parsing. Additionally, it has been observed for the top-down algorithm that Pointer Networks' sequential decoding can be improved by implementing a hierarchical variant, more adequate to model dependency structures. Considering all this, we develop a bottom-up-oriented Hierarchical Pointer Network for the left-to-right parser and propose two novel transition-based alternatives: an approach that parses a sentence in right-to-left order and a variant that does it from the outside in. We empirically test the proposed neural architecture with the different algorithms on a wide variety of languages, outperforming the original approach in practically all of them and setting new state-of-the-art results on the English and Chinese Penn Treebanks for non-contextualized and BERT-based embeddings.
The aim of this paper is to present the utility of the Gorazd: An Old Church Digital Hub for scholars working with Old Romanian and Slavonic texts written on the territory of today’ s Romania. The Gorazd Project was realized during the years 2016–2020 and it includes an Old Church Slavonic Card Index and three Old Church Slavonic lexical databases, among which the largest one is represented by the digitized and updated version of the monumental Lexicon linguæ palæoslovenicæ (vol. I–IV, 1958–1997) composed by the Institute of Slavonic Studies of the Czech Academy of Sciences. As the Gorazd Project uses English as meta-language, its application is not limited to narrowly specialized Slavic philologists, but it is also open for scholars of neighbouring fields. The dictionaries within the Gorazd Digital Hub can serve as a reference tool not just for the oldest attested Slavonic vocabulary and its semantics, but also for the biblical concordance of the Slavonic oldest Bible redaction and the oldest attested Old Church Slavonic morphological forms.
It is proposed to create an information and reference system containing complete and comprehensive information about the scientific and scholarly results achieved by the RAS institutions, scientific departments and researchers in the field of linguistics. The reference system should provide information support for high-quality scientific and methodological guidance from the relevant branch of the Russian Academy of Sciences. The classification of information objects describing scientific and scholarly results includes both traditional forms of writing (publications, reports, dissertations) and the new digital ones (linguistic databases, websites, corpora, accounts, etc.). The preliminary results of creating such a system are described. The functions of the help system are listed. The updated version of the section “Linguistics” of the State rubricator of scientific and technical information is offered.
It is obvious that the ordering distribution of temporal adverbial clauses in advanced Chinese EFL learners writing corpus (ACEFL) and English differs greatly. Based on the theory of dependency grammar, this paper builds two dependency treebanks and uses mean dependency distance (MDD) as an index to measure syntactic difficulty of prepositional and postpositional temporal adverbial clauses by advanced Chinese EFL learners. It is found that: 1) temporal adverbial clauses in Chinese show obvious tendency of preposition, while in English they can be preposed or postposed to the main clauses, but postposition is the dominant order. In contrast, advanced Chinese EFL learners tend to prepose temporal adverbial clauses which is similar to the ordering distribution of their mother tongue; 2) syntactic difficulty of prepositional temporal adverbial clauses in ACEFL is smaller than that of postpositional ones; 3) the main motivation of preposition of temporal adverbial clauses in ACEFL are dual influence of mother tongue and minimization tendency of MDD.
We introduce a method for unsupervised parsing that relies on bootstrapping classifiers to identify if a node dominates a specific span in a sentence. There are two types of classifiers, an inside classifier that acts on a span, and an outside classifier that acts on everything outside of a given span. Through self-training and co-training with the two classifiers, we show that the interplay between them helps improve the accuracy of both, and as a result, effectively parse. A seed bootstrapping technique prepares the data to train these classifiers. Our analyses further validate that such an approach in conjunction with weak supervision using prior branching knowledge of a known language (left/right-branching) and minimal heuristics injects strong inductive bias into the parser, achieving 63.1 F$_1$ on the English (PTB) test set. In addition, we show the effectiveness of our architecture by evaluating on treebanks for Chinese (CTB) and Japanese (KTB) and achieve new state-of-the-art results. Our code and pre-trained models are available at https://github.com/Nickil21/weakly-supervised-parsing.
Recent years have witnessed significant improvement in ASR systems to recognize spoken utterances. However, it is still a challenging task for noisy and out-of-domain data, where substitution and deletion errors are prevalent in the transcribed text. These errors significantly degrade the performance of downstream tasks. In this work, we propose a BERT-style language model, referred to as PhonemeBERT, that learns a joint language model with phoneme sequence and ASR transcript to learn phonetic-aware representations that are robust to ASR errors. We show that PhonemeBERT can be used on downstream tasks using phoneme sequences as additional features, and also in low-resource setup where we only have ASR-transcripts for the downstream tasks with no phoneme information available. We evaluate our approach extensively by generating noisy data for three benchmark datasets - Stanford Sentiment Treebank, TREC and ATIS for sentiment, question and intent classification tasks respectively. The results of the proposed approach beats the state-of-the-art baselines comprehensively on each dataset.
Neural machine translation (NMT) models are typically trained using a softmax cross-entropy loss where the softmax distribution is compared against the gold labels. In low-resource scenarios and NMT models tend to perform poorly because the model training quickly converges to a point where the softmax distribution computed using logits approaches the gold label distribution. Although label smoothing is a well-known solution to address this issue and we further propose to divide the logits by a temperature coefficient greater than one and forcing the softmax distribution to be smoother during training. This makes it harder for the model to quickly over-fit. In our experiments on 11 language pairs in the low-resource Asian Language Treebank dataset and we observed significant improvements in translation quality. Our analysis focuses on finding the right balance of label smoothing and softmax tempering which indicates that they are orthogonal methods. Finally and a study of softmax entropies and gradients reveal the impact of our method on the internal behavior of our NMT models.
In order to achieve deep natural language understanding, syntactic constituent parsing is a vital step, highly demanded by many artificial intelligence systems to process both text and speech. One of the most recent proposals is the use of standard sequence-to-sequence models to perform constituent parsing as a machine translation task, instead of applying task-specific parsers. While they show a competitive performance, these text-to-parse transducers are still lagging behind classic techniques in terms of accuracy, coverage and speed. To close the gap, we here extend the framework of sequence-to-sequence models for constituent parsing, not only by providing a more powerful neural architecture for improving their performance, but also by enlarging their coverage to handle the most complex syntactic phenomena: discontinuous structures. To that end, we design several novel linearizations that can fully produce discontinuities and, for the first time, we test a sequence-to-sequence model on the main discontinuous benchmarks, obtaining competitive results on par with task-specific discontinuous constituent parsers and achieving state-of-the-art scores on the (discontinuous) English Penn Treebank.
The article analyzes transformation of forms of degrees of comparison of adjectives in live television broadcasting. Particular attention is paid to the specific properties of different forms of degrees of comparison of adjectives. To analyze the peculiarities of their use for errors in speech of television journalists, associated with non-compliance with linguistic norms on ways to avoid these errors, to make appropriate recommendations to television journalists. The main method we use is to observe the speech of live TV journalist, we used during the study methods of comparative analysis of comparison of theoretical positions from the work of individual linguists and journalism sat down as well as texts that sounded in the speech of journalists. Our objective is to trace these transformations and develop a certain attitude towards them in our researches of the language of the media and practicing journalists to support positive trends in the development of the broadcasting on TV and give recommendations for overcoming certain negative trends. Improving the live broadcasting of television journalists, in particular the work on deepening the language skills will contribute to the modernization of some trends in the reasonable expediency of the transformation of certain phenomena, modernization of some tendencies concerning the reasonable expedient transformation of separate grammatical phenomena and categories and at braking and in general stopping of processes of transformation of negative unreasonable not expedient. This fully applies primarily to attempts to transform the forms of degrees of comparison of adjectives and this explains importance of the results achieved in these study.
Widespread access to social media ensures that new and emergent coinages are noticed by population masses, attain domains of usage, and strengthen their place within a language. Of such new linguistic constructs, metaphorical neologisms are usually most adopted and frequently used. The present study aims to examine the x + head phrase structure widely used in Turkish and to evaluate its semantic properties. Examples semantically identical to this structure are included in the linguistic database through the Turkish National Corpus. The appearance of such phrases was observed on different social media platforms. Additionally, the opinions of people were sought on the use of this structure in social interactions. Further, the differences between the usage of the mentioned structure on its own and its use within a context were measured via a survey.
We present substructure distribution projection (SubDP), a technique that projects a distribution over structures in one domain to another, by projecting substructure distributions separately. Models for the target domains can be then trained, using the projected distributions as soft silver labels. We evaluate SubDP on zero-shot cross-lingual dependency parsing, taking dependency arcs as substructures: we project the predicted dependency arc distributions in the source language(s) to target language(s), and train a target language parser to fit the resulting distributions. When an English treebank is the only annotation that involves human effort, SubDP achieves better unlabeled attachment score than all prior work on the Universal Dependencies v2.2 (Nivre et al., 2020) test set across eight diverse target languages, as well as the best labeled attachment score on six out of eight languages. In addition, SubDP improves zero-shot cross-lingual dependency parsing with very few (e.g., 50) supervised bitext pairs, across a broader range of target languages.
Controlling the presented forms (or structures) of generated text are as important as controlling the generated contents during neural text generation. It helps to reduce the uncertainty and improve the interpretability of generated text. However, the structures and contents are entangled together and realized simultaneously during text generation, which is challenging for the structure controlling. In this paper, we propose an efficient, straightforward generation framework to control the structure of generated text. A structure-aware transformer (SAT) is proposed to explicitly incorporate multiple types of multi-granularity structure information to guide the text generation with corresponding structure. The structure information is extracted from given sequence template by auxiliary model, and the type of structure for the given template can be learned, represented and imitated. Extensive experiments have been conducted on both Chinese lyrics corpus and English Penn Treebank dataset. Both automatic evaluation metrics and human judgement demonstrate the superior capability of our model in controlling the structure of generated text, and the quality ( like Fluency and Meaningfulness) of the generated text is even better than the state-of-the-arts model.
Text discourse parsing weighs importantly in understanding information flow and argumentative structure in natural language, making it beneficial for downstream tasks. While previous work significantly improves the performance of RST discourse parsing, they are not readily applicable to practical use cases: (1) EDU segmentation is not integrated into most existing tree parsing frameworks, thus it is not straightforward to apply such models on newly-coming data. (2) Most parsers cannot be used in multilingual scenarios, because they are developed only in English. (3) Parsers trained from single-domain treebanks do not generalize well on out-of-domain inputs. In this work, we propose a document-level multilingual RST discourse parsing framework, which conducts EDU segmentation and discourse tree parsing jointly. Moreover, we propose a cross-translation augmentation strategy to enable the framework to support multilingual parsing and improve its domain generality. Experimental results show that our model achieves state-of-the-art performance on document-level multilingual RST parsing in all sub-tasks.
Liking and pleasantness are common concepts in psychological emotion theories, and everyday language related to emotions. Despite obvious similarities between the terms, several empirical and theoretical notions support the idea that pleasantness and liking are cognitively different phenomena, becoming most evident in the context of emotion regulation and art enjoyment. In this study it was investigated whether liking and pleasantness indicate behaviourally measurable differences, not only in the long timespan of emotion regulation, but already within the initial affective responses to visual and auditory stimuli. A cross-modal affective priming protocol was used to assess whether there is a behavioural difference in the response time when providing an affective rating to a liking or pleasantness task. It was hypothesized that the pleasantness task would be faster as it is known to rely on rapid feature detection. Furthermore, an affective priming effect was expected to take place across the sensory modalities and the presentative and non-presentative stimuli. A linear mixed effect analysis indicated a significant priming effect, as well as an interaction effect between the auditory and visual sensory modalities and the affective rating tasks of liking and pleasantness: While liking was rated fastest across modalities, it was significantly faster in vision compared to audition. No significant modality dependent differences between the pleasantness ratings were detected. The results demonstrate that liking and pleasantness rating scales refer to separate processes already within the short time scale of a one to two seconds. Furthermore, the affective priming effect indicates that an affective information transfer takes place across modalities and the types of stimuli applied. Unlike hypothesized, liking rating took place faster across the modalities. This is interpreted to support emotion theoretical notions where liking and disking are crucial properties of emotions perception and homeostatic self-referential information, possibly overriding pleasantness-related feature analysis. Conclusively, the findings provide empirical evidence for a conceptual delineation of common affective processes.
Parsing, i.e., identifying the underlying hierarchical structure of natural language expressions is important for several natural language processing applications. In recent times Machine Learning (ML) approaches have been developed for this study for many languages. Most of the effective techniques require an annotated corpus of the language for training and validation. For the Manipuri language of the Tibeto-Burman family, neither such a corpus nor a grammar framework to automatically analyse and represent the structure of sentences exists yet. This study proposes a Context-Free Grammar (CFG) that provides the framework to represent the structure of Manipuri sentences. This paves the way for parsing Manipuri sentences using CFG-based parsers for various applications and to conveniently build a Treebank for developing ML-based parsers for Manipuri. The rules of the proposed CFG are handcrafted after extensive analysis of the structure of Manipuri sentences. The grammar covers simple, compound, complex and compound-complex sentences. For evaluation, we induce an Earley's parser with the proposed CFG and test it over a collection of sentences that covers the possible varieties of structure. A recognition rate of 83.20% achieved in these experiments indicates the effectiveness of the proposed grammar.
While some heritage languages enjoy large numbers of speakers and vibrant communities, centuries-old and ongoing sociohistorical and sociolinguistic oppression has resulted in the extreme endangerment of many Indigenous languages. To counter this linguistic and cultural loss, a growing number of communities have engaged in language revitalization efforts that are tied to broader objectives of ethnic reclamation and cultural resistance, aiming not only to maintain but also to strengthen what has been lost. Heritage language revitalization is a long-term project that demands change and engagement across many aspects of community life, work that is ripe with tensions and contradictions. This chapter considers three recurrent questions in heritage language revitalization: what efforts should be prioritized in language revitalization, who should take responsibility in revitalizing a language, and how should revitalization efforts navigate the perceived need to establish linguistic norms and standards while concomitantly supporting linguistic diversity. To date, these questions have been described as tensions or problems that reveal conflicting priorities, often the result of historical inequalities, and that frequently hinder language revitalization efforts. Rather than framing these questions as problems, the present chapter considers how communities have responded to these challenges to create new opportunities for collaboration and new approaches that embrace ambiguity and pluralism.
Among the various challenges regarding distance education is the necessity of reducing the student dropout rate. In this sense, the present research aimed to contribute to the design of a lexical database focused on emotions and opinions that can be incorporated into a predictive evasion software. For the database design, we used the Scup tool to collect 150 tweets containing distance education students’ opinions and analyzed them in the light of Martin and White’s Appraisal Framework, along with five resources related to the sentiment Analysis field, which were taken from Liu’s work. In addition, we used the Aulete dictionary to describe the lexical units found in our corpus to better fit them into the analysis categories. Results showed 220 opinion tokens, which were identified and labeled according to their polarity. Moreover, these tokens were included in the domains attitude (judgment and appreciation) and graduation (sharp and strong) from the linguistic framework used. The results also indicated the necessity of another resource to help identify the use of figurative language, slangs, and extralinguistic elements, such as GIFS and emojis.
Recent advancements in language models based on recurrent neural networks and transformers architecture have achieved state-of-the-art results on a wide range of natural language processing tasks such as pos tagging, named entity recognition, and text classification. However, most of these language models are pre-trained in high resource languages like English, German, Spanish. Multi-lingual language models include Indian languages like Hindi, Telugu, Bengali in their training corpus, but they often fail to represent the linguistic features of these languages as they are not the primary language of the study. We introduce HinFlair, which is a language representation model (contextual string embeddings) pre-trained on a large monolingual Hindi corpus. Experiments were conducted on 6 text classification datasets and a Hindi dependency treebank to analyze the performance of these contextualized string embeddings for the Hindi language. Results show that HinFlair outperforms previous state-of-the-art publicly available pre-trained embeddings for downstream tasks like text classification and pos tagging. Also, HinFlair when combined with FastText embeddings outperforms many transformers-based language models trained particularly for the Hindi language.
View of Volume 66, Special Issue, September 2021 The effectiveness of medical evidence is largely dependent on the ability to communicate that evidence to the science-users, mostly patients. Like in many fields of science, also in medicine trust is one of the most important components of doctor-patient interaction. Cultivation of patient trust is, in turn, primarily a linguistic activity, subject to linguistic norms and conventions. Doctor-patient interaction has been at the core of a growing discussion during the past few years, especially in the context of innovations in evidence-based methods and related to the applicability of clinical guidelines derived from those methods. In Italy, this debate resulted in a recent law (n.219/2017), which declares that “the care and trusting relationship between doctor and patient which is based on the informed consent is promoted and enhanced” (art.1) and that “the time of the communication between doctor and patient is a time of care” (art.8). This new kind of perspective on communication between physicians and patients has led to several questions, above all (i) what is the best definition of trust? and (ii) how achieve a trusting relationship? According to a strictly philosophical point of view, it implies how to successfully communicate imperfect evidence and risk to patients who are in a position of epistemic asymmetry with respect to the doctors; it is problematic because it involves a transfer of complex knowledge of risks and uncertainties from experts to laypeople. The paper investigates the difficulties in communicating medical evidence associated with risk and uncertainties of diagnosis and treatment.
In this paper, we present the results of our experiments concerning the zero-shot crosslingual performance of the PERIN sentence-tograph semantic parser. We applied the PTG model trained using the PERIN parser on a 740k-token Czech newspaper corpus to Hungarian. We evaluated the performance of the parser using the official evaluation tool of the MRP 2020 shared task. The gold standard Hungarian annotation was created by manual correction of the output of the parser following the annotation manual of the tectogrammatical level of the Prague Dependency Treebank. An English model trained on a larger one-million-token English newspaper corpus is also available, however, we found that the Czech model performed significantly better on Hungarian input due to the fact that Hungarian is typologically more similar to Czech than to English. We have found that zero-shot transfer of the PTG meaning representation across typologically not-too-distant languages using a neural parser model based on a multilingual contextual language model followed by a manual correction by linguist experts seems to be a viable annotation scenario.
Background: a recurrent linguistic difficulty in the written texts of health professionals lies in the incorrect use of the gerund due to the influence of the English language in Spanish, bad translations, linguistic ignorance and ingrained ―dogmatic― misconceptions. Objective: to base the correct use of the Spanish gerund in scientific writing, as an important and necessary structure of the Spanish language. Methods: a bibliographic review was carried out, 32 texts were consulted ― degree theses, original articles and review articles. After preliminary reading, 20 were selected, taking into account the updating of the bibliography and its relevance to the proposed objective. Results: according to the analysis, the gerund can be correctly used in the most dissimilar situations. The multiplicities of functions granted to this non-personal verbal form by inexperienced in the language have weighed down its functions and have generated fear among those who are unaware of its benefits and possibilities. The authors consider that linguistic improvement in the morphosyntactic knowledge of the language is almost nil in a university context where the priority is the domain of medical sciences, together with the lack of interest of some professionals in the branch for not considering language as the best tool for its intellectual projection and as an inherent aspect of their professional image. Conclusions: it was found that the generality of the consulted authors agree that the gerund can be used in all writing styles, including the scientific one, as long as it is used in correspondence with the linguistic norms.
Constructive interactions through discussion forums allow students to open their horizons and thought processes to acquire more knowledge and develop skills. Thus, discussion forums play an important role in supporting learning. Additionally, the discussion forum provides the content for creating a knowledge repository. It contains discussion threads related to key course topics that are debated by the students. One approach to understanding the student learning experience is through the analysis of the discussion threads. This research proposes the application of discourse analysis and collaborative learning frameworks to discussion forums to gain further insights into the student's learning in a classroom. It is a foray into discourse analysis using in-class discussions. It demonstrates the application of Soller's framework and Penn Discourse Treebank (PDTB) to understand interactions at the discourse and semantic level. It also shows the use of unsupervised automated techniques to diagnose interactions in textual data. In this paper, we present an Integrated Discourse Analysis and Collaborative Learning Skills (IDALS) framework based on in-class discussions. We describe our experiences of applying IDALS framework and evaluating the solution model in a graduate in-class discussion forum. We also highlight the benefits of using visualizations to present the insights to the instructors.
It would be helpful to consider various topic-independent features: syntax, semantics, and discourserelations between text fragments to accurately detect texts containing elements of hatred or enmity.Unfortunately, methods for identifying discourse relations in the texts of social networks are poorlydeveloped. The paper considers the task of classification of discourse relations between two parts ofthe text. The RST Discourse Treebank dataset (LDC2002T07) is used to assess the performance of themethods. Since the size of this dataset is too small for training large language models, the work uses amodel-pre fitting approach. Model pre-fitting is performed on a Reddit user comment dataset. Textsfrom this dataset are labeled automatically. Since automatic labeling is less accurate than manualmarking, we use the multiple-instance learning (MIL) method to train models. A distinctive feature ofmodern language models is the large number of parameters. Using several models at different levels ofsuch a text analyzer requires a lot of resources. Therefore, for the analyzer to work, it is necessary touse high-performance or distributed computing. The use of desktop grid systems can attract andcombine computing resources to solve this type of problem.
Machine learning training methods depend plentifully and intricately on\nhyperparameters, motivating automated strategies for their optimisation. Many\nexisting algorithms restart training for each new hyperparameter choice, at\nconsiderable computational cost. Some hypergradient-based one-pass methods\nexist, but these either cannot be applied to arbitrary optimiser\nhyperparameters (such as learning rates and momenta) or take several times\nlonger to train than their base models. We extend these existing methods to\ndevelop an approximate hypergradient-based hyperparameter optimiser which is\napplicable to any continuous hyperparameter appearing in a differentiable model\nweight update, yet requires only one training episode, with no restarts. We\nalso provide a motivating argument for convergence to the true hypergradient,\nand perform tractable gradient-based optimisation of independent learning rates\nfor each model parameter. Our method performs competitively from varied random\nhyperparameter initialisations on several UCI datasets and Fashion-MNIST (using\na one-layer MLP), Penn Treebank (using an LSTM) and CIFAR-10 (using a\nResNet-18), in time only 2-3x greater than vanilla training.\n
For over thirty years, researchers have developed and analyzed methods for latent tree induction as an approach for unsupervised syntactic parsing. Nonetheless, modern systems still do not perform well enough compared to their supervised counterparts to have any practical use as structural annotation of text. In this work, we present a technique that uses distant supervision in the form of span constraints (i.e. phrase bracketing) to improve performance in unsupervised constituency parsing. Using a relatively small number of span constraints we can substantially improve the output from DIORA, an already competitive unsupervised parsing system. Compared with full parse tree annotation, span constraints can be acquired with minimal effort, such as with a lexicon derived from Wikipedia, to find exact text matches. Our experiments show span constraints based on entities improves constituency parsing on English WSJ Penn Treebank by more than 5 F1. Furthermore, our method extends to any domain where span constraints are easily attainable, and as a case study we demonstrate its effectiveness by parsing biomedical text from the CRAFT dataset.
Machine learning training methods depend plentifully and intricately on hyperparameters, motivating automated strategies for their optimisation. Many existing algorithms restart training for each new hyperparameter choice, at considerable computational cost. Some hypergradient-based one-pass methods exist, but these either cannot be applied to arbitrary optimiser hyperparameters (such as learning rates and momenta) or take several times longer to train than their base models. We extend these existing methods to develop an approximate hypergradient-based hyperparameter optimiser which is applicable to any continuous hyperparameter appearing in a differentiable model weight update, yet requires only one training episode, with no restarts. We also provide a motivating argument for convergence to the true hypergradient, and perform tractable gradient-based optimisation of independent learning rates for each model parameter. Our method performs competitively from varied random hyperparameter initialisations on several UCI datasets and Fashion-MNIST (using a one-layer MLP), Penn Treebank (using an LSTM) and CIFAR-10 (using a ResNet-18), in time only 2-3x greater than vanilla training.
Introduction Studies on fear conditioning have made important contributions to the understanding of affective learning mechanisms as well as its applications (e.g., anxiety disorders, post-traumatic stress disorder). However, central mechanisms of sleep related consolidation of fear memory in humans have been almost neglected by previous studies. Objectives In the current study we aimed to test effects of sleep and a period wakefulness on fear conditioned responses. Methods In our experiment in a group 18 healthy volunteers event-related brain potentials (ERP), heart rate variability (HRV) and behavioral responses were recorded during a fear conditioning procedure presented twice, before daytime sleep (2h) or control intervention (a period of wakefulness) and after. The conditioning procedure involved pairing of a neutral tone (CS+) with a highly unpleasant sound (UCS+). Results Differential conditioning manifested itself in the contingent negative variance (CNV)-like slow ERP component. Both period of sleep and wakefulness resulted in an increased amplitude of the CNV to CS+. But we did not find an interaction effect of Time (Pre-Post) by Intervention (Sleep-Wake), suggesting that sleep did not affect the conditioned response differently as compared to a period of wakefulness. An apparent increase in HRV after a period of wakefulness did not affect fear conditioned responses (CNV and valence ratings). Conclusions To summarize, the data indicate that fear memories are consolidated with the course of time with no beneficial effect of sleep; relearning of fear causes stronger differential responses as measured by slow wave amplitude but not behavior; increase of HRV does not affect fear learning.
<p>Parsing, i.e., identifying the underlying hierarchical structure of natural language expressions is important for several natural language processing applications. In recent times Machine Learning (ML) approaches have been developed for this study for many languages. Most of the effective techniques require an annotated corpus of the language for training and validation. For the Manipuri language of the Tibeto-Burman family, neither such a corpus nor a grammar framework to automatically analyse and represent the structure of sentences exists yet. This study proposes a context-Free Grammar (CFG) that provides the framework to represent the structure of Manipuri sentences. This paves the way for parsing Manipuri sentences using CFG-based parsers for various applications and to conveniently build a Treebank for developing ML-based parsers for Manipuri. The rules of the proposed CFG are handcrafted after extensive analysis of the structure of Manipuri sentences. The grammar covers simple, compound, complex and compound-complex sentences. For evaluation, we induce an Earley&rsquo;s parser with the proposed CFG and test it over a collection of sentences that covers the possible varieties of structure. A recognition rate of 83.20% achieved in these experiments indicates the effectiveness of the proposed grammar.</p>
The linguistic atlas is a research of guebni based on geographic the language used, Arab researchers were interested in the paper linguistic atlases that appeared in France and Germany, which were carried out by Westem researchers such as George Winker, Gillirion and Jacob, Karl Gabberg, and Hans Kewarth, they explained the methods of research and ignored the results of linguistic research of ancient arabs in the collection of language and extrapolation, and identify the tribes that depend on it, but technological advances and the advent of computers have greatly helped the work of the digital linguistic atlas, which is distinct from the paper-based atlantic with a linguistic database, this technology also developed the method of designing linguistic atlases and accomplishing them in terms of time, effort, quality and quality, this type appeared in the early eighties in the west this is done by storing the linguistic atlas of the corsica, and then compiling the linguistic atlas in the basque region by computer.
This paper contributes to the thread of research on the learnability of different dependency annotation schemes: one ('semantic') favouring content words as heads of dependency relations and the other ('syntactic') favouring syntactic heads. Several studies have lent support to the idea that choosing syntactic criteria for assigning heads in dependency trees improves the performance of dependency parsers. This may be explained by postulating that syntactic approaches are generally more learnable. In this study, we test this hypothesis by comparing the performance of five parsing systems (both transition-and graph-based) on a selection of 21 treebanks, each in a 'semantic' variant, represented by standard UD (Universal Dependencies), and a 'syntactic' variant, represented by SUD (Surface-syntactic Universal Dependencies): unlike previously reported experiments, which considered learnability of 'semantic' and 'syntactic' annotations of particular constructions in vitro, the experiments reported here consider whole annotation schemes in vivo. Additionally, we compare these annotation schemes using a range of quantitative syntactic properties, which may also reflect their learnability. The results of the experiments show that SUD tends to be more learnable than UD, but the advantage of one or the other scheme depends on the parser and the corpus in question.
The high memory consumption and computational costs of Recurrent neural network language models (RNNLMs) limit their wider application on resource constrained devices. In recent years, neural network quantization techniques that are capable of producing extremely low-bit compression, for example, binarized RNNLMs, are gaining increasing research interests. Directly training of quantized neural networks is difficult. By formulating quantized RNNLMs training as an optimization problem, this paper presents a novel method to train quantized RNNLMs from scratch using alternating direction methods of multipliers (ADMM). This method can also flexibly adjust the trade-off between the compression rate and model performance using tied low-bit quantization tables. Experiments on two tasks: Penn Treebank (PTB), and Switchboard (SWBD) suggest the proposed ADMM quantization achieved a model size compression factor of up to 31 times over the full precision baseline RNNLMs. Faster convergence of 5 times in model training over the baseline binarized RNNLM quantization was also obtained. Index Terms: Language models, Recurrent neural networks, Quantization, Alternating direction methods of multipliers.
Abstract Background: Nowadays, the mobile app market becomes rapidly increased in world wide. The mobile app marketers have smart enough to understand the requirements and demands of customers and perform their aspirations. They delight them. It provides growth, profitability, and creativity with lot of inventions. The main aim of this research is to analyze the customer interest and preferences of mobile service providers. Methodology: This paper proposed the clustering model named as Hierarchical Flexi-Ensemble Clustering (HFEC). It provides the final result with robustness and improved quality. Before clustering, the unwanted features are removed by using the Genetic Algorithm based on the Collective Materials (GACM) technique. The customer preferences are analyzes with the clustering of mobile usage patterns. Results: The analysis determined that the app usage pattern based on the most frequent word, rating category, rating character count, rating word count and content-based rating in the google play store app dataset. Finally, the results are compared with the existing methods to analyze the superior performance of proposed method. The comparison analysis is estimated based on the based on the average hit rate at different cache sizes. Conclusion: The work is concluded with the app pattern prediction in the form of clustering for app marketing service. From the marketing side, they can analyze the customer preferences and satisfaction.
In this paper, we address the representation of coordinate constructions in\nEnhanced Universal Dependencies (UD), where relevant dependency links are\npropagated from conjunction heads to other conjuncts. English treebanks for\nenhanced UD have been created from gold basic dependencies using a heuristic\nrule-based converter, which propagates only core arguments. With the aim of\ndetermining which set of links should be propagated from a semantic\nperspective, we create a large-scale dataset of manually edited syntax graphs.\nWe identify several systematic errors in the original data, and propose to also\npropagate adjuncts. We observe high inter-annotator agreement for this semantic\nannotation task. Using our new manually verified dataset, we perform the first\nprincipled comparison of rule-based and (partially novel) machine-learning\nbased methods for conjunction propagation for English. We show that learning\npropagation rules is more effective than hand-designing heuristic rules. When\nusing automatic parses, our neural graph-parser based edge predictor\noutperforms the currently predominant pipelinesusing a basic-layer tree parser\nplus converters.\n
Purpose and tasks. The purpose is to actualize the linguistic heritage of S. Karavanskyi as a basis for further prescriptive linguistic research. Among the tasks is the analysis of spelling and lexicographic codification in the works of a linguist. The object of our study is the linguistic heritage of Sviatoslav Karavanskyi, who after more than 30 years of Moscow-Stalin concentration camps and 37 years of American emigration carried, preserved and motivated the specific linguistic norm of the constantly destroyed Ukrainian language and its native speakers. The subject of our research is spelling and lexicographic codification of the first third of the XX-XXI century in the works of S. Karavanskyi. When processing the material, we use the analytical and descriptive method. Conclusions and prospects of the study. Spelling issues in the works of S. Karavanskyi have a substantiated ideological basis, which is to reflect the spelling of specific rather than assimilative (“destructive”) features caused by the occupation and totalitarian regime of the 30-80s of the XX century. Spelling assimilation and the necessity to remove it is to change the phonetic-morphological and syntactic structure of the Ukrainian language, in particular phonetic, morphological, word-formation and syntactic changes. The lexicographic codification of the linguist is evidenced by his two fundamental works: “Practical Dictionary of Synonyms of the Ukrainian Language” and “RussianUkrainian Dictionary of Complex Vocabulary”. The main methodological basis for compiling these dictionaries is the specificity of Ukrainian vocabulary in its resistance to codification in dictionaries of “pseudo-language” imposed on Ukrainians during the ethnocide policy and exposing Soviet lexicography as the main “tool of Ukrainian linguicide”. Among the prospects of our study is a holistic linguistic and political portrait of a linguist and socio-political figure.
This research is aimed to describe the language attitude of the people of Mandar, a migrant community in Desa Baharu Utara, Kotabaru Regency. The community group chosen as the object of the research is the young generation (Generasi Muda or GM) of Mandar. Therefore, the respondents are 40 people in various age groups consisting of children, adolescents, and adults with an age range of 6-45 years. Data collection of language attitudes was carried out using a questionnaire which was supported by field observations at the research location. It was found on the research location that GM is more proficient in Banjarese Language (Bahasa Banjar or BB) than Mandar (Bahasa Mandar or BM). This is based on the reality that BB is a local language with high prestige. On the other hand, BB has a strategic role as a lingua franca, which is the language of communication between ethnic groups in the Kotabaru area. Meanwhile, BM, which is the language of minority migrants from West Sulawesi, tends to be pushed by BB's domination because it has lost its prestige. As a result, BM experiences a shift from time to time which is feared to lead to extinction. The shift occurs at various linguistic levels, both phonemes, morphemes, and lexicon. The results of field observations indicate that the older generation (Generasi Tua or GT) has a more positive attitude towards BM than the GM of Mandar. The language attitudes include 1) pride in using BM, 2) loyalty to BM related to the level of frequency of using BM, and 3) awareness of BM norms related to linguistic norms and social norms related to BM usage situations and domains of use BM.
Labeling data can be an expensive task as it is usually performed manually by\ndomain experts. This is cumbersome for deep learning, as it is dependent on\nlarge labeled datasets. Active learning (AL) is a paradigm that aims to reduce\nlabeling effort by only using the data which the used model deems most\ninformative. Little research has been done on AL in a text classification\nsetting and next to none has involved the more recent, state-of-the-art Natural\nLanguage Processing (NLP) models. Here, we present an empirical study that\ncompares different uncertainty-based algorithms with BERT$_{base}$ as the used\nclassifier. We evaluate the algorithms on two NLP classification datasets:\nStanford Sentiment Treebank and KvK-Frontpages. Additionally, we explore\nheuristics that aim to solve presupposed problems of uncertainty-based AL;\nnamely, that it is unscalable and that it is prone to selecting outliers.\nFurthermore, we explore the influence of the query-pool size on the performance\nof AL. Whereas it was found that the proposed heuristics for AL did not improve\nperformance of AL; our results show that using uncertainty-based AL with\nBERT$_{base}$ outperforms random sampling of data. This difference in\nperformance can decrease as the query-pool size gets larger.\n
The aim of this paper is to offer an insight into the semantic roles of adverbials. The approach is mainly construed around the theory of adverb semantics propounded by Quirk, Greenbaum, Leech, and Svartvik (1985) – grammatical functions and the realisation of semantic roles. The theoretical approach is complemented by a practical analysis of adverbial phrases occurring in social interactions (as well as script-based stage directions) from the TV series “Friends”. The main method used is corpus analysis; in addition, a semi-automated identification of adverbs was performed using both quantitative and qualitative analyses. The tools used were ConcApp software, as well as electronic dictionaries and lexical databases. A quantitative and qualitative analysis of -ly adverbials in the script was carried out to establish certain patterns of adverb occurrence in social interaction. The results reveal a large proportion of subjuncts, in particular emphasisers, intensifier subjuncts and downtoners (approximator) (in Greenbaum et al.’s taxonomy), or, in other taxonomies, speaker-oriented (Jackendoff 1972) / sentence adverbs (Swan 1988) / stance adverbs – attitude and epistemic (Biber et al. 1999). A second important finding is that the –ly adverbs used in this sitcom display high polysemy, including some novel semantic uses peculiar to present-day US English.
In this paper, we address the representation of coordinate constructions in Enhanced Universal Dependencies (UD), where relevant dependency links are propagated from conjunction heads to other conjuncts. English treebanks for enhanced UD have been created from gold basic dependencies using a heuristic rule-based converter, which propagates only core arguments. With the aim of determining which set of links should be propagated from a semantic perspective, we create a large-scale dataset of manually edited syntax graphs. We identify several systematic errors in the original data, and propose to also propagate adjuncts. We observe high inter-annotator agreement for this semantic annotation task. Using our new manually verified dataset, we perform the first principled comparison of rule-based and (partially novel) machine-learning based methods for conjunction propagation for English. We show that learning propagation rules is more effective than hand-designing heuristic rules. When using automatic parses, our neural graph-parser based edge predictor outperforms the currently predominant pipelines using a basic-layer tree parser plus converters.
Cloud-based enterprise search services (e.g., AWS Kendra) have been entrancing big data owners by offering convenient and real-time search solutions to them. However, the problem is that individuals and organizations possessing confidential big data are hesitant to embrace such services due to valid data privacy concerns. In addition, to offer an intelligent search, these services access the user's search history that further jeopardizes his/her privacy. To overcome the privacy problem, the main idea of this research is to separate the intelligence aspect of the search from its pattern matching aspect. According to this idea, the search intelligence is provided by an on-premises edge tier and the shared cloud tier only serves as an exhaustive pattern matching search utility. We propose Smartness at Edge (SAED mechanism that offers intelligence in the form of semantic and personalized search at the edge tier while maintaining privacy of the search on the cloud tier. At the edge tier, SAED uses a knowledge-based lexical database to expand the query and cover its semantics. SAED personalizes the search via an RNN model that can learn the user's interest. A word embedding model is used to retrieve documents based on their semantic relevance to the search query. SAED is generic and can be plugged into existing enterprise search systems and enable them to offer intelligent and privacy-preserving search without enforcing any change on them. Evaluation results on two enterprise search systems under real settings and verified by human users demonstrate that SAED can improve the relevancy of the retrieved results by on average ≈24% for plain-text and ≈75% for encrypted generic datasets.
Lexical substitution is the task of generating meaningful substitutes for a word in a given textual context. Contextual word embedding models have achieved state-of-the-art results in the lexical substitution task by relying on contextual information extracted from the replaced word within the sentence. However, such models do not take into account structured knowledge that exists in external lexical databases. We introduce LexSubCon, an end-to-end lexical substitution framework based on contextual embedding models that can identify highly accurate substitute candidates. This is achieved by combining contextual information with knowledge from structured lexical resources. Our approach involves: (i) introducing a novel mix-up embedding strategy in the creation of the input embedding of the target word through linearly interpolating the pair of the target input embedding and the average embedding of its probable synonyms; (ii) considering the similarity of the sentence-definition embeddings of the target word and its proposed candidates; and, (iii) calculating the effect of each substitution in the semantics of the sentence through a fine-tuned sentence similarity model. Our experiments show that LexSubCon outperforms previous state-of-the-art methods on LS07 and CoInCo benchmark datasets that are widely used for lexical substitution tasks.
Recent impressive improvements in NLP, largely based on the success of contextual neural language models, have been mostly demonstrated on at most a couple dozen high- resource languages. Building language mod- els and, more generally, NLP systems for non- standardized and low-resource languages remains a challenging task. In this work, we fo- cus on North-African colloquial dialectal Arabic written using an extension of the Latin script, called NArabizi, found mostly on social media and messaging communication. In this low-resource scenario with data display- ing a high level of variability, we compare the downstream performance of a character-based language model on part-of-speech tagging and dependency parsing to that of monolingual and multilingual models. We show that a character-based model trained on only 99k sentences of NArabizi and fined-tuned on a small treebank of this language leads to performance close to those obtained with the same architecture pre- trained on large multilingual and monolingual models. Confirming these results a on much larger data set of noisy French user-generated content, we argue that such character-based language models can be an asset for NLP in low-resource and high language variability settings.
Time perception is not veridical, but, rather, it is susceptible to environmental context, like the intrinsic dynamics of moving stimuli. The direction of motion has been reported to affect time perception such that the movement of objects toward an observer is perceived as longer in duration than that of objects away from the observer. This looming-motion-induced time dilation has been explained in terms of an arousal-based or an attentional mechanism (or a combination of both). The current study was interested in which of these two explanations represents a more viable mechanism. With this aim, we investigated how the looming/receding temporal asymmetry is modulated by the emotional contents of stimuli. In two experiments, participants were shown face images expressing three emotions (angry, happy, and neutral) for one of seven target durations (400-1000ms) and performed a temporal bisection task by judging each presentation duration as “short” or “long”. In Experiment 1, the face images were shown in a constant-sized, stationary position. In Experiment 2, the images were expanding (looming) or contracting (receding) in size. In Experiment 1, we found no influence of facial emotion in perceived duration. In Experiment 2, however, looming stimuli were perceived as longer in duration than receding ones, replicating previous findings of the looming-induced time dilation using naturalistic human-face stimuli. More importantly, in Experiment 2 we found an interaction effect between arousal rating of faces and motion direction: The looming/receding asymmetry was pronounced when the arousal of the presented images was rated low, but this asymmetry diminished when arousal was high. These results suggest that (1) affective characteristics of looming stimuli can modulate temporal processing and more specifically, (2) the looming/receding asymmetry is reduced when arousing facial expressions enhance attentional engagement to receding stimuli, supporting the attentional mechanism of the looming-induced time dilation.
The purpose of this study is to examine the orthographic and phonological characteristics of the Yeongsan Sillok(the biography of Yeongsan), published in Jeollabuk-do in the early 20th century. The author of this book is considered to be Jang Bong-seon, an educator from Jeongeup city in Jeollabuk-do. Accordingly, it is expected that this book contains the orthographic characteristics and attitudes toward the language of young intellectuals in Jeollabuk-do in the early 20th century. In Chapter 3, we looked at the orthographic characteristics of this book. The writing characteristics of this book largely follow the characteristics of the 19th century Jeollabuk-do dialect based on the tradition of modern Korean. However, a transitional characteristic of the language transforming into present-day Korean was also present. Although only a few examples have been confirmed, the writing of double consonant letters for tense consonant are gradually similar to the notation method of modern Korean. This can be understood as a dissolution process. At the same time, with the exception of some circumstances of verbs, the tendency to split consonants is widely confirmed, and the modern Korean notation for the /ㄹㄹ/ chain (ㄹㄴ, ​​ㄹㅇ) is gradually changing to ㄹㄹ. Above all, the fact that the notation of ․ or diphthong after sibilants no longer appears in this book is a characteristic feature that differs from data from the Jeollabuk-do region of the same period. This writing trend seems to be related to a set of linguistic norms compiled in the first half of the 20th century. Recalling that the author of this book established a private school in the 1920s and 1930s and devoted himself to educational activities, this assumption is somewhat probable. In Chapter 4, we looked at the phonological characteristics of the Yeongsan Sillok(the biography of Yeongsan). Front-vowelization was very active inside the morpheme, but at the morpheme boundary, it appeared only in the environment behind c. The simple vowelization of jə>e is confirmed throughout the interior and boundary of the morpheme, and it must have been a productive phonological phenomenon in the Jeollabuk-do dialect in the early 20th century, as hypercorrection types also appeared. Regarding the alternation of the ending ‘-a/ə’, when the stem vowel is ‘ø’, there is a high tendency to combine these to ‘-ə’. This is different from the 19th century and modern Jeollabuk-do dialects. In the case of umlauts, only very limited examples were shown. And although t-palatalization is quite actively realized, only a few examples of k-palatalization were shown. Through this realization of phonological phenomena, we were able to confirm whether the young intellectuals in the Jeollabuk-do region in the early 20th century had linguistic attitudes toward the Jeollabuk-do dialect. In this book, the typical phonological phenomenon of the Jeollabuk-do dialect was confirmed only to a very limited extent due to its negative evaluation by the author.