Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Part of speech tagging is a fundamental NLP task often regarded as solved for high-resource languages such as English. Current state-of-the-art models have achieved high accuracy, especially on the news domain. However, when these models are applied to other corpora with different genres, and especially user-generated data from the Web, we see substantial drops in performance. In this work, we study how a state-of-the-art tagging model trained on different genres performs on Web content from unfiltered Reddit forum discussions. More specifically, we use data from multiple sources: OntoNotes, a large benchmark corpus with 'well-edited' text, the English Web Treebank with 5 Web genres, and GUM, with 7 further genres other than Reddit. We report the results when training on different splits of the data, tested on Reddit. Our results show that even small amounts of in-domain data can outperform the contribution of data an order of magnitude larger coming from other Web domains. To make progress on out-of-domain tagging, we also evaluate an ensemble approach using multiple single-genre taggers as input features to a meta-classifier. We present state of the art performance on tagging Reddit data, as well as error analysis of the results of these models, and offer a typology of the most common error types among them, broken down by training corpus.
Social interactions enhance human memories, but little is known about how the neural mechanisms underlying episodic memories are modulated by rewarding outcomes in social interactions. To investigate this, fMRI data were recorded while healthy young adults encoded unfamiliar faces in either a competition or a control task. In the competition task, participants encoded opponents' faces in the rock-paper-scissors game, where trial-by-trial outcomes of Win, Draw, and Lose for participants were shown by facial expressions of opponents (Angry, Neutral, and Happy). In the control task, participants encoded faces by assessing facial expressions. After encoding, participants recognized faces previously learned. Behavioral data showed that emotional valence for opponents' Angry faces as the Win outcome was rated positively in the competition task, whereas the rating for Angry faces was rated negatively in the control task, and that Angry faces were remembered more accurately than Neutral or Happy faces in both tasks. fMRI data showed that activation in the medial orbitofrontal cortex (mOFC) paralleled the pattern of valence ratings, with greater activation for the Win than Draw or Lose conditions of the competition task, and the Angry condition of the control task. Moreover, functional connectivity between the mOFC and hippocampus was increased in Win compared to Angry, and the mOFC-hippocampus functional connectivity predicted individual differences in subsequent memory performance only in Win of the competition task, but not in any other conditions of the two tasks. These results demonstrate that the memory enhancement by context-dependent social rewards involves interactions between reward- and memory-related regions.
To investigate issues that arise in the process of developing a Universal Dependency (UD) treebank for Korean and Japanese, we begin by addressing the typological characteristics of Korean and Japanese. Both Korean and Japanese are agglutinative and head-final languages. And the principle of word segmentation for both languages is different from English, which makes it difficult to apply UD guidelines. Following the typological characteristics of the two languages and the issue of UD application, we review the application of UPOS and DEPREL schemes to the two languages. The annotation principles for AUX, ADJ, DET, ADP and PART are discussed for the UPOS scheme, and the annotation principles for case, aux, iobj, and obl are discussed for the DEPREL scheme.
Self-relevant functional abnormalities and identity disorders constitute the core psychopathological components in borderline personality disorder (BPD). Evidence suggests that appraising the relevance of environmental information to the self may be altered in BPD. However, only a few studies have examined self-relevance (SR) in BPD, and the neural correlates of SR processing has not yet been investigated in this patient group. The current study sought to evaluate brain activation differences between female patients with BPD and healthy controls during SR processing. A task-based fMRI paradigm was applied to evaluate SR processing in 23 female patients with BPD and 23 matched healthy controls. Participants were presented with a set of short sentences and were instructed to rate the stimuli. The differences in fMRI signals between SR rating (task of interest) and valence rating (control task) were examined. During SR rating, participants showed elevated activations of the cortical midline structures (CMS), known to be involved in the processing of self-related stimuli. Furthermore, we observed an elevated activation of the supplementary motor area (SMA) and the regions belonging to the mirror neuron system (MNS). Using whole-brain, seed-based connectivity analysis on the task-based fMRI data, we studied connectivity of networks anchored to the main CMS regions. We found a discrepancy in the connectivity pattern between patients and controls regarding connectivity of the CMS regions with the basal ganglia-thalamus complex. These observations have two main implications: First, they confirm the involvement of the CMS in SR evaluations of our stimuli and add evidence about the involvement of an extended network including the MNS and the SMA in this task. Second, the functional connectivity profile observed in BPD provides evidence for an altered functional interplay between the CMS and the brain regions involved in salience detection and reward evaluation, including the basal ganglia and the thalamus.
The objective of this work is to build an Indonesian morphological analyzer named Aksara that conforms to the Universal Dependencies (UD), especially UD v2. Many works had developed Indonesian morphological analyzer, but as far as we know none conforms to the UD annotation guidelines. In building Aksara we use the same approach with MorphInd, another Indonesian morphological analyzer, that uses finite state compiler named Foma. Aksara has capability to perform four tasks: 1) word segmentation, 2) lemmatization, 3) POS tagging, and 4) morphological features analysis. To evaluate the quality of this tool, we used an Indonesian dependency treebank that conforms to UD v2 as the gold standard. We also compare the performance measures of Aksara with MorphInd, by mapping MorphInd output to CoNNL-U format. The experiment results show that for all the four tasks Aksara outperforms MorphInd. For word segmentation task, Aksara has accuracy of 96.9%, for lemmatization with case-sensitive it has accuracy of 94.83%, for POS tagging it has F1-score of 88.2% and finally for morphological features analysis, among 18 feature-value tags already implemented, nine tags already have F1-score more than 80%.
The LiLa: Linking Latin project aims to build a Knowledge Base of language resources for the study of Latin (corpora, digital lexicons, natural-language-processing tools), based on the Linked Open Data paradigm. In this paper, we discuss the goals and motivation of the project. In particular, we focus on the role played by the lemma as a hub node that holds together the network of linguistic information. The architecture of LiLa is therefore based on lemmas and their morphological properties. The paper illustrates the strategies used to build a collection of Latin lemmas and the first experiments to link them to a set of textual resources (Latin treebanks).
In this paper, we explore self-distillation as a means to improve statistical dependency parsing models for Dutch and German over purely supervised training. Self-distillation (Furlanello et al. 2018) trains a new student model on the output of an existing (weaker) teacher model. In contrast to most previous work on self-distillation, we perform distillation using a large, unannotated corpus. We show that in dependency parsing as sequence labeling (Spoustov´a and Spousta 2010, Strzyz et al. 2019), self-distillation plus finetuning provides large improvements over models that use supervised training. We carry out experiments on the German T¨uBa-D/Z universal dependency (UD) treebank (C¸ ¨oltekin et al. 2017) and the UD conversion of the Dutch Lassy Small treebank (Bouma and van Noord 2017). We find that self-distillation improves German parsing accuracy of a bidirectional LSTM parser from 92.23 to 94.33 Labeled Attachment Score (LAS). Similarly, on Dutch we see improvement from 89.89 to 91.84 LAS.
Hoarding disorder has become an official disorder in the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5). Hoarding disorder affects approximately 1.5% to 5% of the general population, and there is no known literature that has examined the prevalence of hoarding disorder among homeless populations or those living in supported housing, although hoarding problems can jeopardize their housing situation. This study used the Clutter Image Rating to estimate the prevalence of possible hoarding behavior among 660 adults living in supported housing. The results indicate that 18.5% of supported housing residents had hoarding behavior, which is more than three times the prevalence reported in the general population. These results suggest that hoarding behavior and possibly hoarding disorder may be more prevalent among those with histories of homelessness and housing instability, which may be of concern because it may affect both housing and health statuses.
We use Universal Dependencies treebanks to test whether a well-known typological trade-off between word order freedom and richness of morphological marking of core arguments holds within individual languages. Using Russian and German treebank data, we show that the following phenomenon (sometimes dubbed word order freezing) does occur: those sentences where core arguments cannot be distinguished by morphological means (due to case syncretism or other kinds of ambiguity) have more rigid order of subject, verb and object than those where unambiguous morphological marking is present. In ambiguous clauses, word order is more often equal to the one which is default or dominant (most frequent) in the language. While Russian and German differ with respect to how exactly they mark core arguments, the effect of morphological ambiguity is significant in both languages. It is, however, small, suggesting that languages do adapt to the evolutionary pressure on communicative efficiency and avoidance of redundancy, but that the pressure is weak in this particular respect.
In this paper, we present our submission for subtask A of the Common Sense Validation and Explanation (ComVE) shared task. We examine the ability of large-scale pre-trained language models to distinguish commonsense from non-commonsense statements. We also explore the utility of external resources that aim to supplement the world knowledge inherent in such language models, including commonsense knowledge graph embedding models, word concreteness ratings, and text-to-image generation models. We find that such resources provide insignificant gains to the performance of fine-tuned language models. We also provide a qualitative analysis of the limitations of the language model fine-tuned to this task.
The urban sound environment is one of the layers that characterizes a city, and several methodologies are used for its assessment, including the soundwalk approach. However, this approach has been tested mainly with adults. In the work presented here, the aim is to investigate a soundwalk methodology for children, analyzing the sound environment of five different sites of Gothenburg, Sweden, from children's view-point, giving them the opportunity to take action as an active part of society. Both individual assessment of the sound environment and acoustic data were collected. The findings suggested that among significant results, children tended to rank the sound environment as slightly better when lower levels of background noise were present ( L A 90 ). Moreover, traffic dominance ratings appeared as the best predictor among the studied sound sources: when traffic dominated as a sound source, the children rated the sound environment as less good. Additionally, traffic volume appeared as a plausible predictor for sound environment quality judgments, since the higher the traffic volume, the lower the quality of the sound environment. The incorporation of children into urban sound environment research may be able to generate new results in terms of children's understanding of their sound environment. Moreover, sound environment policies can be developed from and for children.
Sentiment analysis, especially for long documents, plausibly requires methods capturing complex linguistics structures. To accommodate this, we propose a novel framework to exploit task-related discourse for the task of sentiment analysis. More specifically, we are combining the largescale, sentiment-dependent MEGA-DT treebank with a novel neural architecture for sentiment prediction, based on a hybrid TreeLSTM hierarchical attention model. Experiments show that our framework using sentiment-related discourse augmentations for sentiment prediction enhances the overall performance for long documents, even beyond previous approaches using well-established discourse parsers trained on human annotated data. We show that a simple ensemble approach can further enhance performance by selectively using discourse, depending on the document length.
Arabic dependency parsers perform poorly compared to parsers of other languages. There is little research on improving the performance of Arabic parsers. However, recent research has shown slight improvements in the performance of dependency parsers by utilizing the lexical level of a dependency treebank. To our knowledge, no previous study has studied the effect of utilizing the syntactic level. In this study, we empirically investigated the impact of varying the set of dependency relations on the performance of Arabic dependency parsers. The results were compared to those of previous studies, and showed that having an appropriate set of dependency relations could improve the performance of an Arabic dependency parser.
Summary In this work a novel framework for modeling role and task allocation in Cooperative Heterogeneous Multi‐Robot Systems (CHMRSs) is presented. This framework encodes a CHMRS as a set of multidimensional relational structures (MDRSs). This set of structure defines collaborative tasks through both temporal and spatial relations between processes of heterogeneous robots. These relations are enriched with tensors which allow for geometrical reasoning about collaborative tasks. A learning schema is also proposed in order to derive the components of each MDRS. According to this schema, the components are learnt from data reporting the situated history of the processes executed by the team of robots. Data are organized as a multirobot collaboration treebank (MRCT) in order to support learning. Moreover, a generative approach, based on a probabilistic model, is combined together with nonnegative tensor decomposition (NTD) for both building the tensors and estimating latent knowledge. Preliminary evaluation of the performance of this framework is performed in simulation with three heterogeneous robots, namely, two Unmanned Ground Vehicles (UGVs) and one Unmanned Aerial Vehicle (UAV).
In this paper we present a method for identifying and analyzing adnominal possessive constructions in 66 Universal Dependencies treebanks. We classify adpossessive constructions in terms of their morphological type (locus of marking) and present a workflow for detecting and analyzing them typologically. Based on a preliminary evaluation, the algorithm works fairly reliably in adpossessive constructions that are morphologically marked. However, it performs rather poorly in adpossessive constructions that are not marked morphologically, so-called zero-marked constructions, because of difficulties in identifying these constructions with the current annotation. We also discuss different types of variation in annotation in different treebanks for the same language and for treebanks of closely related languages. The research focuses on one well-circumscribed and universal construction in the hope of generating more interest in using UD for cross-linguistic comparison and for contributing towards developing yet more consistent annotation of constructions in the UD annotation scheme.
A study of the comprehensive reading, its importance and the basic value of an educational institution in the province of Manabí was conducted, it was investigated what is the impact on students, in which it benefits them to understand what is found in the pages of a text, how to broaden its criticality, the concentration of reading, writing and oral skills, to have a better communication and how reading influenced values, it was elaborated in a research way with the help of ICT, not practical, with the inductive method - deductive, finally, the linguistic norms, pauses, emphasis, vowel sounds that are used when reading are analyzed, such as strategies that can be used for images, graphics, mental or conceptual maps, together with teacher support inside or outside the classroom, coordinate the ways of studying the topics because each student thinks and analyzes differently.
Abstract The COVID-19 pandemic has dramatically changed the nature of our social interactions. In order to understand how protective equipment and distancing measures influence the ability to comprehend others’ emotions and, thus, to effectively interact with others, we carried out an online survey across the Italian population during the first pandemic peak. Participants were shown static facial expressions (Angry, Happy and Neutral) covered by a sanitary mask or by a scarf. They were asked to evaluate the expressed emotions as well as to assess the degree to which one would adopt physical and social distancing measures for each stimulus. Results demonstrate that, despite the covering of the lower-face, participants correctly recognized the facial expressions of emotions with a polarizing effect on emotional valence ratings found in females. Noticeably, while females’ ratings for physical and social distancing were driven by the emotional content of the stimuli, males were influenced by the “covered” condition. The results also show the impact of the pandemic on anxiety and fear experienced by participants. Taken together, our results offer novel insights on the impact of the COVID-19 pandemic on social interactions, providing a deeper understanding of the way people react to different kinds of protective face covering.
There is empirical evidence that expected yet not current affect predicts decisions. However, common research designs in affective decision-making show consistent methodological problems (e.g., conceptualization of different emotion concepts; measuring only emotional valence, but not arousal). We developed a gambling task that systematically varied learning experience, average feedback balance and feedback consistency. In Experiment 1 we studied whether predecisional current affect or expected affect predict recurrent gambling responses. Furthermore, we exploratively examined how affective information is represented on a neuronal level in Experiment 2. Expected and current valence and arousal ratings as well as Blood Oxygen Level Dependent (BOLD) responses were analyzed using a within-subject design. We used a generalized mixed effect model to predict gambling responses with the different affect variables. Results suggest a guiding function of expected valence for decisions. In the anticipation period, we found activity in brain areas previously associated with valence-general processing (e.g., anterior cingulate cortex, nucleus accumbens, thalamus) mostly independent of contextual factors. These findings are discussed in the context of the idea of a valence-general affective work-space, a goal-directed account of emotions, and the hypothesis that current affect might be used to form expectations of future outcomes. In conclusion, expected valence seems to be the best predictor of recurrent decisions in gambling tasks.
Nowadays there are a lot of modern technologies in electronic lexicography: speech synthesis technology, cross-referencing between dictionary modules, spell-checking functions, etc. The increasing availability of online information has necessitated intensive research in the area of automatic text summarization within the Natural Language Processing community. Belarusian scientists are also interested in this sphere and new lexicographical approaches for creating a linguistic database are shown in the paper. The authors present English-Belarusian-Russian electronic dictionary TechLex. This is the project of the 2 nd English Department and the Department of Software for Information Systems and Technologies of the Belarusian National Technical University. The linguistic database of the dictionary is compiled not by the traditional method of processing a large number of paper dictionaries and combining the received translations, but by sequential processing of scientific and technical English-language periodicals. While the designing the dictionary the authors have taken into account the analysis of modern electronic multilingual translation dictionaries and created a client-server application in the Java programming language. The client part of the system contains a mobile application for the Android operating system, which has been tested on tablets and smartphones with different screen diagonals. The interface of the TechLex dictionary is designed taking into account the possibility of adding new subject areas and filling them with appropriate lexical material. The main advantage of our dictionary is that it is the first technical multilingual electronic dictionary having a Belarusian version.
The robustness of text classifier based on the deep neural networks can be improved by correcting adversarial spelling mistakes. The main challenge brought by these mistakes is that it is difficult for the text classifier to recognize the words in the text correctly. Existing works to deal with these mistakes can be divided into two types: optimization of training data and reconstruction of text classification model. However, these works are not suitable for the text classification model that has been deployed, as retraining or reconstruction is always involved. To address the above problem, we propose a two-step spelling correction model, which consists of a misspelled word detector and a misspelled word correcter, referred to as DE-CO. Specifically, we use a detector to recognize misspelled words in the text, use a corrector to correct the misspelled words, and then feed the correction results into the downstream classifier. In this way, without reconstruction or retraining, the normal recognition of words by the text classifier can be guaranteed. We evaluate DE-CO on adversarial examples generated from the Stanford Sentiment Treebank (SST). The classification accuracy of the downstream classifier is improved in a variety of attack scenarios, which demonstrates that DE-CO improves the robustness of the text classifier.
Temporal models based on recurrent neural networks have proven to be quite powerful in a wide variety of applications, including language modeling and speech processing. However, training these models often relies on backpropagation through time (BPTT), which entails unfolding the network over many time steps, making the process of conducting credit assignment considerably more challenging. Furthermore, the nature of backpropagation itself does not permit the use of nondifferentiable activation functions and is inherently sequential, making parallelization of the underlying training process difficult. Here, we propose the parallel temporal neural coding network (P-TNCN), a biologically inspired model trained by the learning algorithm we call local representation alignment. It aims to resolve the difficulties and problems that plague recurrent networks trained by BPTT. The architecture requires neither unrolling in time nor the derivatives of its internal activation functions. We compare our model and learning procedure with other BPTT alternatives (which also tend to be computationally expensive), including real-time recurrent learning, echo state networks, and unbiased online recurrent optimization. We show that it outperforms these on-sequence modeling benchmarks such as Bouncing MNIST, a new benchmark we denote as Bouncing NotMNIST, and Penn Treebank. Notably, our approach can, in some instances, outperform full BPTT as well as variants such as sparse attentive backtracking. Significantly, the hidden unit correction phase of P-TNCN allows it to adapt to new data sets even if its synaptic weights are held fixed (zero-shot adaptation) and facilitates retention of prior generative knowledge when faced with a task sequence. We present results that show the P-TNCN's ability to conduct zero-shot adaptation and online continual sequence modeling.
Generating novel design concepts is a cornerstone for producing innovative products. Although many methods have been proposed for supporting the task, their performance depends on human ability. The ultimate goal of this research is to build a method supporting designers to generate novel design concepts with the knowledge of human creativity. Toward the goal, this research assumes that the more distant two function concepts chosen, the more novel idea would be come up with by the combination of the two concepts. Based on the assumption, this paper introduces a notion of novelty potential of the combination of two function concepts, and builds a method to assess it by the function similarity. Some alternative methods are proposed to calculate it with the integration of a lexical database for natural language called WordNet and a distributional semantics method called word2vec. They are verified with an evaluation experiment which performs correlation analysis between the human’s evaluation of the novelty potential and the proposed method’s assessment of function similarity. This paper discusses which method matches a human sense most, and its possibility for design concept generation based on the results of the experiment.
Aims: This study provides new insights into Arabic-English code-switching with particular reference to verb insertion. It aims to identify (1) patterns of English verb insertion into Arabic; (2) factors affecting them. We offer an alternative to previous studies’ conclusions regarding a supposed lack of English verbs integrated morphologically into Arabic, which is claimed to result from incongruence between Arabic and English verb systems. Methodology: We employ the Matrix Language Frame (MLF) model and the 4-M model. Data and analysis: The data comprise 14,414 clauses obtained from interviews with students at the American University in Cairo. Data were analyzed quantitatively. Findings: Most (80.17%) of inserted verbs were inflected with Arabic tense, gender, and number prefixes showing morphological integration into Arabic. We distinguished four recurrent patterns in verb insertion: (1) complete morphological integration in the present tense; (2) incomplete assimilation of forms requiring the use of the plural suffix -u; (3) lack of morphological integration in the past tense; and (4) lack of suffixation of Arabic clitics to English verbs. Originality: This is the first study focusing on verb insertion in Arabic-English code-switching based on empirical data collected in Egypt. It offers different findings on verb patterns and their explanation compared with other quantitative studies based on the MLF model. We propose to look beyond incongruence between Arabic and English as a factor determining verb patterns to include linguistic convention. Thus, we hypothesize that verb insertion might be controlled by linguistic norms accepted and perpetuated in a given speech community. Significance: Contrary to previous claims, our results show that patterns of verb insertion in Arabic-English code-switching are consistent with the MLF model. Hence, the study contributes evidence for the MLF model and its explanatory value.
A cache-inspired approach is proposed for neural language models (LMs) to improve long-range dependency and better predict rare words from long contexts.This approach is a simpler alternative to attention-based pointer mechanism that enables neural LMs to reproduce words from recent history.Without using attention and mixture structure, the method only involves appending extra tokens that represent words in history to the output layer of a neural LM and modifying training supervisions accordingly.A memory-augmentation unit is introduced to learn words that are particularly likely to repeat.We experiment with both recurrent neural network-and Transformer-based LMs.Perplexity evaluation on Penn Treebank and WikiText-2 shows the proposed model outperforms both LSTM and LSTM with attention-based pointer mechanism and is more effective on rare words.N -best rescoring experiments on Switchboard indicate that it benefits both very rare and frequent words.However, it is challenging for the proposed model as well as two other models with attention-based pointer mechanism to obtain good overall WER reductions.
Pointing devices are the primary media of interac-tion between humans and computers. The three most popular pointing devices used in computers (both portable and non-portable) are mouse, touchpad and nubs (joystick). They have their different advantages and use cases while being targeted to different user groups. The aim of this study was to investigate whether the aforementioned pointing devices have different effects on human valence. A total of 12 participants were recruited for the experiment. Each participant completed a pointing reaction test with every pointing device aforementioned, where they selected as many randomly appearing circles as possible in a given amount of time. Then, subjective ratings of emotional valence and arousal were collected, and the effects of the pointing device used on these ratings were investigated. Our study shows that the valence rating of using the mouse were significantly higher in challenging scenarios, compared to the likes of touchpad and nub.
OBJECTIVES: To investigate whether tethered swimming (TS) performed 8 minutes before a 50-m freestyle swimming sprint could be an effective postactivation potentiation method to improve performance in young swimmers. METHODS: Fourteen regional-level male adolescent swimmers (age 13.0 [2.0] y; height 161.1 [12.4] cm; body mass 52.5 [9.5] kg) underwent 2 trial conditions in a randomized and counterbalanced order (1 experimental [TS], 1 control) on different days. During the experimental session, the participants performed a standard warm-up of 1200 m followed by a TS exercise, which consisted of 3 × 10-second maximal efforts of TS with 1-minute rests between bouts. In the control condition, the warm-up phase was immediately followed by 200 m at a moderate pace (same duration as the TS in the experimental session). Performance (time trial); biomechanical (stroke length), physiological (blood lactate concentrations), and psychophysiological (ratings of perceived exertion) variables; and countermovement-jump (CMJ) flight time were collected. RESULTS: TS warm-up had no significant effect on 50-m swimming performance (P =.27), postexercise ratings of perceived exertion, stroke length, or CMJ flight time (P ≥.05). Blood lactate concentrations significantly increased at the end of the warm-up in the TS condition only (interaction effect: F1.91,29.91 = 4.91, P =.01, η2 =.27) and after the 50-m trial in both conditions (F1.57,20.41 = 62.39, P =.001, η2 =.82). CONCLUSIONS: The present study demonstrated that 3 × 10-second TS exercises performed 8 minutes prior to the event did not affect ratings of perceived exertion, stroke length, or CMJ flight time. In addition, tethered swimming did not affect 50-m freestyle sprint performance in young swimmers.
Abstract Syntactic parsing is an important topic in the field of Mongolian language information processing. Compared with English and Chinese dependency parsing, Mongolian dependency parsing is still at the beginning stage. Mongolian syntactic parsing lack of Treebank resources seriously, under such conditions, a high quality syntactic parser cannot be developed by statistical methods simply. Aiming at the characteristics that Mongolian language has rich morphological features, this paper presented a rule and statistics-based dependency parsing model using Mongolian Dependency Treebank as training and evaluation data. The morphological and syntactic rules are represented using complex features and unification operations. The statistical model is represented using lexical dependent probability. This model has now achieved accuracies of 77.18%, 69.42% and 95.44% for the unlabelled annotation score, the labeled annotation score and the head word annotation score respectively.
Transformer-based pre-trained language models (PLMs) have dramatically improved the state of the art in NLP across many tasks. This has led to substantial interest in analyzing the syntactic knowledge PLMs learn. Previous approaches to this question have been limited, mostly using test suites or probes. Here, we propose a novel fully unsupervised parsing approach that extracts constituency trees from PLM attention heads. We rank transformer attention heads based on their inherent properties, and create an ensemble of high-ranking heads to produce the final tree. Our method is adaptable to low-resource languages, as it does not rely on development sets, which can be expensive to annotate. Our experiments show that the proposed method often outperform existing approaches if there is no development set present. Our unsupervised parser can also be used as a tool to analyze the grammars PLMs learn implicitly. For this, we use the parse trees induced by our method to train a neural PCFG and compare it to a grammar derived from a human-annotated treebank.
Neural sequence model, though widely used for modeling sequential data such as the language model, has sequential recency bias (Kuncoro et al. 2018) to the local context, limiting its full potential to capture long-distance context. To address this problem, this paper proposes augmenting sequence models with a span-based neural buffer that efficiently represents long-distance context, allowing a gate policy network to make interpolated predictions from both the neural buffer and the underlying sequence model. Training this policy network to utilize long-distance context is however challenging due to the simple sentence dominance problem (Marvin and Linzen 2018). To alleviate this problem, we propose a novel training algorithm that combines an annealed maximum likelihood estimation with an intrinsic reward-driven reinforcement learning. Sequence models with the proposed span-based neural buffer significantly improve the state-of-the-art perplexities on the benchmark Penn Treebank and WikiText-2 datasets to 43.9 and 35.2 respectively. We conduct extensive analysis and confirm that the proposed architecture and the training algorithm both contribute to the improvements.
Corporate credit ratings (CRs) are closely related to companies’ cost of debt financing. Recent research has drawn wide attention to how nonfinancial as well as financial factors may affect ratings. By manually collecting information about the profiles of chief financial officers (CFOs) of US companies, we examine the effect of CFOs’ accounting expertise on corporate CRs. The results show that firms with accounting expert CFOs are more likely to receive higher CRs and that the effect of CFOs’ accounting expertise on the ratings is more pronounced for firms with higher default risk, suggesting that the accounting expertise of CFOs may be an important factor that affects CRs. Moreover, we find a dynamic relation between accounting expert CFOs and CRs such that a downgrade in a firm’s CR in a prior year affects the subsequent selection of an accounting expert CFO.
In order to effectively respond to the increased linguistic and cultural diversity in the U.S. schools and close the consistently documented achievement gap between culturally and linguistically diverse (CLD) students and mainstream students, teachers need to take an asset-based approach and be able to draw on CLD students’ entire funds of linguistic knowledge. However, few studies have examined CLD students’ linguistic choices in multiple discursive spaces with different linguistic norms, values and practices. This article addresses this research gap through a case study of Elif, a Turkish American student and her linguistic boundary crossing experiences within and across three discursive spaces: her home, her Turkish heritage language school, and her mainstream school. Through in-depth analysis of interviews, observations, and field notes, the study revealed that Elif experienced different linguistic environments and boundary types. She negotiated experiences that ranged from smooth to managed to insurmountable boundaries. Finally, translanguaging practices acted as a key boundary object that mediated sociocultural discontinuities in the Turkish heritage language school, and facilitated Elif’s experiences between Turkish dominant and English dominant discursive spaces.
Our work on the automatic detection of English discourse connectives in the Penn Discourse Treebank (PDTB) shows that syntactic information from the Universal Dependencies (UD) framework is a viable alternative to that from the Penn Treebank (PTB) framework. In fact, we found minor increases when comparing between the use of gold standard PTB part-of-speech (POS) tag information and automatically parsed UD information. The former has traditionally been used for the task but there are now much more UD corpora and in many more languages than that available in the PTB framework. As such, this finding is promising for areas in discourse parsing such as in multilingual as well as under production settings, where gold standard PTB information may be scarce.
BACKGROUND: Analyze intrarater and interrater reliability for evaluating endoscopic images of velopharyngeal (VP) physiology. METHOD: Speakers produced 9 speech stimuli representing 4 stimulus types: sustained phonemes, repetitions of "puh," single words, and short phrases. The 37-speaker participants included 16 patients with VP dysfunction and 21 control participants. Five raters independently rated the video images for degree of VP opening, location of opening, and pattern of closure. Outcome measures included intrarater and interrater measures of reliability and the effects of raters and stimulus type on ratings. RESULTS: Intrarater reliability was acceptable, and ratings were logically consistent. Fixed effects regression coefficients for the patient and the control groups showed that raters were a significant source of variability for degree of opening and pattern of closing. Stimulus type was not a significant source of variation for any metric for the controls, but stimulus type was a significant determinant for degree of opening for patients. The degree of opening was larger for sustained phonemes than for the other speech stimuli. Ratings for degree of opening were most similar for repeated "puh." CONCLUSIONS: Interrater reliability needs to be improved so that the assessment procedure produces more consistent findings among clinicians, thus strengthening our evidence base for this procedure. Interrater additional research is needed to understand how the stimulus affects ratings of VP physiology, to identify stimuli that yield the most useful clinical information, and to understand how training affects the ratings of VP physiology.
Increasing popularity of electronic dictionaries, ontologies, thesauri and lexical databases makes them an effective tool for language learning purposes (Dash, 2013; Fellbaum, 2010; Miller & Fellbaum, 1992; Shimodaira et al, 2006; Sun et al, 2011). The aim of this research is to study the educational potential of electronic lexical database for the English language WordNet (Miller, 1995; Fellbaum, 1998) and electronic thesauri for the Russian language RuWordNet (Loukachevitch, 2011; Loukachevitch, Lashevich, 2016) in teaching English as a foreign language. In this research the authors focus on teaching colours, in particular a colour term white, as colours represent complex linguistic and culture-specific phenomena, reflected in the “cultural memory” of people, the very concept of ‘colour’ being a function of language and culture (Wierzbicka, 2006). Thus, understanding colours helps students both study a foreign language and learn its history and culture.This is a mixed method study based on comparison of synonyms for colour term white in WordNet and RuWordNet. The main relation among words in these thesauri is synonymy. Based on their meanings, words are grouped into unordered sets of synonyms expressing one underlying concept (synsets). This allows to consider specific senses of words and semantic relations. Firstly, Russian students studying English as a foreign language were asked to analyse the meanings and synonyms for colour term white in RuWordNet. Secondly, the students compared the representation of colour term white in WordNet. Additionally, they examined set phrases with the adjective white in English dictionaries. Then a qualitative method (interviewing) was used to reveal the students’ perceptions of studying English by means of electronic dictionaries, thesauri and lexical databases.The study allowed to claim that electronic thesauri and lexical databases are of educational value, increasing students’ linguistic awareness and language proficiency.
The research work is dealt with the culture of speech. Culture of speech is identified by language of speech, social surrounding, language and psycology, language and pragmatics. As a result of culture of speech, personal culture, human quality, linguistic knowledge of a person is realized. Several types of antropolinguistic analyses are mentioned. Public speech (in auditorium, in crowd) shows the social aspect of communicative, linguistic norms and humans morality in speaking expresses wisdom and psycholinguistic aspect of speaker. Literary norm, functional grammar, cognitive pragmatics, literary language are thoroughly explained.
OBJECTIVE: To design and evaluate the effectiveness of a stimulus material in eliciting the N400 event related potential (ERP). DESIGN: A set of 700 semantically congruent and incongruent sentences was developed in accordance with current linguistic norms, and validated with an electroencephalography (EEG) study, in which the influence of age and gender on the N400 ERP magnitude was analysed. STUDY SAMPLE: Forty-five normal-hearing subjects (19-57 years, 21 females) participated in the EEG study. RESULTS: The stimulus material used in the EEG study elicited a robust N400 ERP, with a morphology consistent with the literature. Results also showed no statistically significant effect of age or gender on the N400 magnitude. CONCLUSIONS: The material presented in this paper constitutes the largest complete stimulus set suitable for both auditory and text-based N400 experiments. This material may help facilitate the efficient implementation of future N400 ERP studies, as well as promote standardisation and consistency across studies.
This paper represents the development of the Myanmar Named Entity Recognition (NER) system using Conditional Random Fields (CRFs). In order to develop the system, a manually annotated Named Entities (NEs) corpus - collected from Myanmar news websites and Asia Language Treebank(ALT)-Parallel-Corpus has been used. We compare the performance of the system getting syllable-based input to the one getting character-based input. We observed that training data has more impact on the performance of the system. The experimental results show that the syllable-based system performs better than the character-based system. It achieves that Precision, Recall and F1-score values of 93.62%, 91.64% and 92.62% respectively.
Scene graph is a graph representation that explicitly represents high-level semantic knowledge of an image such as objects, attributes of objects and relationships between objects. Various tasks have been proposed for the scene graph, but the problem is that they have a limited vocabulary and biased information due to their own hypothesis. Therefore, results of each task are not generalizable and difficult to be applied to other down-stream tasks. In this paper, we propose Entity Synset Alignment(ESA), which is a method to create a general scene graph by aligning various semantic knowledge efficiently to solve this bias problem. The ESA uses a large-scale lexical database, WordNet and Intersection of Union (IoU) to align the object labels in multiple scene graphs/semantic knowledge. In experiment, the integrated scene graph is applied to the image-caption retrieval task as a downstream task. We confirm that integrating multiple scene graphs helps to get better representations of images.
This paper analyses data to address a specific linguistic problem, i.e. the acquisition of the modification potential of the three more or less synonymous Dutch degree modifiers heel, erg and zeer, all meaning 'very', which show syntactic differences in modification potential. It continues the research reported on in The analysis makes crucial use of linguistic applications developed in the CLARIN infrastructure, in particular the treebank search applications PaQu (Parse and Query) and GrETEL Version 4.00. The analysis benefits from the use of parsed corpora (treebanks) in combination with the search and analysis options offered by PaQu and GrETEL. Earlier work showed that despite little data for zeer modifying adpositional phrases adult speakers end up with a generalised modification potential for this word. In this paper, I extend the dataset considered, and find more (but still little) data for this phenomenon. However, I also find a similar amount of data that form counterexamples to the non-generalisation of the modification potential of heel. I argue that the examples with heel concern constructions with idiosyncratic semantics and therefore are not counted as evidence for the general rule of modification. I suggest a simple statistical analysis to account for the fact that children 'learn' that heel cannot modify verbs or adpositions though there is no explicit evidence for this and they are not explicitly taught so.
BACKGROUND: Instrumental activities of daily living (IADL) impairment can begin in mild cognitive impairment (MCI), and is the core criteria for diagnosing dementia in both Alzheimer's (AD) and Parkinson's (PD) diseases. The Functional Activities Questionnaire (FAQ) has high discriminative power for dementia and MCI in older age populations, but is influenced by demographic factors. It is currently unclear whether the FAQ is suitable for assessing cognitive-associated IADL in non-demented PD patients, as motor disorders may affect ratings. OBJECTIVE: To compare IADL profiles in MCI patients with PD (PD-MCI) and AD (AD-MCI) and to verify the discriminative ability of the FAQ for MCI in patients with (PD-MCI) and without (AD-MCI) additional motor impairment. METHODS: Data of 42 patients each of PD-MCI, AD-MCI, PD cognitively normal (PD-CN), and healthy controls (HC), matched according to age, gender, education, and global cognitive impairment were analyzed. ANCOVA and binary regressions were used to examine the relationship between the FAQ scores and groups. FAQ cut-offs for PD-MCI (versus PD-NC) and AD-MCI (versus HC) were separately identified using receiver operating characteristic analyses. RESULTS: FAQ total score did not differentiate between MCI groups. PD-MCI subjects had greater difficulties with tax records and traveling while AD-MCI individuals were more impaired in managing finances and remembering appointments. Classification accuracy of the FAQ was good for diagnosing AD-MCI (69%, cut-off ≥1) compared to HC, and sufficient for differentiating PD-MCI (38.1%, cut-off ≥3) from PD-CN. CONCLUSION: The FAQ task profiles and classification accuracy differed between MCI related to PD and AD.
Using incorrect worked examples during mathematics instruction can improve student learning. However, teachers worry that students may confuse correct and incorrect examples over time, and memory research supports this fear. To examine if this forgetting occurs, we had undergraduates rate the correctness of correct and incorrect worked examples immediately and one week later (Experiment 1). Previously studied incorrect examples were rated as slightly more correct after the delay, but this did not affect ratings of unstudied examples or problem-solving accuracy. In Experiment 2, we more closely mimicked how incorrect worked examples are used in classroom settings. Again, we found only small changes in students’ memory for studied worked examples after the delay, and no changes for unstudied examples or problem-solving accuracy. Our findings suggest the costs of teaching with incorrect worked examples are limited to the specific studied problems, and do not affect learning of the underlying mathematical rule.
The aim of this paper is to retrieve the most relevant expansion words for expanding the initial query of the user in order to enhance the outcomes of web search results. Query expansion plays a major role in reformulating a user’s initial query to a one more pertinent to the user’s intended meaning. The reformulated query is then used to obtain more appropriate outcomes from a large amount of information on the web. The proposed semantic query expansion technique uses Wikipedia and WordNet as data sources. Wikipedia is taken as a base for all query expansions because it is one of the most diversified and relevant databases available on the web. To further improve the proposed query expansion technique,WordNet—a lexical database—is used as the as another data source because the synonyms (synsets) of the query term provided by it can be quite useful for query expansion. The proposed expansion technique successfully combines the two data sources to retrieve the most relevant expansion terms from the data sources in response to the user’s original query. The proposed work has been divided into four phases: (1) extraction of relevant words from Wikipedia (2) extraction of relevant words from WordNet (3) merging of the expansion terms obtained from Wikipedia and WordNet, and (4) query formulation by combining the expansion terms using Boolean operators. This reformulated query is then fired on the web to find the desired result. The Experimental result shows a significant improvement in information retrieval using query expansion.
This journal article follows the research line opened on the search for semantic primes’ exponents in Old English within the frame of the Natural Semantic Metalanguage theory (Goddard 1997, 2012; Goddard and Wierzbicka 2002). The aim of this study is to complete the line of research on prime identification opened on the category Actions, events, movement, contact by establishing the Old English exponent of the prime DO. With this purpose, this paper discusses the adequacy of different OE verbs as possible prime exponent on the basis of textual frequency, morphology, semantics and syntactic complementation. Relevant data of analysis have been retrieved mainly from the lexical database of Old English Nerthus, the Dictionary of Old English (Healey et al. 2018) and the Dictionary of Old English Corpus (Healey et al. 2009).
Neural machine translation (NMT) models are typically trained using a softmax cross-entropy loss where the softmax distribution is compared against smoothed gold labels. In low-resource scenarios, NMT models tend to over-fit because the softmax distribution quickly approaches the gold label distribution. To address this issue, we propose to divide the logits by a temperature coefficient, prior to applying softmax, during training. In our experiments on 11 language pairs in the Asian Language Treebank dataset and the WMT 2019 English-to-German translation task, we observed significant improvements in translation quality by up to 3.9 BLEU points. Furthermore, softmax tempering makes the greedy search to be as good as beam search decoding in terms of translation quality, enabling 1.5 to 3.5 times speed-up. We also study the impact of softmax tempering on multilingual NMT and recurrently stacked NMT, both of which aim to reduce the NMT model size by parameter sharing thereby verifying the utility of temperature in developing compact NMT models. Finally, an analysis of softmax entropies and gradients reveal the impact of our method on the internal behavior of NMT models.
The article analyzes differences in the description of discourse relations in corpus research, in particular with the reference to the use of discourse markers – expressions that tie together subsequent fragments of the text and provide information about the nature of these relations. The text presents three concepts of the description of explicitness and implicitness of the content: Rhetorical Structure Theory, Penn Discourse Treebank and the author’s original proposal and indicates the consequences of each solution. The analysis of relations with particles as metatexual expressions defined in accordance with The Nest Dictionary of Polish reveals the possibility of expressing explicitness as a representation of elements of informational structure shaped by the use of a given particle, and implicitness as a lack of representation of certain elements of this type.