Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Cross-lingual dependency parsing approaches have been employed to develop dependency parsers for the languages for which little or no treebanks are available using the treebanks of other languages. A language for which the cross-lingual parser is developed is usually referred to as the target language and the language whose treebank is used to train the cross-lingual parser model is referred to as the source language. The cross-lingual parsing approaches for dependency parsing may be broadly classified into three categories: model transfer, annotation projection, and treebank translation. This survey provides an overview of the various aspects of the model transfer approach of cross-lingual dependency parsing. In this survey, we present a classification of the model transfer approaches based on the different aspects of the method. We discuss some of the challenges associated with cross-lingual parsing and the techniques used to address these challenges. In order to address the difference in vocabulary between two languages, some approaches use only non-lexical features of the words to train the models while others use shared representations of the words. Some approaches address the morphological differences by chunk-level transfer rather than word-level transfer. The syntactic differences between the source and target languages are sometimes addressed by transforming the source language treebanks or by combining the resources of multiple source languages. Besides cross-lingual transfer parser models may be developed for a specific target language or it may be trained to parse sentences of multiple languages. With respect to the above-mentioned aspects, we look at the different ways in which the methods can be classified. We further classify and discuss the different approaches from the perspective of the corresponding aspects. We also demonstrate the performance of the transferred models under different settings corresponding to the classification aspects on a common dataset.
In this paper we propose a novel data augmentation approach where guided outputs of a language generation model, e.g. GPT-2, when labeled, can improve the performance of text classifiers through an active learning process. We transform the data generation task into an optimization problem which maximizes the usefulness of the generated output, using Monte Carlo Tree Search (MCTS) as the optimization strategy and incorporating entropy as one of the optimization criteria. We test our approach against a Non-Guided Data Generation (NGDG) process that does not optimize for a reward function. Starting with a small set of data, our results show an increased performance with MCTS of 26% on the TREC-6 Questions dataset, and 10% on the Stanford Sentiment Treebank SST-2 dataset. Compared with NGDG, we are able to achieve increases of 3% and 5% on TREC-6 and SST-2.
Several studies have suggested that females and males differ in reward behaviors and their underlying neural circuitry. Whether human sex differences extend across neural and behavioral levels for both rewards and punishments remains unclear. We studied a community sample of 221 young women and men who performed a monetary incentive task known to engage the mesoaccumbal pathway and salience network. Both stimulus salience (behavioral relevance) and valence (win vs loss) varied during the task. In response to high- vs low-salience stimuli presented during the monetary incentive task, men showed greater subjective arousal ratings, behavioral accuracy and skin conductance responses (P < 0.006, Hedges' effect size g = 0.38 to 0.46). In a subsample studied with functional magnetic resonance imaging (n = 44), men exhibited greater responsiveness to stimulus salience in the nucleus accumbens, midbrain, anterior insula and dorsal anterior cingulate cortex (P < 0.02, g = 0.86 to 1.7). Behavioral, autonomic and neural sensitivity to the valence of stimuli did not differ by sex, indicating that responses to rewards vs punishments were similar in women and men. These results reveal novel and robust sex differences in reward- and punishment-related traits, behavior, autonomic activity and neural responses. These convergent results suggest a neurobehavioral basis for sexual dimorphism observed in the reward system, including reward-related disorders.
This qualitative study investigates the lived experiences of interculturality among international students enrolled in a top-rated comprehensive university in Shanghai and a key provincial university in China’s hinterland. Based on an analysis of narrative interviews with 20 international students from different sociocultural backgrounds, the study examines how the international students’ intercultural experiences are simultaneously constrained and facilitated by the linguacultural resources distributed among interactants in the universities and cities. The results suggest that the international students’ intercultural experiences differ in substantial ways, mediated by varied distributions of sociolinguistic resources which are closely associated with their academic socialisation through English, local interactional norms, relevant career opportunities, and linguistic norms for everyday interaction with both Chinese students and local residents in the two universities. It is also found that international students play agentive roles to construct and negotiate scales among the sociolinguistic resources at their disposal, in order to create meanings for and reflect upon their lived experiences of interculturality in relation to their previous and prospective life trajectories.
We trained a model to distinguish an extreme high arousal, unpleasant drink from regular drinks based on a range of implicit behavioral and physiological responses to naturalistic tasting. The trained model predicted arousal ratings of regular drinks, highlighting the possibility to estimate affective experience without having to rely on subjective ratings.
The training of Deep Neural Networks (DNNs) brings enormous memory requirements and computational complexity, which makes it a challenge to train DNN models on resource-constrained devices. Training DNNs with reduced-precision data representation is crucial to mitigate this problem. In this article, we conduct a thorough investigation on training DNNs with low-bit posit numbers, a Type-III universal number (Unum). Through a comprehensive analysis of quantization with various data formats, it is demonstrated that the posit format shows great potential to be employed in the training of DNNs. Moreover, a DNN training framework using 8-bit posit is proposed with a novel tensor-wise scaling scheme. The experiments show the same performance as the state-of-the-art (SOTA) across multiple datasets (MNIST, CIFAR-10, ImageNet, and Penn Treebank) and model architectures (LeNet-5, AlexNet, ResNet, MobileNet-V2, and LSTM). We further design an energy-efficient hardware prototype for our framework. Compared to the standard floating-point counterpart, our design achieves a reduction of 68, 51, and 75 percent in terms of area, power, and memory capacity, respectively.
Federated learning (FL) is a privacy-preserving technique for training a vast amount of decentralized data and making inferences on mobile devices. As a typical language modeling problem, mobile keyboard prediction aims at suggesting a probable next word or phrase and facilitating the human-machine interaction in a virtual keyboard of the smartphone or laptop. Mobile keyboard prediction with FL hopes to satisfy the growing demand that high-level data privacy be preserved in artificial intelligence applications even with the distributed models training. However, there are two major problems in the federated optimization for the prediction: (1) aggregating model parameters on the server-side and (2) reducing communication costs caused by model weights collection. To address the above issues, traditional FL methods simply use averaging aggregation or ignore communication costs. We propose a novel Federated Mediation (FedMed) framework with the adaptive aggregation, mediation incentive scheme, and topK strategy to address the model aggregation and communication costs. The performance is evaluated in terms of perplexity and communication rounds. Experiments are conducted on three datasets (i.e., Penn Treebank, WikiText-2, and Yelp) and the results demonstrate that our FedMed framework achieves robust performance and outperforms baseline approaches.
Abstract The rise of world Englishes has challenged the emphasis on native-speaker accents and cultures in English language teaching. This study aimed to investigate the representation of world Englishes and cultures in three global language teaching textbooks, namely Interchange, English Result, and American English File. The textbooks were subjected to content analysis regarding their reference to Inner, Outer, and Expanding Circles’ varieties and their associated cultural contents. Kachru ( The alchemy of English: The spread functions and models of non-native Englishes. Oxford: Pergamon, 1986) notion of Concentric Circles, and the categorization proposed by Pfister and Borzelli ( Unterrichtspraxis 10:102–108, 1977) functioned as a framework to see which aspects of each culture (social, personal, religion/arts/humanities, political systems and institutions, and environmental concerns) were addressed in these textbooks. Findings revealed that most of the references to the three circles and cultural elements embodied in the textbooks were toward Inner Circle countries in American English File. Furthermore, in Interchange and English Result series, reference to Outer and Expanding Circles’ varieties and cultural elements were comparatively more evident. However, all the three textbook series mostly represented Inner-Circle accents. These findings have implications for materials developers to adopt an EIL-aware approach and to avoid the sole representation of native speakers’ linguistic norms and cultures in ELT textbooks.
The chapter addresses the concept of linguistic norm in the tradition of classical grammar and rhetoric, paying special attention to activities concerning standardization processes in the Romance languages. Since a clear distinction between a prescriptive and a descriptive point of view is not given in "traditional grammar", the latter is manifested in the form of grammatical treatises which often also aimed at offering norms for "correct" language use. As a consequence thereof, our contribution will be concerned with aspects relating to the realm of the history of language sciences and, at least partially, to the history of rhetoric. The period taken into consideration ranges from Latin antiquity (Cicero, Quintilian) to the middle of the 17th century (Vaugelas). The topics to be discussed were selected with regard to the significance of the respective protagonists in the history of ideas in (Latin and) Romance language standardization.
In this study we evaluate the convergent validity of a new graphical self-report tool (the EmojiGrid) for the affective appraisal of perceived touch events. The EmojiGrid is a square grid labeled with facial icons (emoji) showing different levels of valence and arousal. The EmojiGrid is language independent and efficient (a single click suffices to report both valence and arousal), making it a practical instrument for studies on affective appraisal. We previously showed that participants can intuitively and reliably report their affective appraisal (valence and arousal) of visual, auditory and olfactory stimuli using the EmojiGrid, even without additional (verbal) instructions. However, because touch events can be bidirectional and dynamic, these previous results cannot be generalized to the touch domain. In this study, participants reported their affective appraisal of video clips showing different interpersonal (social) and object-based touch events, using either the validated 9-point SAM (Self-Assessment Mannikin) scale or the EmojiGrid. The valence ratings obtained with the EmojiGrid and the SAM are in excellent agreement. The arousal ratings show good agreement for object-based touch and moderate agreement for social touch. For social touch and at more extreme levels of valence, the EmojiGrid appears more sensitive to arousal than the SAM. We conclude that the EmojiGrid can also serve as a valid and efficient graphical self-report instrument to measure human affective response to a wide range of tactile signals.
Lexical simplification (LS) aims to replace complex words in a given sentence with their simpler alternatives of equivalent meaning. Recently unsupervised lexical simplification approaches only rely on the complex word itself regardless of the given sentence to generate candidate substitutions, which will inevitably produce a large number of spurious candidates. We present a simple LS approach that makes use of the Bidirectional Encoder Representations from Transformers (BERT) which can consider both the given sentence and the complex word during generating candidate substitutions for the complex word. Specifically, we mask the complex word of the original sentence for feeding into the BERT to predict the masked token. The predicted results will be used as candidate substitutions. Despite being entirely unsupervised, experimental results show that our approach obtains obvious improvement compared with these baselines leveraging linguistic databases and parallel corpus, outperforming the state-of-the-art by more than 12 Accuracy points on three well-known benchmarks.
Machine learning (ML) has progressed rapidly during the past decade and ML models have been deployed in various real-world applications. Meanwhile, machine learning models have been shown to be vulnerable to various security and privacy attacks. One attack that has attracted a great deal of attention recently is the backdoor attack. Specifically, the adversary poisons the target model training set, to mislead any input with an added secret trigger to a target class, while keeping the accuracy for original inputs unchanged. Previous backdoor attacks mainly focus on computer vision tasks. In this paper, we present the first systematic investigation of the backdoor attack against models designed for natural language processing (NLP) tasks. Specifically, we propose three methods to construct triggers in the NLP setting, including Char-level, Word-level, and Sentence-level triggers. Our Attacks achieve an almost perfect success rate without jeopardizing the original model utility. For instance, using the word-level triggers, our backdoor attack achieves 100% backdoor accuracy with only a drop of 0.18%, 1.26%, and 0.19% in the models utility, for the IMDB, Amazon, and Stanford Sentiment Treebank datasets, respectively.
BACKGROUND: When the COVID-19 pandemic restricted visitation between intensive care unit patients and their families, the virtual intensive care unit (vICU) in our large tertiary hospital was adapted to facilitate virtual family visitation. The objective of this paper is to document findings from interviews conducted with family members on three categories: (1) feelings experienced during the visit, (2) barriers, challenges or concerns faced using this service, and (3) opportunities for improvements. METHODS: Family members were interviewed postvisit via phone. For category 1 (feelings), automated analysis in Python using the Valence Aware Dictionary for sentiment Reasoner package produced weighted valence (extent of positive, negative or neutral emotive connotations) of the interviewees' word choices. Outputs were compared with a manual coder's valence ratings to assess reliability. Two raters conducted inductive thematic analysis on the notes from these interviews to analyse categories 2 (barriers) and 3 (opportunities). RESULTS: Valence-based and manual sentiment analysis of 230 comments received on feelings showed over 86% positive sentiments (88.2% and 86.8%, respectively) with some neutral (7.3% and 6.8%) and negative (4.5% and 6.4%) sentiments. The qualitative analysis of data from 57 participants who commented on barriers showed four primary concerns: inability to communicate due to patient status (44% of respondents); technical difficulties (35%); lack of touch and physical presence (11%); and frequency and clarity of communications with the care team (11%). Suggested improvements from 59 participants included: on demand access (51%); improved communication with the care team (17%); improved scheduling processes (10%); and improved system feedback and technical capabilities (17%). CONCLUSIONS: Use of vICU for remote family visitations evoked happiness, joy, gratitude and relief and a sense of closure for those who lost loved ones. Identified areas for concern and improvement should be addressed in future implementations of telecritical care for this purpose.
It is now a common practice to compare models of human language processing by comparing how well they predict behavioral and neural measures of processing difficulty, such as reading times, on corpora of rich naturalistic linguistic materials. However, many of these corpora, which are based on naturally-occurring text, do not contain many of the low-frequency syntactic constructions that are often required to distinguish between processing theories. Here we describe a new corpus consisting of English texts edited to contain many low-frequency syntactic constructions while still sounding fluent to native speakers. The corpus is annotated with hand-corrected Penn Treebank-style parse trees and includes self-paced reading time data and aligned audio recordings. We give an overview of the content of the corpus, review recent work using the corpus, and release the data.
Emotion research typically searches for consistency and specificity in physiological activity across instances of an emotion category, such as anger or fear, yet studies to date have observed more variation than expected. In the present study, we adopt an alternative approach, searching inductively for structure within variation, both within and across participants. Following a novel, physiologically-triggered experience sampling procedure, participants' self-reports and peripheral physiological activity were recorded when substantial changes in cardiac activity occurred in the absence of movement. Unsupervised clustering analyses revealed variability in the number and nature of patterns of physiological activity that recurred within individuals, as well as in the affect ratings and emotion labels associated with each pattern. There were also broad patterns that recurred across individuals. These findings support a constructionist account of emotion which, drawing on Darwin, proposes that emotion categories are populations of variable instances tied to situation-specific needs.
assumptions about relevant dimensions. We employed the RC method to visualize mental representations of self and examined their relationships with traits related to self-image. For this purpose, 110 participants (70 women) performed a two-image forced choice RC task to generate a classification image of self (self-CI). Participants perceived their self-CIs as bearing a stronger resemblance to themselves than did CIs of others (filler-CIs). Valence ratings of participants who performed the RC task (RC sample) and of 30 independent raters both showed positive correlations with self-esteem, explicit self-evaluation, and extraversion. Moreover, valence ratings of independent raters were negatively correlated with social anxiety symptoms. On the other hand, valence ratings of the RC sample and independent raters were not correlated with depression symptoms, trait anxiety, or social desirability. The results imply that mental representations of self can be properly visualized by using the RC method.
Most syntactic dependency parsing models may fall into one of two categories: transition- and graph-based models. The former models enjoy high inference efficiency with linear time complexity, but they rely on the stacking or re-ranking of partially-built parse trees to build a complete parse tree and are stuck with slower training for the necessity of dynamic oracle training. The latter, graph-based models, may boast better performance but are unfortunately marred by polynomial time inference. In this paper, we propose a novel parsing order objective, resulting in a novel dependency parsing model capable of both global (in sentence scope) feature extraction as in graph models and linear time inference as in transitional models. The proposed global greedy parser only uses two arc-building actions, left and right arcs, for projective parsing. When equipped with two extra non-projective arc-building actions, the proposed parser may also smoothly support non-projective parsing. Using multiple benchmark treebanks, including the Penn Treebank (PTB), the CoNLL-X treebanks, and the Universal Dependency Treebanks, we evaluate our parser and demonstrate that the proposed novel parser achieves good performance with faster training and decoding.
One way music is thought to convey emotion is by mimicking acoustic features of affective human vocalizations [Juslin and Laukka (2003). Psychol. Bull. 129(5), 770-814]. Regarding fear, it has been informally noted that music for scary scenes in films frequently exhibits a "scream-like" character. Here, this proposition is formally tested. This paper reports acoustic analyses for four categories of audio stimuli: screams, non-screaming vocalizations, scream-like music, and non-scream-like music. Valence and arousal ratings were also collected. Results support the hypothesis that a key feature of human screams (roughness) is imitated by scream-like music and could potentially signal danger through both music and the voice.
Linguistic Melanesia is the linguistically diverse area centred on the island of New Guinea. At its core the area is dominated by genealogically diverse Papuan languages, but also takes in a large number of languages of the Austronesian family. Austronesian languages show a progressive convergence on the linguistic norms of Papuan languages the closer they are to New Guinea. This attenuation of Austronesian features to Papuan ones results in concentric circles of linguistic features clustering around New Guinea. Contact between Papuan languages has also led to smaller zones in which competing forces of convergence and divergence can be seen at play.
This paper explores the knowledge of linguistic structure learned by large artificial neural networks, trained via self-supervision, whereby the model simply tries to predict a masked word in a given context. Human language communication is via sequences of words, but language understanding requires constructing rich hierarchical structures that are never observed explicitly. The mechanisms for this have been a prime mystery of human language acquisition, while engineering work has mainly proceeded by supervised learning on treebanks of sentences hand labeled for this latent structure. However, we demonstrate that modern deep contextual language models learn major aspects of this structure, without any explicit supervision. We develop methods for identifying linguistic hierarchical structure emergent in artificial neural networks and demonstrate that components in these models focus on syntactic grammatical relationships and anaphoric coreference. Indeed, we show that a linear transformation of learned embeddings in these models captures parse tree distances to a surprising degree, allowing approximate reconstruction of the sentence tree structures normally assumed by linguists. These results help explain why these models have brought such large improvements across many language-understanding tasks.
Treebanks are valuable linguistic resources that include the syntactic\nstructure of a language sentence in addition to POS-tags and morphological\nfeatures. They are mainly utilized in modeling statistical parsers. Although\nthe statistical natural language parser has recently become more accurate for\nlanguages such as English, those for the Arabic language still have low\naccuracy. The purpose of this paper is to construct a new Arabic dependency\ntreebank based on the traditional Arabic grammatical theory and the\ncharacteristics of the Arabic language, to investigate their effects on the\naccuracy of statistical parsers. The proposed Arabic dependency treebank,\ncalled I3rab, contrasts with existing Arabic dependency treebanks in two main\nconcepts. The first concept is the approach of determining the main word of the\nsentence, and the second concept is the representation of the joined and covert\npronouns. To evaluate I3rab, we compared its performance against a subset of\nPrague Arabic Dependency Treebank that shares a comparable level of details.\nThe conducted experiments show that the percentage improvement reached up to\n7.5% in UAS and 18.8% in LAS.\n
AlpinoGraph is a graph-based search engine which provides treebank search using SQL database technology coupled with the Cypher query language for graphs. In the paper, we show that AlpinoGraph is a very powerful and very flexible approach towards treebank search. At the same time, AlpinoGraph is efficient. Currently, AlpinoGraph is applicable for all standard Dutch treebanks. We compare the Cypher queries in AlpinoGraph with the XPath queries used in earlier treebank search applications for the same treebanks. We also present a pre-processing technique which speeds up query processing dramatically in some cases, and is applicable beyond AlpinoGraph.
An exploration of the physiological correlates of subjective emotional states has theoretical and practical significance. Previous studies have reported that subjective valence and arousal correspond to facial electromyography (EMG) and electrodermal activity (EDA), respectively, across stimuli. However, the reported results were inconsistent, no study investigated subjective-physiological concordance across time, and measures of arousal remain controversial. To investigate these issues, while healthy adults (n = 20) viewed emotional films, we assessed overall and continuous ratings of valence and arousal and recorded EMG from the corrugator supercilii and zygomatic major, EDA from the palms and forehead, and nose-tip temperature. The corrugator and zygomatic EMG were negatively and positively associated with valence ratings, respectively, across stimuli and time. EDA (both sites) and nose-tip temperature were positively and negatively associated with arousal ratings, respectively, across stimuli and time. It is concluded that subjective emotional valence and arousal dynamics have specific physiological correlates.
We present scalable Universal Dependency (UD) treebank synthesis techniques that exploit advances in language representation modeling which leverage vast amounts of unlabeled generalpurpose multilingual text. We introduce a data augmentation technique that uses synthetic treebanks to improve production-grade parsers. The synthetic treebanks are generated using a state-of-the-art biaffine parser adapted with pretrained Transformer models, such as Multilingual BERT (M-BERT). The new parser improves LAS by up to two points on seven languages. The production models' LAS performance improves as the augmented treebanks scale in size, surpassing performance of production models trained on originally annotated UD treebanks.
Syntactic parsing is a highly linguistic processing task whose parser requires training on treebanks from the expensive human annotation. As it is unlikely to obtain a treebank for every human language, in this work, we propose an effective cross-lingual UD parsing framework for transferring parser from only one source monolingual treebank to any other target languages without treebank available. To reach satisfactory parsing accuracy among quite different languages, we introduce two language modeling tasks into dependency parsing as multi-tasking. Assuming only unlabeled data from target languages plus the source treebank can be exploited together, we adopt a self-training strategy for further performance improvement in terms of our multi-task framework. Our proposed cross-lingual parsers are implemented for English, Chinese, and 22 UD treebanks. The empirical study shows that our cross-lingual parsers yield promising results for all target languages, for the first time, approaching the parser performance which is trained in its own target treebank.
The present article presents some challenges posed by lemmatization and PoS tagging of Latin, with reference to the ongoing work to revise the Latin Dependency Treebank. Current options available for lemmatization and morphological analysis of Latin are reviewed and discussed. The pipeline to annotate the morphological layer of the Latin Dependency Treebank is shown to consist of three main steps: (i) tokenization/sentence split, which is performed via a documented rule-based algorithm, (ii) pre-population by means of COMBO, a state-of-the-art joint lemmatizer, PoS tagger, and parser trained on the data of the Latin Dependency Treebank 2.1, and (iii) manual error correction informed by the attempt to identify and document lemmatization and morphology annotation rules.
Historical linguistics, whether synchronic or diachronic, is by definition based on corpora.Since we do not have access to the intuitions of native speakers we can only test linguistic hypotheses about historical languages by systematically collating information from our corpus of texts.For questions that typically concern linguists, this often means identifying every occurrence of a particular phenomenon in the corpus, analysing, classifying and counting the occurrences and then using this for testing hypotheses about the structure of the language.This can be done manually, but this is time-consuming and error-prone.As Haug (2015) points out, while reading the text and manually collating information from it is essential for hypothesis formation it is much less useful for hypothesis testing.Even if the text is in electronic form, it is easy to overlook an example, record it incorrectly or fail to apply test criteria consistently over time.This paper focuses on treebanks, which are corpora that have been annotated with morphosyntactic information so that we can extract linguistic structures like 'verb with an accusative noun'.High-quality treebanks for a range of historical languages now exist and are widely used in historical linguistic research.This includes treebanks that follow the Penn-style of annotation, e.g. the Penn-Helsinki
This paper reports on the analysis and annotation of Multiword Expressions in the Irish Universal Dependency Treebank. We provide a linguistic discussion around decisions on how to appropri- ately label Irish MWEs using the compound, flat and fixed dependency relation labels within the framework of the Universal Dependencies annotation guidelines. We discuss some nuances of the Irish language that pose challenges for assigning these UD labels and provide this report in support of the Irish UD annotation guidelines. With this we hope to ensure consistency in annotation across the dataset and provide a basis for future MWE annotation for Irish.
In this dataset, we introduce MS-TR a Morphologically Enriched Sentiment Treebank, which was implemented for training Recursive Deep Models to address compositional sentiment analysis in Turkish.
Two iconic twentieth-century print dictionaries provide a sampling of dictionary front matter in the Soviet–Russian lexicographic tradition; they demonstrate how front matter was used for overt and covert messaging about the linguistic norm. First is the one-volume bilingual English–Russian dictionary of Vladimir Karlovich Müller, and second is the one-volume monolingual Russian dictionary of Sergei Ivanovich Ozhegov; later it was published with Ozhegov and Nataliia Iul’evna Shvedova listed as co-authors. Both dictionaries were revised and reprinted over many decades. Only two editions of each dictionary are in focus, although there were more than sixty editions of Müller, 23 of Ozhegov, and several of Ozhegov and Shvedova.
Tae Hwan Oh, Ji Yoon Han, Hyonsu Choe, Seokwon Park, Han He, Jinho D. Choi, Na-Rae Han, Jena D. Hwang, Hansaem Kim. Proceedings of the 16th International Conference on Parsing Technologies and the IWPT 2020 Shared Task on Parsing into Enhanced Universal Dependencies. 2020.
We introduce a new symmetric measure (called pos ) that utilises the non-symmetric KL cpos 3 measure We can set a threshold for this new measure so that a pair of treebanks can be considered harmonious in their annotation if pos does not surpass the threshold. For the calculation of the threshold, we estimate the effects of (i) the size variation, and (ii) the genre variation in the considered pair of treebanks. The estimations are based on data from treebanks of distinct language families, making the threshold less dependent on the properties of individual languages. We demonstrate the utility of the proposed measure by listing the treebanks in Universal Dependencies version 2.5 (UDv2.5) (Zeman et al., 2019) data that are annotated consistently with other treebanks of the same language. However, the measure could be used to assess inter-treebank annotation consistency under other (non-UD) annotation guidelines as well.
Pretrained multilingual contextual representations have shown great success, but due to the limits of their pretraining data, their benefits do not apply equally to all language varieties. This presents a challenge for language varieties unfamiliar to these models, whose labeled and unlabeled data is too limited to train a monolingual model effectively. We propose the use of additional language-specific pretraining and vocabulary augmentation to adapt multilingual models to low-resource settings. Using dependency parsing of four diverse low-resource language varieties as a case study, we show that these methods significantly improve performance over baselines, especially in the lowestresource cases, and demonstrate the importance of the relationship between such models' pretraining data and target language varieties.
Text structuring is a fundamental step in NLG, especially when generating multi-sentential text. With the goal of fostering more general and data-driven approaches to text structuring, we propose the new and domain-independent NLG task of structuring and ordering a (possibly large) set of EDUs. We then present a solution for this task that combines neural dependency tree induction with pointer networks and can be trained on large discourse treebanks that have only recently become available. Further, we propose a new evaluation metric that is arguably more suitable for our new task compared to existing content ordering metrics. Finally, we empirically show that our approach outperforms competitive alternatives on the proposed measure and is equivalent in performance with respect to previously established measures.
Abstract In this paper, we present a novel lemmatization method based on a sequence-to-sequence neural network architecture and morphosyntactic context representation. In the proposed method, our context-sensitive lemmatizer generates the lemma one character at a time based on the surface form characters and its morphosyntactic features obtained from a morphological tagger. We argue that a sliding window context representation suffers from sparseness, while in majority of cases the morphosyntactic features of a word bring enough information to resolve lemma ambiguities while keeping the context representation dense and more practical for machine learning systems. Additionally, we study two different data augmentation methods utilizing autoencoder training and morphological transducers especially beneficial for low-resource languages. We evaluate our lemmatizer on 52 different languages and 76 different treebanks, showing that our system outperforms all latest baseline systems. Compared to the best overall baseline, UDPipe Future, our system outperforms it on 62 out of 76 treebanks reducing errors on average by 19% relative. The lemmatizer together with all trained models is made available as a part of the Turku-neural-parsing-pipeline under the Apache 2.0 license.
A recent advance in monolingual dependency parsing is the idea of a treebank embedding vector, which allows all treebanks for a particular language to be used as training data while at the same time allowing the model to prefer training data from one treebank over others and to select the preferred treebank at test time. We build on this idea by 1) introducing a method to predict a treebank vector for sentences that do not come from a treebank used in training, and 2) exploring what happens when we move away from predefined treebank embedding vectors during test time and instead devise tailored interpolations. We show that 1) there are interpolated vectors that are superior to the predefined ones, and 2) treebank vectors can be predicted with sufficient accuracy, for nine out of ten test languages, to match the performance of an oracle approach that knows the most suitable predefined treebank embedding for the test set.
International audience
Multiword Expression (MWE) has been a pain in the neck, especially in determining its word-classes in syntactic treebank. Previous work had proposed annotation guidelines for Indonesian MWEs that align to the Penn Treebank (PTB) format. However, we think that their proposed annotation still needs improvements. Therefore, this study proposes a new annotation guideline in labeling Indonesian MWE that conforms to PTB format. Moreover, we also revised the MWE annotation of an existing Indonesian constituency treebank consisting of 1030 sentences to conform to the new guidelines. To evaluate the revised treebank's quality, we built an Indonesian constituency parser model using the revised treebank and Stanford parser. The experiments show that the resulting parser has an F1-score of 69.97%.
The question whether a constitutive linguistic norm can be prescriptive is central to the debate on the normativity of meaning. Recently, the author has attempted to defend an affirmative answer, pointing to how speakers sporadically invoke constitutive linguistic norms in the service of linguistic calibration. Such invocations are clearly prescriptive. However, they are only appropriate if the invoked norms are applicable to the addressed speaker. But that can only be the case if the speaker herself generally accepts them. This qualification has led critics to argue that if an addressed speaker’s acceptance is a necessary condition for legitimate prescriptions (and reproach for failure to adhere to them), then the account becomes unable to underwrite actual normativity. Moreover, critics argue, a danger of vicious circularity arises from the calibration account. This paper shows that once a vantage point within the calibration practice is accepted, the criticisms lose their force. It then explores why a theorist might reject such a perspective and suggests, as a plausible candidate, implicit Humean assumptions about the proper explanation of (linguistic) action. The paper ends by sketching a way forward for the debate on the normativity of meaning in light of this diagnosis.