Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The article presents a new resource for A2-B2 learners of Lithuanian as L2 to improve their lexical competence and language production skills. The lexical database is a lexicographic application of the Lithuanian Pedagogic Corpus which was used both to develop headword lists and to collect word usage information. For this study, we adopt the inductive procedure of Corpus Pattern Analysis which was partly automated using the Lithuanian Sketch Grammar in Sketch Engine. We explain the model for pattern recognition and description, sense division, the selection of examples and give some details concerning the user interface.
This paper presents the foundations, procedures, tests and first results of a dependency treebank of the Spanish Sign Language (LSE). Dependency syntax offers many advantages over other alternatives for the systematic and exhaustive syntactic analysis of a corpus. Nevertheless, the visual modality that is characteristic of sign languages poses unique challenges for their syntactic analysis, among which the most prominent is the simultaneity of expression: both hands, face and other non-manual components. Taking into account these and other particularities of sign languages, the paper explores the main difficulties faced when one tries to apply some usual categories and relations from the syntactic analysis of spoken and written languages to LSE.
The study aims to shed light on the modern forms of linguistic dictionaries and show the influence of the Internet at the present time on them. These lexical changes that the Internet has caused since the late nineties must be understood primarily as a form of additional development of the dictionary format, which meets the needs and requirements of users and tries to adapt them to the principles of dictionaries, which in turn led to the emergence of a modern and very sophisticated dictionary of the Internet. This lexical product is often called the online dictionary, electronic dictionary, online database or lexical database. The results of the numerous survey conducted by the German Language Institute (IDS) in Mannheim show that searching for an Internet dictionary is much easier, more useful, and more exciting. It is interesting that users can make a significant contribution to entering and modifying existing words in the electronic dictionary and influencing the presence of vocabulary online through their desires or ideas or even through their comments. The search for traditional dictionaries is often replaced by electronic dictionaries and lexical databases. Therefore, the introduction of electronic dictionaries and the use of lexical databases in an educational curriculum is imperative and necessary, as it must be emphasized that user interaction is one of the important methodological foundations for the development of electronic تهدف الدراسة إلي تسليط الضوء علي الأشکال الحديثة للمعاجم اللغوية وإظهار تأثير الإنترنت في الوقت الحاضر عليها، حيث تُبرز الدراسة التأثيرات الکبيرة والعميقة للإنترنت علي علم المعاجم العريق. إن هذه المتغييرات المعجمية التي تسبب فيها الإنترنت منذ أواخر التسعينيات ، يجب أن تُفهم في المقام الأول على أنها شکل من أشکال التطوير الإضافي لشکل القاموس، يلبي احتياجات المستخدمين ومتطلباتهم ويحاول تطويعها لمباديء علم المعاجم، الأمر الذي أدى بدوره إلى ظهور معجم الإنترنت الحديث والمتطور للغاية. وغالبًا ما يطلق على هذا المنتج المعجمي معجم الإنترنت أو القاموس الإلکتروني أو قاعدة البيانات على الإنترنت أو قاعدة البيانات المعجمية. وتُظهر نتائج المسح العديدة التي أجراها معهد اللغة الألمانية في مدينة مانهايم أن البحث عن قاموس إنترنت أسهل بکثير وأکثر إفادة وأکثر إثارة. ومن المثير للاهتمام أن المستخدمين يمکنهم أن يسهموا مساهمة کبيرة في إدخال وتعديل الکلمات الحالية في المعجم الإلکتروني والتأثير في وجود المفردات عبر الإنترنت من خلال رغباتهم أو أفکارهم أو حتى من خلال تعليقاتهم. ومما لا شک فيه أنه يمکن إثبات وجود اختلافات کبيرة بين المعجم التقليدي والمعجم الرقمي أو المعجم العادي بمساعدة الکمبيوتر، لأن المعلومات في المعجم التقليدي تعتمد أساساً علي البحث الأبجدي عن المفردات، أما المعجم الإلکتروني فغالبا ما يتعلق باستراتيجيات البحث المختلفه التي يوضحها البحث. ومن المهم في هذا السياق التأکيد على أن القواميس والقواعد هي عنصر مهم في تعلم اللغات الأجنبية ، بغض النظر عن العمر ومستوى اللغة. ومع ذلک ، فقد لوحظ لبعض الوقت أن الأساليب التقليدية لتدريس قواعد اللغة والقواميس التقليدية لم تعد تتکيف مع دروس تعلم اللغات. وغالبا ما يتم استبدال البحث في القواميس التقليدية بالقواميس الإلکترونية وقواعد البيانات المعجمية. لذلک، أصبح إدخال القواميس الإلکترونية واستخدام قواعد البيانات المعجمية في منهج تعليمي أمرًا حتمياً وضروريًا، حيث أنه يجب التأکيد على أن تفاعل المستخدم هي واحدة من الأسس المنهجية الهامة لتطوير المعاجم الإلکترونية.
This dataset in CSV format contains all books from the web http://books.toscrape.com which has been got using web scraping method in November 2020. The CSV file has 12 columns called each of them like: title, image, rating, description, category, UPC, producttype, priceextax, priceincltax, tax, availability, numberreviews. The project was born as a practice for a subject of the Master of Science (MSc) of Data Science at the Universitat Oberta de Catalunya (UOC).
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages.
OBJECTIVE: We investigated the relations between psychopathic traits and fear enjoyment. METHOD: In Study 1, 140 undergraduate participants (62 men, 78 women) watched the footage of video game play meant to induce either excitement or fear, rating each on positive/negative adjectives. In Study 2, 150 undergraduate participants (94 women, 56 men) rated valence (positive/negative) of 20 sets of morphed surprise/fear photos. RESULTS: In Study 1, participants with higher levels of psychopathy rated the fear video as less negative and more positive. In Study 2, valence ratings became more negative as fear information increased (fear-laden faces were rated more negatively than surprise-laden faces). As well, there were significant interactions between psychopathy and morph level in predicting valence with psychopathic traits being associated with giving higher positivity ratings to fear-laden faces. CONCLUSIONS: The results of these two studies suggest that people with psychopathic traits have a more positive interpretation of the experience of fear, which could extend to evaluations of others' experiences of fear.
We present our contribution to the EvaLatin shared task, which is the first\nevaluation campaign devoted to the evaluation of NLP tools for Latin. We\nsubmitted a system based on UDPipe 2.0, one of the winners of the CoNLL 2018\nShared Task, The 2018 Shared Task on Extrinsic Parser Evaluation and SIGMORPHON\n2019 Shared Task. Our system places first by a wide margin both in\nlemmatization and POS tagging in the open modality, where additional supervised\ndata is allowed, in which case we utilize all Universal Dependency Latin\ntreebanks. In the closed modality, where only the EvaLatin training data is\nallowed, our system achieves the best performance in lemmatization and in\nclassical subtask of POS tagging, while reaching second place in cross-genre\nand cross-time settings. In the ablation experiments, we also evaluate the\ninfluence of BERT and XLM-RoBERTa contextualized embeddings, and the treebank\nencodings of the different flavors of Latin treebanks.\n
Foreign language education primarily aims to cultivate learners’ competence to communicate in an additional language. However, the meaning of communication competence is not entirely transparent, especially given the current neoliberal valorization of communication in the knowledge economy. The meaning of communication can be scrutinized in two contradictory trends observed in language education: the exclusive focus on teaching English as a global language, signifying a homogenizing trend, and increased scholarly attention to the heterogeneity of linguistic forms and practices. This article examines how communication competence is differentially understood by policymakers and corporate workers in Japan. The authors examine a government report that evaluated the attainment of educational goals for coping with globalization and contrasting it with interview data drawn from another study on the communicative experiences of Japanese transnational workers in Asia. Political discourse analysis and content analysis reveal the paradoxical nature of what can be called neoliberal communication competence, which on the one hand conflates global communication with use of the four measurable skills in English to transmit information and, on the other hand, challenges linguistic norms, foregrounding plurilingualism and co‐constructed interactional competence. Transformation of policies and pedagogies can be pursued by appropriating neoliberal communication competence for achieving broader educational goals.
Implicit relation classification on Penn Discourse TreeBank (PDTB) 2.0 is a common benchmark task for evaluating the understanding of discourse relations. However, the lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in the literature. In this work, we highlight these inconsistencies and propose an improved evaluation protocol. Paired with this protocol, we report strong baseline results from pretrained sentence encoders, which set the new state-of-the-art for PDTB 2.0. Furthermore, this work is the first to explore fine-grained relation classification on PDTB 3.0. We expect our work to serve as a point of comparison for future work, and also as an initiative to discuss models of larger context and possible data augmentations for downstream transferability.
Both syntactic and semantic structures are key linguistic contextual clues, in which parsing the latter has been well shown beneficial from parsing the former. However, few works ever made an attempt to let semantic parsing help syntactic parsing. As linguistic representation formalisms, both syntax and semantics may be represented in either span (constituent/phrase) or dependency, on both of which joint learning was also seldom explored. In this paper, we propose a novel joint model of syntactic and semantic parsing on both span and dependency representations, which incorporates syntactic information effectively in the encoder of neural network and benefits from two representation formalisms in a uniform way. The experiments show that semantics and syntax can benefit each other by optimizing joint objectives. Our single model achieves new state-of-the-art or competitive results on both span and dependency semantic parsing on Propbank benchmarks and both dependency and constituent syntactic parsing on Penn Treebank.
Recommender systems often involve multi-aspect factors. For example, when shopping for shoes online, consumers usually look through their images, ratings, and product's reviews before making their decisions. To learn multi-aspect factors, many context-aware models have been developed based on tensor factorizations. However, existing models assume multilinear structures in the tensor data, thus failing to capture nonlinear feature interactions. To fill this gap, we propose a novel nonlinear tensor machine, which combines deep neural networks and tensor algebra to capture nonlinear interactions among multi-aspect factors. We further consider adversarial learning to assist the training of our model. Extensive experiments demonstrate the effectiveness of the proposed model.
In this paper, we propose MCNN-ReMGU model based on multi-window convolution and residual-connected minimal gated unit (MGU) network for the natural language word prediction. First, the convolution kernels with different sizes are used to extract the local feature information of different graininess between the word sequences. Then, the extracted features are fed to the residual-connected MGU network. Finally, the prediction results are output by the SoftMax layer. Through the residual-connection processing of MGU network in the model, not only the problems of vanishing gradient and network degradation are effectively solved, but also the long-term dependence between word sequences is effectively extracted to predict the next word accurately. Meanwhile, the introduction of the convolution kernel in a convolutional neural network (CNN) enables the feature information between word sequences to be extracted more fully. The experimental results on the Penn Treebank and WikiText-2 datasets show that the proposed method has certain advantages in the word prediction task.
We propose a transition-based approach that, by training a single model, can efficiently parse any input sentence with both constituent and dependency trees, supporting both continuous/projective and discontinuous/non-projective syntactic structures. To that end, we develop a Pointer Network architecture with two separate task-specific decoders and a common encoder, and follow a multitask learning strategy to jointly train them. The resulting quadratic system, not only becomes the first parser that can jointly produce both unrestricted constituent and dependency trees from a single model, but also proves that both syntactic formalisms can benefit from each other during training, achieving state-of-the-art accuracies in several widely-used benchmarks such as the continuous English and Chinese Penn Treebanks, as well as the discontinuous German NEGRA and TIGER datasets.
Even though Automatic Speech Recognition (ASR) systems significantly improved over the last decade, they still introduce a lot of errors when they transcribe voice to text. One of the most common reasons for these errors is phonetic confusion between similar-sounding expressions. As a result, ASR transcriptions often contain "quasi-oronyms", i.e., words or phrases that sound similar to the source ones, but that have completely different semantics (e.g., "win" instead of "when" or "accessible on defecting" instead of "accessible and affecting"). These errors significantly affect the performance of downstream Natural Language Understanding (NLU) models (e.g., intent classification, slot filling, etc.) and impair user experience. To make NLU models more robust to such errors, we propose novel phonetic-aware text representations. Specifically, we represent ASR transcriptions at the phoneme level, aiming to capture pronunciation similarities, which are typically neglected in word-level representations (e.g., word embeddings). To train and evaluate our phoneme representations, we generate noisy ASR transcriptions of four existing datasets - Stanford Sentiment Treebank, SQuAD, TREC Question Classification and Subjectivity Analysis - and show that common neural network architectures exploiting the proposed phoneme representations can effectively handle noisy transcriptions and significantly outperform state-of-the-art baselines. Finally, we confirm these results by testing our models on real utterances spoken to the Alexa virtual assistant.
本研究依据以谓词为核心的块依存语法构建块依存树库,在句内和句间寻找谓词所支配的组块,利用汉语中组块和组块间的依存关系补全缺省部分,明确谓词支配关系。目前共标注2199篇文本,涵盖百科、新闻两个领域,共约187万字语料。本文简述了块依存语法的原则,并对组块及其依存关系进行了定义。将详细介绍标注流程、标注一致率、数据分布等情况。基于现有的树库,本研究发现汉语中有约25%的小句是非自足的,约有88%的核心谓词可支配1~3个从属成分。
Arabic diacritics play a significant role in distinguishing words with the same orthography but different meanings, pronunciations, and syntactic functions. The presence of Arabic diacritics can be useful in many natural language processing applications, such as text-to-speech tasks, machine translation, and part-of-speech tagging. This article discusses the use of bidirectional long short-term memory neural networks with conditional random fields for Arabic diacritization. This approach requires no morphological analyzers, dictionary, or feature engineering, but rather uses a sequence-to-sequence schema. The input is a sequence of characters that constitute the sentence, and the output consists of the corresponding diacritic(s) for each character in that sentence. The performance of the proposed approach was examined using four datasets with different sizes and genres, namely, the King Abdulaziz City for Science and Technology text-to-speech (KACST TTS) dataset, the Holy Quran, Sahih Al-Bukhary, and the Penn Arabic Treebank (ATB). For training, 60% of the sentences were randomly selected from each dataset, 20% were selected for validation, and 20% were selected for testing. The trained models achieved diacritic error rates of 3.41%, 1.34%, 1.57%, and 2.13% and word error rates of 14.46%, 4.92%, 5.65%, and 8.43% on the KACST TTS, Holy Quran, Sahih Al-Bukhary, and ATB datasets, respectively. Comparison of the proposed method with those used in other studies and existing systems revealed that its results are comparable to or better than those of the state-of-the-art methods.
This chapter provides an overview of the contribution the Prague School of Linguistics has made to the study of linguistic norm. Departing from a functionalist approach to the analysis of language, the Prague School shaped and theorized the influential notions of language culture (jazykova kultura, Sprachkultur) and language cultivation (Sprachpflege) at a time when, in Western Europe, prescriptive approaches to language were considered unworthy of scientific attention. Distancing themselves from "purist" and "school-grammar" conceptions of language maintenance and based on the case of Czech language culture, the representatives of the Prague School advocated for the cultivation of the written (standard) language in functional terms in order to achieve a "stable" yet "elastic" concept of norm.
Dual language immersion (DLI) programs have emerged in the U.S. as effective ways to bring together language minority and language majority speakers in school settings with the goal of bilingualism and bi-literacy for all. However, the proliferation of these programs has raised concerns regarding issues of inequity and dissimilar power dynamics in these spaces (Cervantes-Soon, 2014 Cervantes-Soon, C. G. 2014. “A Critical Look at Dual Language Immersion in the New Latin@ Diaspora.” Bilingual Research Journal 37 (1): 64–82.[Taylor & Francis Online], [Google Scholar], “A Critical Look at Dual Language Immersion in the New Latin@ Diaspora.” Bilingual Research Journal 37 (1): 64–82; Flores, 2016, Do Black Lives matter in Bilingual Education [Web log post]. Accessed May 1, 2017. https://educationallinguist.wordpress.com/2016/09/11/do-black-lives-matter-in-bilingual-education/; Valdes, 1997, “Dual language immersion programs: A cautionary note concerning the education of language-minority students.” Harvard Educational Review 67: 391–430, 2018, “Analyzing the curricularization of language in two-way immersion education: Restating two cautionary notes.” Bilingual Research Journal). With this in mind, this study aims to shed light on the intricate social processes at work in DLI contexts. In particular, this paper examines first, how notions of language use, race, and ethnicity are socially constructed and intersect in DLI settings; and second, it explores how these ideas are discerned and re-shaped by young children into their own social and linguistic norms. Employing qualitative research methods, this year-long ethnographic case study uses the intersectional lens of raciolinguistics (Alim, Rickford & Ball, 2016 Alim, H. S., J. R. Rickford, and A. F. Ball, eds. 2016. Raciolinguistics: how Language Shapes our Ideas About Race. New York, NY: Oxford University Press.[Crossref], [Google Scholar], Raciolinguistics: how language shapes our ideas about race. New York, NY: Oxford University Press; Rosa & Flores, 2017, “Unsettling race and language: Toward a raciolinguistic perspective.” Language in Society 46 (5): 621–647), to examine the intricate cross-cutting dynamics at play in bilingual spaces. The exploration of these ideas helps to illuminate the ways in which language practices and interactions are shaped by social constructions from a very early age. Furthermore, it contributes to understandings of social perceptions and relations in multilingual/multicultural/multiethnic contemporary school settings.
Delivery of best-practice care for posttraumatic stress disorder (PTSD) is a priority for clinicians working with active duty military personnel and veterans. The PTSD Clinicians Exchange, an Internet-based intervention, was designed to assist in disseminating clinically relevant information and resources that support delivery of key practices endorsed in the Veterans Administration (VA)-Department of Defense (DoD) Clinical Practice Guidelines (CPG) for the Management of Posttraumatic Stress. We conducted a randomized controlled trial to examine the effectiveness of the Clinicians Exchange intervention in increasing familiarity and perceived benefits of 26 CPG-related and emerging practices. The intervention consisted of ongoing access to an Internet resource featuring best-in-class resources for practices, self-management of burnout, and biweekly e-mail reminders highlighting selected practices. Mental health clinicians (N = 605) were recruited from three service sectors (VA, DoD, community); 32.7% of participants assigned to the Internet intervention accessed the site to view resources. Individuals who were offered the intervention increased their practice familiarity ratings significantly more than those assigned to a newsletter-only control condition, d = 0.27, p =.005. From baseline to 12-months, mean familiarity ratings of clinicians in the intervention group increased from 3.0 to 3.4 on scale of 1 (not at all) to 5 (extremely); mean ratings for the control group were 3.2 at both assessments. Clinicians generally viewed the CPG practices favorably, rating them as likely to benefit their clients. The results suggest that Internet-based resources may aid more comprehensive efforts to disseminate CPGs, but increasing clinician engagement will be important.
The Italian Sign Language (LIS) is the natural language used by the Italian Deaf community. This paper discusses the application of the Universal Dependencies (UD) format to the syntactic annotation of a LIS corpus. This investigation aims in particular at contributing to sign language research by addressing the challenges that the visual-manual modality of LIS creates generally in linguistic annotation and specifically in segmentation and syntactic analysis. We addressed two case studies from the storytelling domain first segmented on the ELAN platform, and second syntactically annotated using CoNLL-U format.
Cupping therapy has recently gained public attention and is widely used in many regions. Some patients are resistant to being treated with cupping therapy, as visually unpleasant marks on the skin may elicit negative reactions. This study aimed to identify the cognitive and emotional components of cupping therapy. Twenty-five healthy volunteers were presented with emotionally evocative visual stimuli representing fear, disgust, happiness, neutral emotion, and cupping, along with control images. Participants evaluated the valence and arousal level of each stimulus. Before the experiment, they completed the Fear of Pain Questionnaire-III. In two-dimensional affective space, emotional arousal increases as hedonic valence ratings become increasingly pleasant or unpleasant. Cupping therapy images were more unpleasant and more arousing than the control images. Cluster analysis showed that the response to cupping therapy images had emotional characteristics similar to those for fear images. Individuals with a greater fear of pain rated cupping therapy images as more unpleasant and more arousing. Psychophysical analysis showed that individuals experienced unpleasant and aroused emotional states in response to the cupping therapy images. Our findings suggest that cupping therapy might be associated with unpleasant-defensive motivation and motivational activation. Determining the emotional components of cupping therapy would help clinicians and researchers to understand the intrinsic effects of cupping therapy.
As the number, size, and complexity of building construction projects increase, code compliance checking becomes more challenging because of the time-consuming, costly, and error-prone nature of a manual checking process. A fully automated code compliance checking would be desirable in facilitating a more efficient, cost effective, and human error-proof code checking. Such automation requires automated information extraction from building designs and building codes, and automated information transformation to a format that allows automated reasoning. Natural language processing (NLP) is an important technology to support such automated processing of building codes, because building codes are represented in natural language texts. Part-of-speech (POS) tagging, as an important basis of NLP tasks, must have a high performance to ensure the quality of the automated processing of building codes in such a compliance checking system. However, no systematic testing of existing POS taggers on domain specific building codes data have been performed. To address this gap, the authors analyzed the performance of seven state-of-the-at POS taggers on tagging building codes and compared their results to a manually-labeled gold standard. The authors aim to: (1) find the best performing tagger in terms of accuracy, and (2) identify common sources of errors. In providing the POS tags, the authors used the Penn Treebank tagset, which is a widely used tagset with a proper balance between conciseness and information richness. An average accuracy of 88.80% was found on the testing data. The Standford coreNLP tagger outperformed the other taggers in the experiment. Common sources of errors were identified to be: (1) word ambiguity, (2) rare words, and (3) unique meaning of common English words in the construction context. The found result of machine taggers on building codes calls for performance improvement, such as error-fixing transformational rules and machine taggers that are trained on building codes.
Noun phrases convey key information in communication and are of interest in NLP tasks. A base NP is defined as the headword and left-hand side modifiers of a noun phrase. In this thesis, we identify base NPs in Universal Dependencies treebanks in English and French using an RNN architecture.The data of this thesis consist of three multi-layered treebanks in which each sentence is annotated in both constituency and dependency formalisms. To build our training data, we find base NPs in the constituency layers and project them onto the dependency layer by labeling corresponding tokens. For input features, we devised 18 configurations of features available in UD annotation. We train RNN models with LSTM and GRU cells with different numbers of epochs on these configurations of features.Tested on monolingual and bilingual test sets, our models delivered satisfactory token-based F1 scores (92.70% on English, 94.87% on French, 94.29% on bilingual test set). The most predicative configuration of features is found out to be pos_dep_parent_child_morph, which covers 1) dependency relations between the current token, its syntactic head, its leftmost and rightmost syntactic dependents; 2) PoS tags of these tokens; and 3) morphological features of the current token.
In order to extract the semantic and grammatical information of sentences more effectively, this paper proposes a sentence sentiment classification method based on Self-supervised and Self-attention mechanism (SS-SAtt-BiLSTM). In this method, BiLSTM network is used to extract the feature of text context relationship, and self-supervised (SS) learning mode is introduced into the supervised sentence representation model. The sentence itself is used as the label data information of current words, and an improved self-attention mechanism (SA) is used to calculate the attention weight of each moment. The experimental results of MR and Stanford sentient treebank (sst-5) data sets show that this method reduces the dependence on tagged data, and the improved self-attention mechanism enables the model to learn more key features of sentences and improve the classification performance.
Affective responses to music have been shown to be influenced by the psychoacoustic features of the acoustic signal, learned associations between musical features and emotions, and familiarity with a musical system through exposure. The present article reports two experiments investigating whether short-term exposure has an effect on valence and consonance ratings of unfamiliar musical chords from the Bohlen-Pierce system, which are not based on a traditional Western musical scale. In a pre- and post-test design, exposure to positive, negative and neutral chord types was manipulated to test for an effect of exposure on liking. In this paradigm, short-term (“mere”) exposure to unfamiliar chords produced an increase only in valence ratings for negative chords. In neither experiment did it produce an increase in valence or pleasantness ratings for other chord types. Contrast effects for some chord types were found in both experiments, suggesting that a chord’s affect (i.e., affective response to the chord) might be emphasised when the chord is preceded by a stimulus with a contrasting affect. The results confirmed those of a previous study showing that psychoacoustic features play an important role in the perception of music. The findings are discussed in light of their psychological and musical implications.
Semantic Role Labelling (SRL) is the process of automatically finding the semantic roles of terms in a sentence. It is an essential task towards creating a machine-meaningful representation of textual information. One public linguistic resource commonly used for this task is the FrameNet Project. FrameNet is a human and machine-readable lexical database containing a considerable number of annotated sentences, those annotations link sentence fragments to semantic frames. However, while the annotations across all the documents covered in the dataset link to most of the frames, a large group of frames lack annotations in the documents pointing to them. In this paper, we present a data augmentation method for FrameNet documents that increases by over 13% the total number of annotations. Our approach relies on lexical, syntactic, and semantic aspects of the sentences to provide additional annotations. We evaluate the proposed augmentation method by comparing the performance of a state-of-the-art semantic-role-labelling system, trained using a dataset with and without augmentation.
Introduction. High-quality language education in technical universities requires its interdisciplinary relation to the content of highly specialised subjects corresponding to the training programmes aimed at instructing the future specialists. Educational materials in a foreign language are highly productive if they emphasise the terminology and professional vocabulary authentic to the current state of the scientific field. The aim of the study presented in the article was to assess the validity of the lexical material delivered in the course “English for Business Communication”, to determine the selection criteria for this vocabulary as well as the methods for its assimilation and practical application. Methodology and research methods. The applied corpus software enabled to obtain quantitative indicators of the distribution of foreign-language business vocabulary in the given training course. The lexical material being currently offered to students and the professional thesaurus identified via linguistic databases was compared with the use of comparative analysis and synthesis. Results and scientific novelty. The lexical units (terms, set expressions), which are the most active in the business sphere, were identified on the basis of its frequency. The authors established the correlation between them and educational vocabulary, both from the perspective of its integration into the course without block concentration throughout the course of university training, and from the perspective of the variety of methods used to practice this vocabulary. It is concluded that the applied educational material needs to be substantially adjusted. The vocabulary does not completely reflect the realities of the business communication sphere and the distribution of active vocational vocabulary regulated by methodological guidelines does not entirely contribute to its strong assimilation. According to the authors, the necessary changes to the approaches and methods for selecting and compiling lexical material and to the methodology for designing a foreign language course should be made on the basis of integrating pedagogical and linguistic knowledge, in particular, the methodology of teaching foreign languages and the corpus linguistics. Practical significance. The ways of integrating corpus programs in the process of developing the content of language disciplines, which are part of the main educational program of technical universities, are demonstrated as one of the methods to increase the effectiveness of teaching foreign languages to students of non-linguistic specialties.
Facebook has recently gained popularity among young, digitally literate and predominantly urban Pakistanis. Such social networking sites allow users the freedom to express themselves using usernames, visuals and topics of their own choice. In this article, I examine how Pakistani Facebook users mobilize such resources in their identity work. Using Multimodal Discourse Analysis, I investigate how Pakistani women construct their gender identities on Facebook using visual and linguistic resources. The results revealed the significant impact of Facebook on the socio-cultural and linguistic norms of discourse in Pakistan that enables women to challenge established communication models while they simultaneously reinforce traditional gender models.
This paper presents the structure of the LiLa Knowledge Base, i.e. a collection of multifarious linguistic resources for Latin described with the same vocabulary of knowledge description and interlinked according to the principles of the so-called Linked Data paradigm. Following its highly lexically based nature, the core of the LiLa Knowledge Base consists of a large collection of Latin lemmas, serving as the backbone to achieve interoperability between the resources, by linking all those entries in lexical resources and tokens in corpora that point to the same lemma. After detailing the architecture supporting LiLa, the paper particularly focusses on how we approach the challenges raised by harmonizing different strategies of lemmatization that can be found in linguistic resources for Latin. As an example of the process to connect a linguistic resource to LiLa, the inclusion in the Knowledge Base of a dependency treebank is described and evaluated.
Language users and learners are sensitive to distributional information in their environment, which enables them to extract regularities that occur in the language input that they are exposed to. This process is referred to as statistical learning. While the statistical learning phonotactic literature thoroughly investigates the learning of overall phonotactics in specific languages, little is known about cases where different phonological systems coexist within a single language. The Japanese lexicon is generally classified into four lexical strata according to the etymological status of each word (Itô & Mester, 1995, 1999, 2001). Although each stratum includes the internal phonological similarity in the Japanese language as a whole, there are also distinctive phonological properties. A recent study suggests that language users should be able to learn phonotactics of each sublexicon based on the same kind of statistical probabilities that computers analyse from language users’ accumulated lexicons (Morita, 2018). This thesis examines whether second-language (L2) learners can learn the loanword phonotactics/phonology of Japanese through experience of using and/or passive exposure to Japanese lexical stratification. Using two loanword phonological regularities (categorical and gradient rules) as a case study, two fully-crossed perceptual experiments involving English- speaking learners of Japanese, native speakers of Japanese, and English-speaking monolinguals are presented. The first experiment explores listeners’ phonotactic/phonological knowledge of nativised loanwords in Japanese using a well-formedness task which shows the adaptation of English final consonants in monosyllabic words. Listeners judge whether the pronunciation they hear is how the word would be pronounced if it was a Japanese word, rating how confident they are on a scale of 1-5. This study shows that L2 learners learn categorical rules, but not gradient patterns. This study also confirms that loanword phonotactics and overall phonotactics make separate contributions to perceived well-formedness. L2 learners access and make use of the sublexicon-specific probabilities of Japanese during the task. The second perceptual experiment is designed to support the findings in the first experiment, by testing for discrimination of non-native consonantal contrasts. Even under high memory demand, L2 learners show the ability to discriminate non-native consonantal contrasts (i.e., CVCV/CVCCV) effectively enough to support findings in the first experiment. These results suggest that L2 learners can implicitly detect the statistical structure of a language’s sublexicon phonology over the course of acquiring a natural language. However, while native speakers of Japanese learn a gradient rule, L2 learners of Japanese do not. A potential explanation for the differences in gradient rule learning is that the vocabulary size of the target language might play a crucial role. This remains an open question. In addition, the present work provides a basis for future investigation into whether L2 learners of Japanese, whose native language is other than English, are able to learn Japanese loanword phonotactics/phonology. L1 English-L2 Japanese speakers might gain advantage in perceiving the English input which inevitably overlaps with the phonological form of the host language.
The goal of this special issue of Critical Multilingualism Studies “National Standards – Local Varieties: A Cross-Linguistic Discussion on Regional Variation in L2 Studies” is to incite a conversation on how topics such as linguistic norms and variation, dominant practices, ideologies, identities, and politics surrounding languages are discussed from a view outside of the dominant centers of linguistic norms.
Because of its focus on the past and on historical languages, the classics is a discipline that is particularly interested in translations and text alignment. Starting from a diachronic perspective, this contribution demonstrates how issues related to text alignment, present since antiquity, can be approached from a different angle and with entirely new opportunities thank to tools and methods developed in the field of digital humanities. By comparing examples from antiquity (e.g. Origen’s Hexapla from the third century CE) with modern projects based on treebanking and dependency grammar (e.g. the Ancient Greek and Latin Dependency Treebank [AGLDT] as part of the Perseus Digital Library from Tufts University), we shall present some new approaches and their potentials. In doing so, we shall also examine what status English has in these projects and how the different languages involved in each of them interact with English and/or with each other.
We study the effect of rich supertag features in greedy transition-based dependency parsing. While previous studies have shown that sparse boolean features representing the 1-best supertag of a word can improve parsing accuracy, we show that we can get further improvements by adding a continuous vector representation of the entire supertag distribution for a word. In this way, we achieve the best results for greedy transition-based parsing with supertag features with $88.6\%$ LAS and $90.9\%$ UASon the English Penn Treebank converted to Stanford Dependencies.
This is an introduction to the proposed theme, in which the importance of sociolinguistic studies for the teaching, acquisition and learning of languages is emphasized. In addition, each text of the material is presented, starting with interviews with significant and current representatives of the variation sociolinguistics (Francisco Moreno Fernández and Juan Manuel Hernández Campoy) from the Hispanic and Anglo-Saxon spheres, respectively; then, it discusses the ten articles that deal with the theme from two perspectives: linguistic attitudes and beliefs of speakers and linguistic norms and policies. Finally, the reviews of two books related to the Special issue are commented: The Routledge handbook of Spanish as a heritage language, edited by Kim Potowsky, 2018, New York, Routledge publisher, and La trastienda de la enseñanza de lenguas extranjeras, by Francisco García Marcos, 2018, from the Interlingua collection of Editora Comares de Granada / Spain. The presentation is an invitation to readers to enjoy reading the Special issue.
This paper analyzes linguistic deficit discourse as it emerges in language gap research, gets appropriated by language gap foundations, and is reported in the media. Through intertextual analysis, we show how language deficit ideologies combine with neoliberal logic to normalize the marginalization of minoritized families, linguistic and sociolinguistic hierarchies, and the privileging of White middle-class (socio)linguistic norms. Language gap discourse turns parents into scapegoats by blaming them for the linguistic deficiencies of their children and low-income families are encouraged to misrecognize the inherent value of their communication abilities. IN the process, social processes that engender economic and educational inequality are obfuscated. Rather than attempting to find real answers to real problems, language gap discourse emphasizes a quick fix solution (filling your kids up with words) instead of engaging with the real causes of educational inequity.
We propose the Graph2Graph Transformer architecture for conditioning on and predicting arbitrary graphs, and apply it to the challenging task of transition-based dependency parsing. After proposing two novel Transformer models of transition-based dependency parsing as strong baselines, we show that adding the proposed mechanisms for conditioning on and predicting graphs of Graph2Graph Transformer results in significant improvements, both with and without BERT pre-training. The novel baselines and their integration with Graph2Graph Transformer significantly outperform the state-of-the-art in traditional transition-based dependency parsing on both English Penn Treebank, and 13 languages of Universal Dependencies Treebanks.
HDT-UD, the largest German UD treebank by a large margin, as well as the German-LIT treebank, currently do not analyze preposition-determiner contractions such as zum (= zu dem, “to the”) as multi-word tokens, which is inconsistent both with UD guidelines as well as other German UD corpora (GSD and PUD). In this paper, we show that harmonizing corpora with regard to this highly frequent phenomenon using a lookup-table based approach leads to a considerable increase in automatic parsing performance.
The present study investigates the relationship between two features of dependencies, namely, dependency distances and dependency frequencies. The study is based on the analysis of a parallel dependency treebank that includes 10 Indo-European languages. Two corresponding random dependency treebanks are generated as baselines for comparison. After computing the values of dependency distances and their frequencies in these treebanks, for each lan-guage, we fit four functions, namely quadratic, exponent, logarithm, and power-law func-tions, to its original and random datasets. The preliminary result shows that there is a rela-tion between the two dependency features for all 10 Indo-European languages. The relation can be further formalized as a power-law function which can distinguish the observed data from randomly generated datasets.
We propose a method for unsupervised parsing based on the linguistic notion of a constituency test. One type of constituency test involves modifying the sentence via some transformation (e.g. replacing the span with a pronoun) and then judging the result (e.g. checking if it is grammatical). Motivated by this idea, we design an unsupervised parser by specifying a set of transformations and using an unsupervised neural acceptability model to make grammaticality decisions. To produce a tree given a sentence, we score each span by aggregating its constituency test judgments, and we choose the binary tree with the highest total score. While this approach already achieves performance in the range of current methods, we further improve accuracy by fine-tuning the grammaticality model through a refinement procedure, where we alternate between improving the estimated trees and improving the grammaticality model. The refined model achieves 62.8 F1 on the Penn Treebank test set, an absolute improvement of 7.6 points over the previous best published result.
With the rapid development of big data and deep learning, breakthroughs have been made in phonetic and textual research, the two fundamental attributes of language. Language is an essential medium of information exchange in teaching activity. The aim is to promote the transformation of the training mode and content of translation major and the application of the translation service industry in various fields. Based on previous research, the SCN-LSTM (Skip Convolutional Network and Long Short Term Memory) translation model of deep learning neural network is constructed by learning and training the real dataset and the public PTB (Penn Treebank Dataset). The feasibility of the model's performance, translation quality, and adaptability in practical teaching is analyzed to provide a theoretical basis for the research and application of the SCN-LSTM translation model in English teaching. The results show that the capability of the neural network for translation teaching is nearly one times higher than that of the traditional N-tuple translation model, and the fusion model performs much better than the single model, translation quality, and teaching effect. To be specific, the accuracy of the SCN-LSTM translation model based on deep learning neural network is 95.21%, the degree of translation confusion is reduced by 39.21% compared with that of the LSTM (Long Short Term Memory) model, and the adaptability is 0.4 times that of the N-tuple model. With the highest level of satisfaction in practical teaching evaluation, the SCN-LSTM translation model has achieved a favorable effect on the translation teaching of the English major. In summary, the performance and quality of the translation model are improved significantly by learning the language characteristics in translations by teachers and students, providing ideas for applying machine translation in professional translation teaching.
Berkeley Neural Parser model for Korean using the Sejong treebank Berkeley Neural Parser https://github.com/nikitakit/self-attentive-parser COLLINS-SJTREE.prm for evalb is available at http://doi.org/10.5281/zenodo.1004604