Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Traumatic experiences are associated with increased emotional arousal. Overnight consolidation strengthens the episodic content of emotional memories, but it is still unclear how sleep influences the associated arousal response. To investigate this question, we compared the effects of sleep and wake on psychophysiological and subjective reactivity during emotional memory retrieval. Participants provided affective ratings for negative and neutral images while heart rate deceleration (HRD) and skin conductance responses (SCRs) were monitored. Following a 12-hour delay of sleep or wakefulness, participants completed an image recognition task where HRD, SCRs and affective ratings were recorded again. HRD responses to previously-encoded (“old”) negative images were preserved after sleep but diminished after wakefulness. No between-group difference in HRD was observed for novel negative images at recognition, indicating that the effects of sleep for old images were not driven by a generalised overnight increase in visceral activity, or circadian factors. No significant effects of sleep were observed for SCRs or subjective ratings. Our data suggest that cardiac arousal experienced at the time of encoding is sensitive to plasticity-promoting processes during sleep in a similar manner to episodic aspects of emotional memory.
This paper closely analysed and depicted the role of the hyponymy that is used in the lexical database called the Word Net to relate the words semantically with a relation... Is a kind of.... The relation has very important role for disambiguating the senses of a polysemy word in any natural language and is, therefore, massively used in knowledge-based Word Sense Disambiguation (WSD) approaches. However, the limitation of this relation is not so cared about. In many cases, the use of the hypernymy does not only mean nothing for disambiguation but also it increases only the computational effort for the system. The improper use of hypernymy is very costly. In some cases, which are described later, the use of this relation even decreases the accuracy of the WSD approaches.
This paper focuses on parsing processes and principles, which are essential tasks for machines to understand syntactic and semantic structures of a sentence. Machine analysis procedures for Chinese sentences, including word segmentation, and part-of-speech tagging and parsing, were visually represented using Cparser, a rule-based constituency parser developed by Peking University. Then we explained how linguistic knowledge was embodied in Cparser lexical, syntactic and semantic components, and discussed their complex interplay that allows automatic parsing. As a practical example, a Chinese textbook treebank is also constructed using Cparser. According to the theoretical and practical discussion in this paper, Peking University Cparser can easily to reflect and modify undated linguistic knowledge and is expected to be widely used as an analysis and verification tool for Chinese grammar research.
This article discusses the modes of defining cultural content in teaching Polish as a foreign language. It proposes strategies for introducing cultural knowledge and knowledge about culture in the practice of teaching Polish as a non-native language. It seems appropriate to offer a division into:
 
 The cultural semantics of lexemes, set phrases and expressions;
 The culture (meaning the correct use) of expression and the linguistic norm; and
 Culture content and areas conveyed to foreigners through the language they are learning.
 
 This discussion fits the debate regarding the language which we teach to foreigners. In it, researchers have offered examples of semantic distortions (mostly shifts), and the dangers of private or rather extra-curricular linguistic contact with a language’s native users. It also carries reflections on linguistic regionalisation and its presence in the practice of teaching Polish as a non-native language.
This is an introduction to the proposed theme, in which the importance of sociolinguistic studies for the teaching, acquisition and learning of languages is emphasized. In addition, each text of the material is presented, starting with interviews with significant and current representatives of the variation sociolinguistics (Francisco Moreno Fernández and Juan Manuel Hernández Campoy) from the Hispanic and Anglo-Saxon spheres, respectively; then, it discusses the ten articles that deal with the theme from two perspectives: linguistic attitudes and beliefs of speakers and linguistic norms and policies. Finally, the reviews of two books related to the Special issue are commented: The Routledge handbook of Spanish as a heritage language, edited by Kim Potowsky, 2018, New York, Routledge publisher, and La trastienda de la enseñanza de lenguas extranjeras, by Francisco García Marcos, 2018, from the Interlingua collection of Editora Comares de Granada / Spain. The presentation is an invitation to readers to enjoy reading the Special issue.
RST-based discourse parsing is an important NLP task with numerous downstream\napplications, such as summarization, machine translation and opinion mining. In\nthis paper, we demonstrate a simple, yet highly accurate discourse parser,\nincorporating recent contextual language models. Our parser establishes the new\nstate-of-the-art (SOTA) performance for predicting structure and nuclearity on\ntwo key RST datasets, RST-DT and Instr-DT. We further demonstrate that\npretraining our parser on the recently available large-scale "silver-standard"\ndiscourse treebank MEGA-DT provides even larger performance benefits,\nsuggesting a novel and promising research direction in the field of discourse\nanalysis.\n
Abstract This study uses treebanking to investigate how spoken language infiltrated legal Latin in early medieval Italy. The documents used are always formulaic, but they also always contain a ‘free’ part where the case in question is described in free prose. This paper uses this difference to measure how ten linguistic features, representative of the evolution that took place between Classical and Late Latin, are distributed between the formulaic and free parts. Some variants are attested equally often in both parts of the documents, while perceptually or conceptually salient variants appear to be preserved in their conservative form mainly in the formulaic parts.
Current methods of cross-lingual parser transfer focus on predicting the best\nparser for a low-resource target language globally, that is, "at treebank\nlevel". In this work, we propose and argue for a novel cross-lingual transfer\nparadigm: instance-level parser selection (ILPS), and present a\nproof-of-concept study focused on instance-level selection in the framework of\ndelexicalized parser transfer. We start from an empirical observation that\ndifferent source parsers are the best choice for different Universal POS\nsequences in the target language. We then propose to predict the best parser at\nthe instance level. To this end, we train a supervised regression model, based\non the Transformer architecture, to predict parser accuracies for individual\nPOS-sequences. We compare ILPS against two strong single-best parser selection\nbaselines (SBPS): (1) a model that compares POS n-gram distributions between\nthe source and target languages (KL) and (2) a model that selects the source\nbased on the similarity between manually created language vectors encoding\nsyntactic properties of languages (L2V). The results from our extensive\nevaluation, coupling 42 source parsers and 20 diverse low-resource test\nlanguages, show that ILPS outperforms KL and L2V on 13/20 and 14/20 test\nlanguages, respectively. Further, we show that by predicting the best parser\n"at the treebank level" (SBPS), using the aggregation of predictions from our\ninstance-level model, we outperform the same baselines on 17/20 and 16/20 test\nlanguages.\n
Context: Parkinson’s disease (PD) is a neurodegenerative disease caused by degeneration of the dopaminesynthesizing cells of the mesostriatal-mesocortical neuronal pathway,which affects motor pathway in basal ganglia (BG). Neuropsychological studies showed that degeneration of dopamine neuroreceptor also affects nigrostriatal and mesocortical limbic system which is associated with emotional processing in PD. However, very few studies have identified deficit in selective attention in patients with PD patients except in patients with PD-MCI (PD-Mild Cognitive Impairment) or PD-D (PD-Dementia). Thus, the present study examined the effect of emotion on attentional processing in PD and matched control. Emotional flanker task was designed by using pictures selected from the International Affective Picture System (IAPS) based on their normative valence ratings. Results revealed that attentional processing of emotional images were slower in PD patients in comparison to matched healthy control.
Der Beitrag reflektiert sprachliche Normen und den Umgang mit innersprachlicher Variation im Deutschunterricht. Das Inklusionsparadigma fordert hier von den Lehrkräften eine gesteigerte Sprachreflexionskompetenz, die Exklusionsrisiken erkennt. Am Beispiel von Sprachideologien in Bezug auf Bildungs- und Standardsprache im Deutschunterricht wird die Problemlage illustriert, und es werden Schlussfolgerungen für die Lehrer*innenbildung gezogen. The article reflects linguistic norms and linguistic variation in German teaching and learning at school. The call for inclusion demands from teachers a greater ability to reflect language use and by that, identify risks of exclusion. Based on the analysis of language ideologies with a focus on academic and standard language in school, the article draws conclusions regarding teacher training under the conditions of inclusion.
In this paper, we aim at improving the study of Latin in three ways: 1) by providing better visualizations of syntagma and structure for both research and the classroom, 2) by supporting a high-level search interface for corpus exploration, and 3) by improving the accuracy of taggers and parsers. To achieve this, we introduce a new linguistic description called Intelligenti Pauca, an alternative to Universal Dependencies for under-resourced languages. We show the key differences between the two linguistic descriptions, how the structure of Intelligenti Pauca favours our goals, and the effect it has on parsing accuracy for the Index Tomisticus Treebank.
We study the effect of rich supertag features in greedy transition-based dependency parsing. While previous studies have shown that sparse boolean features representing the 1-best supertag of a word can improve parsing accuracy, we show that we can get further improvements by adding a continuous vector representation of the entire supertag distribution for a word. In this way, we achieve the best results for greedy transition-based parsing with supertag features with $88.6\%$ LAS and $90.9\%$ UASon the English Penn Treebank converted to Stanford Dependencies.
The article attempts to conduct a primary analysis of the consequences of digital transformation for heritage languages which make up the cultural and historical legacy of individual ethnic communities. In a multilingual society, such a study requires an integrated approach, which takes into account the sociolinguistic parameters of various target audiences, communication channels aimed to disseminate and transfer information, discourse analyses of lin-guistic means, as well as extralinguistic factors impacting the development of different environments. It is equally important to study the specificity of the socio-cultural interaction between communicants in the professional sphere, which primarily indicates the institutional status of participants in communication, as well as their observance / nonobservance of linguistic norms. The latter seems extremely important with regard to heritage languages and their linguistic status in institutional discourse. In many respects, observance / nonobservance of linguistic norms makes it possible, on the one hand, to define the linguistic portrait of the communicant and, on the other hand, to assess the survival of national identity. Both aspects are central across various types of institutional discourse, including political, marketing, ad-vertising discourse etc. The analysis of the institutional aspects of cross-cultural and cross-lingual communication is carried out using an etiological approach that allows to determine the degree of importance of sociolinguistic parameters to achieve adequacy of socio-cultural interaction of representatives of different linguocultures. It is performed indirectly using vari-ous language pairs, in the context of heritage bilingualism, as well as interpersonal interaction. The article also expounds consequences of the global turn towards digital transformation affecting the overall knowledge in liberal arts and human sciences in general and cross-lingual and cross-cultural communication in particular. The study discusses areas of application of heritage language resources such as locus branding, image making, reports of scientific and technical achievements, etc. The article concludes by inferring the need to preserve linguistic diversity and its teleological use in various types of institutional discourse.
Current methods of cross-lingual parser transfer focus on predicting the best parser for a lowresource target language globally, that is, "at treebank level". In this work, we propose and argue for a novel cross-lingual transfer paradigm: instance-level parser selection (ILPS), and present a proof-of-concept study focused on instance-level selection in the framework of delexicalized parser transfer. Our work is motivated by an empirical observation that different source parsers are the best choice for different Universal POS-sequences (i.e., UPOS sentences) in the target language. We then propose to predict the best parser at the instance level. To this end, we train a supervised regression model, based on the Transformer architecture, to predict parser accuracies for individual POS-sequences. We compare ILPS against two strong single-best parser selection baselines (SBPS): (1) a model that compares POS n-gram distributions between the source and target languages (KL) and ( The results from our extensive evaluation, coupling 42 source parsers and 20 diverse low-resource test languages, show that ILPS outperforms KL and L2V on 13/20 and 14/20 test languages, respectively. Further, we show that by predicting the best parser "at treebank level" (SBPS), using the aggregation of predictions from our instance-level model, we outperform the same baselines on 17/20 and 16/20 test languages.
Some fragments of Epicharmus are examined in the context of contemporary scholarly discourse, especially with regard to textual and literary criticism, grammar and stylistics as they developed in the Sicily of his time. The interaction of the comedy of Epicharmus with scholarship is complex and involves a reflection and revision of contemporary ideas, but also a process of literary differentiation. This new genre was created through a critical interaction with other genres as part of a process of reinterpretation and literary exegesis. Furthermore, the fragments discussed offer the possibility to speculate on the handling of linguistic norms and also literary standards in pre- and early classical Sicily. They illustrate the interaction between different genres and the way they were integrated into text, metatext and performance to create tension and comic effects.
This chapter provides an account of the Epicurean theory of language, focusing in particular on Epicurus’ account of the origins of language, as detailed at <italic>Ep. Hdt</italic>. 75–6. It identifies two forms of linguistic naturalism (‘functional’ and ‘referential’) in Epicurus’ account of the first stage of linguistic phylogeny. It goes on to describe the implications of the advent of the second, conventionalist stage for Epicurus’ linguistic naturalism. The chapter suggests that Lucretius (like Epicurus before him) may be considered a latter-day συνειδών, enlarging and improving the language via the introduction and development of new expressions for new philosophical concepts. Finally, it considers how, if at all, Epicurean linguistic norms may have been grounded in Epicurean linguistic naturalism.
Cet article propose d’analyser les apports d’un modèle de langue pré-entraîné de type BERT (bidirectional encoder representations from transformers) à l’analyse syntaxique en constituants discontinus en anglais (PTB, Penn Treebank). Pour cela, nous réalisons une comparaison des erreurs d’un analyseur syntaxique dans deux configurations (i) avec un accès à BERT affiné lors de l’apprentissage (ii) sans accès à BERT (modèle n’utilisant que les données d’entraînement). Cette comparaison s’appuie sur la construction d’une suite de tests que nous rendons publique. Nous annotons les phrases de la section de validation du Penn Treebank avec des informations sur les phénomènes syntaxiques à l’origine des discontinuités. Ces annotations nous permettent de réaliser une évaluation fine des capacités syntaxiques de l’analyseur pour chaque phénomène cible. Nous montrons que malgré l’apport de BERT à la qualité des analyses (jusqu’à 95 en F1 ), certains phénomènes complexes ne sont toujours pas analysés de manière satisfaisante.
Cet article propose d’analyser les apports d’un modele de langue pre-entraine de type BERT (bidirectional encoder representations from transformers) a l’analyse syntaxique en constituants discontinus en anglais (PTB, Penn Treebank). Pour cela, nous realisons une comparaison des erreurs d’un analyseur syntaxique dans deux configurations (i) avec un acces a BERT affine lors de l’apprentissage (ii) sans acces a BERT (modele n’utilisant que les donnees d’entrainement). Cette comparaison s’appuie sur la construction d’une suite de tests que nous rendons publique. Nous annotons les phrases de la section de validation du Penn Treebank avec des informations sur les phenomenes syntaxiques a l’origine des discontinuites. Ces annotations nous permettent de realiser une evaluation fine des capacites syntaxiques de l’analyseur pour chaque phenomene cible. Nous montrons que malgre l’apport de BERT a la qualite des analyses (jusqu’a 95 en F1 ), certains phenomenes complexes ne sont toujours pas analyses de maniere satisfaisante.
This paper is part of the project Between Lexicon and Grammar (2016–2018), supported by the Grant Agency of the Czech Republic, reg. no. 16-07473S. This project is a follow-up of the project entitled The Grammar-Based Treebank of Czech (2013–2015,cf. Skoumalova et al. 2014; Petkevic et al. 2015a, 2015b) and devoted to automatic parsing driven by a formal HPSG-like grammar of Czech.
This paper explores the possibility of improving the performance of\nspecialized parsers for pre-modern Slavic by training them on data from\ndifferent related varieties. Because of their linguistic heterogeneity,\npre-modern Slavic varieties are treated as low-resource historical languages,\nwhereby cross-dialectal treebank data may be exploited to overcome data\nscarcity and attempt the training of a variety-agnostic parser. Previous\nexperiments on early Slavic dependency parsing are discussed, particularly with\nregard to their ability to tackle different orthographic, regional and\nstylistic features. A generic pre-modern Slavic parser and two specialized\nparsers -- one for East Slavic and one for South Slavic -- are trained using\njPTDP (Nguyen & Verspoor 2018), a neural network model for joint part-of-speech\n(POS) tagging and dependency parsing which had shown promising results on a\nnumber of Universal Dependency (UD) treebanks, including Old Church Slavonic\n(OCS). With these experiments, a new state of the art is obtained for both OCS\n(83.79\\% unlabelled attachment score (UAS) and 78.43\\% labelled attachement\nscore (LAS)) and Old East Slavic (OES) (85.7\\% UAS and 80.16\\% LAS).\n
We present a bracketing-based encoding that can be used to represent any 2-planar dependency tree over a sentence of length n as a sequence of n labels, hence providing almost total coverage of crossing arcs in sequence labeling parsing. First, we show that existing bracketing encodings for parsing as labeling can only handle a very mild extension of projective trees. Second, we overcome this limitation by taking into account the well-known property of 2-planarity, which is present in the vast majority of dependency syntactic structures in treebanks, i.e., the arcs of a dependency tree can be split into two planes such that arcs in a given plane do not cross. We take advantage of this property to design a method that balances the brackets and that encodes the arcs belonging to each of those planes, allowing for almost unrestricted non-projectivity ( 99.9% coverage) in sequence labeling parsing. The experiments show that our linearizations improve over the accuracy of the original bracketing encoding in highly non-projective treebanks (on average by 0.4 LAS), while achieving a similar speed. Also, they are especially suitable when PoS tags are not used as input parameters to the models.
This paper explores the possibility of improving the performance of specialized parsers for pre-modern Slavic by training them on data from different related varieties. Because of their linguistic heterogeneity, pre-modern Slavic varieties are treated as low-resource historical languages, whereby cross-dialectal treebank data may be exploited to overcome data scarcity and attempt the training of a variety-agnostic parser. Previous experiments on early Slavic dependency parsing are discussed, particularly with regard to their ability to tackle different orthographic, regional and stylistic features. A generic pre-modern Slavic parser and two specialized parsers -- one for East Slavic and one for South Slavic -- are trained using jPTDP (Nguyen & Verspoor 2018), a neural network model for joint part-of-speech (POS) tagging and dependency parsing which had shown promising results on a number of Universal Dependency (UD) treebanks, including Old Church Slavonic (OCS). With these experiments, a new state of the art is obtained for both OCS (83.79\% unlabelled attachment score (UAS) and 78.43\% labelled attachement score (LAS)) and Old East Slavic (OES) (85.7\% UAS and 80.16\% LAS).
In this paper, we introduce experiment results of a Vietnamese sentence parser which is built by using the Chomsky’s subcategorization theory and PDCG (Probabilistic Definite Clause Grammar). The efficiency of this subcategorized PDCG parser has been proved by experiments, in which, we have built by hand a Treebank with 1000 syntactic structures of Vietnamese training sentences, and used different testing datasets to evaluate the results. As a result, the precisions, recalls and F-measures of these experiments are over 98%.
Neste artigo, buscamos descrever as concepcoes de lingua e norma linguistica que sao veiculadas em gramaticas produzidas pela Real Academia de Espanola (RAE) – instituicao fundada na Espanha, no inicio do seculo XVIII, cuja missao principal e a “defesa da unidade da lingua”. Foram feitas analises de excertos de dois manuais da instituicao, a saber: Esbozo de una nueva gramatica de la lengua espanola (1982) e Nueva gramatica de la lengua espanola - Manual (2010). Na analise, buscamos identificar de qual concepcao de Norma Linguistica o discurso normativo da RAE mais se aproxima, observando tanto os capitulos introdutorios quanto um capitulo com descricoes propriamente ditas. Apresentaremos, aqui, os resultados obtidos ao longo do projeto, que mostraram uma aproximacao maior a visao mais prescritiva do termo norma, apesar de todo o projeto de pan-hispanismo que teve como evento importante a publicacao dessas duas obras. ABSTRACT: In this paper, we describe how concepts such as ‘language’ and ‘linguistic norm’ are presented on the grammars produced by Real Academia Espanola (RAE) – an institution founded in Spain in the 18th century who claims to have the mission of “defending the unity of the language”. We made an analysis of two books published by RAE: Esbozo de una nueva gramatica de la lengua espanola (1982) and Nueva gramatica de la lengua espanola - Manual (2010). In the analysis, we seek to identify from which conception of Linguistic Norm the normative discourse of RAE comes closest and we have done that by studying the introductory chapters and a chapter with grammatical content from both publications. Here, we present the results of this research, which showed that the content of both of RAE’s books has more closeness to a more prescriptivist vision of said concepts, even though they are a direct result of the whole pan-hispanic politics promoted by the academy. KEYWORDS: Linguistic norm; Grammaticography; Real Academia Espanola; Spanish.
This article offers a review of the most relevant general dictionaries of Spanish. In doing so, we consider briefly the relationship between linguistic norms, standardization and dictionaries, highlighting the concept of "general dictionary" as most relevant to the issue of normativity. Thereafter, we take a look at the historical origins and emergence of Spanish general lexicography. Finally, we pay attention to more recent developments in the field as there are the lexicographic implications and consequences of pluricentricity which implies a new mode of codification as well as new developments due to digitalization.
Graphs have become an increasingly important means of representing data, for instance, when communicating data on climate change. However, graph characteristics might significantly affect graph comprehension. The goal of the present work was to test whether the marking forms usually depicted on line-graphs, can have an impact on graph evaluation. As past work suggests that triangular forms might be related to threat, we compared the effect of triangular marking forms with other symbols (triangles, circles, squares, rhombi, and asterisks) on subjective assessments. Participants in Study 1 ( N = 314) received 5 different line-graphs about climate change, each of them using one out of 5 marking forms. In Study 1, the threat and arousal ratings of the graphs with triangular marking shapes were not higher than those with the other marking symbols. Participants in Study 2 ( N = 279) received the same graphs, yet without labels and indeed rated the graphs with triangle point markers as more threatening. Testing whether local rather than global spatial attention would lead to an impact of marker shape in climate graphs, Study 3 ( N = 307) documented that a task demanding to process a specific data-point on the graph (rather than just the line graph as a whole) did not lead to an effect either. These results suggest that marking symbols can principally affect threat and arousal ratings but not in the context of climate change. Hence, in graphs on climate change, choice of point markers does not have to take potential side-effects on threat and arousal into account. These seem to be restricted to the processing of graphs where form aspects face less competition from the content domain on judgments.
Because of its focus on the past and on historical languages, the classics is a discipline that is particularly interested in translations and text alignment. Starting from a diachronic perspective, this contribution demonstrates how issues related to text alignment, present since antiquity, can be approached from a different angle and with entirely new opportunities thank to tools and methods developed in the field of digital humanities. By comparing examples from antiquity (e.g. Origen’s Hexapla from the third century CE) with modern projects based on treebanking and dependency grammar (e.g. the Ancient Greek and Latin Dependency Treebank [AGLDT] as part of the Perseus Digital Library from Tufts University), we shall present some new approaches and their potentials. In doing so, we shall also examine what status English has in these projects and how the different languages involved in each of them interact with English and/or with each other.
We propose an approach for generating an accurate and consistent PropBank-annotated corpus, given a FrameNet-annotated corpus which has an underlying dependency annotation layer, namely, a parallel Universal Dependencies (UD) treebank. The PropBank annotation layer of such a multi-layer corpus can be semi-automatically derived from the existing FrameNet and UD annotation layers, by providing a mapping configuration from lexical units in [a non-English language] FrameNet to [English language] PropBank predicates, and a mapping configuration from FrameNet frame elements to PropBank semantic arguments for the given pair of a FrameNet frame and a PropBank predicate. The latter mapping generally depends on the underlying UD syntactic relations. To demonstrate our approach, we use Latvian FrameNet, annotated on top of Latvian UD Treebank, for generating Latvian PropBank in compliance with the Universal Propositions approach.
Resumo: Neste artigo, abordamos os conceitos de erro linguístico surgidos com a relativização do conceito de norma linguística, que, após os estudos da Linguística Moderna, não mais diz respeito unicamente às regras prescritas pela gramática normativa. Defendemos a tese de que o erro linguístico existe e está diretamente relacionado às normas linguísticas exigidas para o contexto de emprego da língua. Este artigo tem os seguintes objetivos: a) demonstrar que linguistas de perspectivas teóricas distintas reconhecem a existência do erro linguístico; b) demonstrar que o erro linguístico não exclui o reconhecimento da variação linguística, mas o endossa; c) apresentar os diferentes entendimentos sobre a questão do erro linguístico; e d) fomentar uma rediscussão sobre o conceito de erro linguístico. Metodologicamente, este é um trabalho bibliográfico de caráter qualitativo. A pesquisa bibliográfica implica a análise ou a resolução de um problema, recorrendo a referenciais teóricos enquanto fontes importantes para a pesquisa. A pesquisa qualitativa preocupa-se com a dimensão descritiva do fenômeno, ocupando-se primordialmente com o(s) processo(s), sem ignorar os resultados e os produtos. Concluímos que as discussões que se dão no campo da Linguística Aplicada sobre o erro no emprego da língua ocorrem basicamente no âmbito da terminologia, e não no âmbito do que constitui o erro em língua. Palavras-chave: norma linguística; erro linguístico; correção linguística. Abstract: This article aims to discuss the concepts of linguistic error which have arisen from the relativization of linguistic norm definition, since under the light of Modern Linguistic studies, linguistic norm does not only refer to prescriptive grammar’s rules anymore. We defend that linguistic error does exist and it is directly related to linguistic norms required for the language usage context. This article has the following objectives: a) demonstrate that linguists from different perspectives recognize the existence of linguistic error; b) expose that linguistic error does not exclude linguistic variation recognition; c) present different understandings about linguistic error and d) foment a rediscussion concerning the concept of linguistic error. This is a qualitative research based on bibliographic data. Bibliographic research implies the analysis or the resolution of a problem, taking theoretical references into consideration, whereas the qualitative research considers the descriptive dimension of the phenomenon and is concerned primarily with the process itself, without taking into account the results and products. We conclude that discussions about error in language usage, which take place in the Applied Linguistics field of studies, basically occur in the terminology aspect, not in what error in language use is constituted of. Keywords: linguistic norm; linguistic error; linguistic correction.
Different attitudes towards the definition of linguistic variations are noticeable in contemporary teaching of French as a foreign language. The difference depends on whether it is viewed from the perspective of traditionally understood exemplary linguistic levels le bon usage, from the standpoint of propositions – the “basic French language” the SGAV methodology is based on; or with the emergence of a communicative approach in foreign language teaching from the perspective of a moderately conservative and flexible attitude towards conservative linguistic norms. In order to study how linguistic variations in the teaching of French as a foreign language are treated, our research was based on an analysis of the use of different forms for expressing a direct partial question with interrogative adverbs où, quand, comment and pourquoi in seven selected methods published from the middle of the last century until the present day. The results of the qualitative and quantitative analyses of the abovementioned interrogative forms in the selected methods, illustrate significant oscillations in terms of the representation of interrogative forms derived from inversion and those which, contrary to the conservative linguistic norm and in accordance with the general tendencies in the French spoken language, preserve the word order characteristic of a statement. The obtained results indicate an expressed need for a more consistent pragmatic definition of linguistic variation in the contemporary teaching of French as a foreign language.
BACKGROUND: Instrumental activities of daily living (IADL) impairment can begin in mild cognitive impairment (MCI), and is the core criteria for diagnosing dementia in both Alzheimer's (AD) and Parkinson's (PD) diseases. The Functional Activities Questionnaire (FAQ) has high discriminative power for dementia and MCI in older age populations, but is influenced by demographic factors. It is currently unclear whether the FAQ is suitable for assessing cognitive-associated IADL in non-demented PD patients, as motor disorders may affect ratings. OBJECTIVE: To compare IADL profiles in MCI patients with PD (PD-MCI) and AD (AD-MCI) and to verify the discriminative ability of the FAQ for MCI in patients with (PD-MCI) and without (AD-MCI) additional motor impairment. METHODS: Data of 42 patients each of PD-MCI, AD-MCI, PD cognitively normal (PD-CN), and healthy controls (HC), matched according to age, gender, education, and global cognitive impairment were analyzed. ANCOVA and binary regressions were used to examine the relationship between the FAQ scores and groups. FAQ cut-offs for PD-MCI (versus PD-NC) and AD-MCI (versus HC) were separately identified using receiver operating characteristic analyses. RESULTS: FAQ total score did not differentiate between MCI groups. PD-MCI subjects had greater difficulties with tax records and traveling while AD-MCI individuals were more impaired in managing finances and remembering appointments. Classification accuracy of the FAQ was good for diagnosing AD-MCI (69%, cut-off ≥1) compared to HC, and sufficient for differentiating PD-MCI (38.1%, cut-off ≥3) from PD-CN. CONCLUSION: The FAQ task profiles and classification accuracy differed between MCI related to PD and AD.
We introduce a novel chart-based algorithm for span-based parsing of\ndiscontinuous constituency trees of block degree two, including ill-nested\nstructures. In particular, we show that we can build variants of our parser\nwith smaller search spaces and time complexities ranging from $\\mathcal O(n^6)$\ndown to $\\mathcal O(n^3)$. The cubic time variant covers 98\\% of constituents\nobserved in linguistic treebanks while having the same complexity as continuous\nconstituency parsers. We evaluate our approach on German and English treebanks\n(Negra, Tiger and Discontinuous PTB) and report state-of-the-art results in the\nfully supervised setting. We also experiment with pre-trained word embeddings\nand \\bert{}-based neural networks.\n
Abstract Individuals, who score high in self-reported intolerance of uncertainty (IU), tend to find uncertainty anxiety-provoking. IU has been reliably associated with disrupted threat extinction. However, it remains unclear whether IU would be related to disrupted extinction to other arousing stimuli that are not threatening (i.e., rewarding). We addressed this question by conducting a reward associative learning task with acquisition and extinction training phases ( n = 58). Throughout the associative learning task, we recorded valence ratings (i.e. liking), skin conductance response (SCR) (i.e. sweating), and corrugator supercilii activity (i.e. brow muscle indicative or negative and positive affect) to learned reward and neutral cues. During acquisition training with partial reward reinforcement, higher IU was associated with greater corrugator supercilii activity to neutral compared to reward cues. IU was not related to valence ratings or SCR’s during the acquisition or extinction training phases. These preliminary results suggest that IU-related deficits during extinction may be limited to situations with threat. The findings further our conceptual understanding of IU’s role in the associative learning and extinction of reward, and in relation to the processing of threat and reward more generally.
espanolEn este trabajo, argumentamos que en los actuales treebanks que aplican el formalismo de las Dependencias Universales, la anotacion de los reflexivos espanoles es un problema sin resolver, que afecta claramente a la precision y consistencia de los parsers actuales. Evaluamos diferentes propuestas para afinar las diferentes categorias y discutimos los problemas pendientes. Creemos que la solucion para estos problemas se puede encontrar en una anotacion en multiples niveles, combinando la anotacion de la relacion de dependencia y de las caracteristicas (features) de los tokens, en lugar de ampliar el numero de categorias en un solo nivel de anotacion. Aplicamos la propuesta a la version espanola del treebank UD AnCora (v2.5) y proporcionamos una tabla de conversion categorizada que se puede ejecutar mediante un script Python. EnglishIn this paper, we argue that in current Universal Dependencies treebanks, the annotation of Spanish reflexives is an unsolved problem, which clearly affects the accuracy and consistency of current parsers. We evaluate different proposals for fine-tuning the various categories, and discuss remaining open issues. We believe that the solution for these issues could lie in a multi-layered way of annotating the characteristics, combining annotation of the dependency relation and of the so-called token features, rather than in expanding the number of categories on one layer. We apply this proposal to the v2.5 Spanish UD AnCora treebank and provide a categorized conversion table that can be run with a Python script.
Abstract This paper illustrates how enriched diachronic treebank data can shed new light on an old and vexed topic, even when that topic is primarily morphological and semantic in nature rather than syntactic. The topic is the rise of the Russian po delimitatives, a change seen as crucial in most accounts of the history of Russian aspect, since it represents a major step in generalising the derivational aspect system. Earlier accounts concur that the po delimitatives spread fairly recently, too recently for the development to be connected to the loss of the aorist tense, which also had delimitative readings with atelic verbs. Using treebank data from the Tromsø Old Russian and Old Church Slavonic Treebank, enriched with tags for derivational morphology and semantics, I show that the po delimitatives were not marginal even in the earliest Slavic sources, either in terms of frequency or semantics, and that they first complemented and then competed with the delimitative aorists. It can thus be claimed that the exotic po delimitatives grew organically out of the old Indo-European inflectional aspect system.
Universal Dependencies conversion of the Late Latin Charter Treebank 2
Investigating the neural correlates of arousal in non-offending paedophiles provides an opportunity to focus on the dysfunctional process underlying paedophilic predilections. The study aimed to investigate the neural correlates of sexual arousal related to paedophilic preference. Functional magnetic resonance imaging (fMRI) and phallometric testing were used to determine sexual arousal related brain activity while listening to a series of auditory narratives including neutral, adult sexual, and child sexual content. Participants included 12 non-offending paedophiles and a control sample of 12 males with teleiophilic preferences. Blood oxygen level-dependent fMRI, phallometric response and subjective arousal ratings related to auditory stimuli. Subjective arousal ratings showed expected direction of preference for adult and child sexual content between control group and non-offending paedophiles respectively. Phallometric response during adult sexual content was higher in the control group although there was no difference between samples during the child sexual content. Imaging contrasts of the child relative to adult sexual content revealed functional activity in the cerebellum specific to non-offending paedophiles.
In order to explore the possibility of leveraging discourse information for the identiffication of argumentative components and relations we add a new annotation layer to a subset of the Discourse Dependency TreeBank for Scientiffic Abstracts (SciDTB). [1] We introduce a ffine-grained annotation schema aimed at capturing information that accounts for the specifficities of the scientiffic discourse, including the type of evidence that is offered to support a statement (e.g., background information, experimental data or interpretation of results). [1] Yang, A., Li, S.: SciDTB: Discourse dependency TreeBank for scientiffc abstracts. In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (ACL 2018) (Volume 2: Short Papers). pp. 444{449. Association for Computational Linguistics, Melbourne, Australia (Jul 2018)
Preprint for the paper that was submitted to the workshop 'Treebanks and Linguistic Theories 2020' on October 27-28, 2020.
A wide variety of transition-based algorithms are currently used for dependency parsers. Empirical studies have shown that performance varies across different treebanks in such a way that one algorithm outperforms another on one treebank and the reverse is true for a different treebank. There is often no discernible reason for what causes one algorithm to be more suitable for a certain treebank and less so for another. In this paper we shed some light on this by introducing the concept of an algorithm's inherent dependency displacement distribution. This characterises the bias of the algorithm in terms of dependency displacement, which quantify both distance and direction of syntactic relations. We show that the similarity of an algorithm's inherent distribution to a treebank's displacement distribution is clearly correlated to the algorithm's parsing performance on that treebank, specifically with highly significant and substantial correlations for the predominant sentence lengths in Universal Dependency treebanks. We also obtain results which show a more discrete analysis of dependency displacement does not result in any meaningful correlations.
Using incorrect worked examples during mathematics instruction can improve student learning. However, teachers worry that students may confuse correct and incorrect examples over time, and memory research supports this fear. To examine if this forgetting occurs, we had undergraduates rate the correctness of correct and incorrect worked examples immediately and one week later (Experiment 1). Previously studied incorrect examples were rated as slightly more correct after the delay, but this did not affect ratings of unstudied examples or problem-solving accuracy. In Experiment 2, we more closely mimicked how incorrect worked examples are used in classroom settings. Again, we found only small changes in students’ memory for studied worked examples after the delay, and no changes for unstudied examples or problem-solving accuracy. Our findings suggest the costs of teaching with incorrect worked examples are limited to the specific studied problems, and do not affect learning of the underlying mathematical rule.
We experience our neighborhoods and cities through our streets. An essential question is what street features contribute to a superior pedestrian experience? Further, can interacting with urban streetscapes improve cognitive functioning, as interacting with nature does? To answer these questions, large urban street perception databases are enormously helpful. Toward this end, we propose a crowdsourcing method for measuring the pedestrian qualities of streets – preference, imageability, complexity, enclosure, human scale, transparency, and order, all of which are important design dimensions for pedestrian experience. We obtained 556 street images from Google Street View by sampling two sidewalk images from 278 geo-coordinates in Chicago. For each dimension, over 58 (SD=2.5) Amazon Mechanical Turk workers (469 in total) completed an image rating task for the 556 images. In each trial, participants were shown 12 images in a 4x3 grid and asked to choose four images that they evaluate highly on that dimension. The probability of selecting each image for a given dimension across participants was used to quantify how much that image represented that dimension. To test the inter-rater reliability, we randomly split participants into two groups 2000 times and calculated rank-correlations between the measures from each group. We found that the split-half correlations were high for walkability (0.86±0.01; M±SD), preference (0.83±0.02), imageability (0.78±0.02), complexity (0.79±0.02), enclosure (0.80±0.02), and transparency (0.88±0.01) and modest for human scale (0.45±0.05) and order (0.57±0.05). To test whether the measures are internally consistent, we randomly split two images from the same geo-coordinate into two bins (of 278 images) 2000 times, calculated rank-correlations, and found modest to low correlation values, range: [0.06, 0.37]. These results suggest that our method can be used to efficiently measure pedestrian street qualities from non-experts, but with some limitations, to build a large database for studying street features affecting urban street pedestrian experiences.
Language users and learners are sensitive to distributional information in their environment, which enables them to extract regularities that occur in the language input that they are exposed to. This process is referred to as statistical learning. While the statistical learning phonotactic literature thoroughly investigates the learning of overall phonotactics in specific languages, little is known about cases where different phonological systems coexist within a single language. The Japanese lexicon is generally classified into four lexical strata according to the etymological status of each word (Itô & Mester, 1995, 1999, 2001). Although each stratum includes the internal phonological similarity in the Japanese language as a whole, there are also distinctive phonological properties. A recent study suggests that language users should be able to learn phonotactics of each sublexicon based on the same kind of statistical probabilities that computers analyse from language users’ accumulated lexicons (Morita, 2018). This thesis examines whether second-language (L2) learners can learn the loanword phonotactics/phonology of Japanese through experience of using and/or passive exposure to Japanese lexical stratification. Using two loanword phonological regularities (categorical and gradient rules) as a case study, two fully-crossed perceptual experiments involving English- speaking learners of Japanese, native speakers of Japanese, and English-speaking monolinguals are presented. The first experiment explores listeners’ phonotactic/phonological knowledge of nativised loanwords in Japanese using a well-formedness task which shows the adaptation of English final consonants in monosyllabic words. Listeners judge whether the pronunciation they hear is how the word would be pronounced if it was a Japanese word, rating how confident they are on a scale of 1-5. This study shows that L2 learners learn categorical rules, but not gradient patterns. This study also confirms that loanword phonotactics and overall phonotactics make separate contributions to perceived well-formedness. L2 learners access and make use of the sublexicon-specific probabilities of Japanese during the task. The second perceptual experiment is designed to support the findings in the first experiment, by testing for discrimination of non-native consonantal contrasts. Even under high memory demand, L2 learners show the ability to discriminate non-native consonantal contrasts (i.e., CVCV/CVCCV) effectively enough to support findings in the first experiment. These results suggest that L2 learners can implicitly detect the statistical structure of a language’s sublexicon phonology over the course of acquiring a natural language. However, while native speakers of Japanese learn a gradient rule, L2 learners of Japanese do not. A potential explanation for the differences in gradient rule learning is that the vocabulary size of the target language might play a crucial role. This remains an open question. In addition, the present work provides a basis for future investigation into whether L2 learners of Japanese, whose native language is other than English, are able to learn Japanese loanword phonotactics/phonology. L1 English-L2 Japanese speakers might gain advantage in perceiving the English input which inevitably overlaps with the phonological form of the host language.
This research will focus on the use of Javanese communication related to the form of family greetings that can be seen from the completeness of its elements. Javanese communication forms of family greetings are divided into three, namely: complete greeting forms, incomplete greeting forms, and a combination of complete greeting forms and incomplete greeting forms. Whereas based on the meanings and meanings of language communication, the form of family greetings can be in the form of self-names, kinship terms, paraban, national titles, adjective transpositions, and beatings. Factors that influence Javanese communication in the form of family greetings are the position of parents towards their children viewed from various aspects of course higher, but related to the use of the form of greeting it turns out that its use often shows a respectful form of greeting. This can be related to the role of the first person as a parent whose obligation is to educate and direct their children to be good children, who have good manners and can respect others and also their own parents. Other things that affect the form of family greetings are the first person, second person, third person, the meaning of the speaker, the color of the emotion, the tone of the speech, the subject, speech sequence, form of discourse, speech facilities, speech scenes, speech environment, and linguistic norms.
 
 Abstrak
 Penelitian ini akan difokuskan pada penggunaan komunikasi bahasa Jawa yang berkaitan dengan bentuk sapaan keluarga yangdapat dilihat dari kelengkapan unsur-unsurnya. Komunikasi bahasa Jawa bentuk sapaan keluarga dibedakan menjadi tiga, yaitu: bentuk sapaan lengkap, bentuk sapaan tak lengkap, dan gabungan bentuk sapaan lengkap dan bentuk sapaan tak lengkap. Sedangkan berdasarkan makna dan artinya komunikasi Bahasa bentuk sapaan keluarga dapat berupa nama diri, istilah kekerabatan, paraban, gelar kebangsawaan, transposisi ajektif dan poyokan. Faktoryang mempengaruhi komunikasi Bahasa Jawa dalam bentuk sapaan keluargaadalah posisi orang tua terhadap anak-anaknya dilihat dari berbagai segi tentunya lebih tinggi, namun berkaitan dengan pemakaian bentuk sapaan ternyata sering sekali penggunaannya justru menunjukan bentuk sapaan yang hormat. Hal ini dapat dikaitkan dengan peran orang pertama sebagai orang tua yang salah satu kewajibannya adalah mendidik dan mengarahkan anak-anaknya agar menjadi anak yang baik, yang memiliki sopan santun dan dapat menghormati orang lain dan juga orang tuanya sendiri. Hal lain yang mempengaruhi bentuk sapaan keluarga adalah orang pertama, orang kedua, orang ketiga, maksud penutur, warna emosi, nada suasana bicara, pokok pembicaraan, urutan bicara, bentuk wacana, sarana tutur, adegan tutur, lingkungan tutur, dan norma kebahasaan.
 Kata kunci: bentuk sapaan, komunikasi Bahasa Jawa
Introduction. High-quality language education in technical universities requires its interdisciplinary relation to the content of highly specialised subjects corresponding to the training programmes aimed at instructing the future specialists. Educational materials in a foreign language are highly productive if they emphasise the terminology and professional vocabulary authentic to the current state of the scientific field. The aim of the study presented in the article was to assess the validity of the lexical material delivered in the course “English for Business Communication”, to determine the selection criteria for this vocabulary as well as the methods for its assimilation and practical application. Methodology and research methods. The applied corpus software enabled to obtain quantitative indicators of the distribution of foreign-language business vocabulary in the given training course. The lexical material being currently offered to students and the professional thesaurus identified via linguistic databases was compared with the use of comparative analysis and synthesis. Results and scientific novelty. The lexical units (terms, set expressions), which are the most active in the business sphere, were identified on the basis of its frequency. The authors established the correlation between them and educational vocabulary, both from the perspective of its integration into the course without block concentration throughout the course of university training, and from the perspective of the variety of methods used to practice this vocabulary. It is concluded that the applied educational material needs to be substantially adjusted. The vocabulary does not completely reflect the realities of the business communication sphere and the distribution of active vocational vocabulary regulated by methodological guidelines does not entirely contribute to its strong assimilation. According to the authors, the necessary changes to the approaches and methods for selecting and compiling lexical material and to the methodology for designing a foreign language course should be made on the basis of integrating pedagogical and linguistic knowledge, in particular, the methodology of teaching foreign languages and the corpus linguistics. Practical significance. The ways of integrating corpus programs in the process of developing the content of language disciplines, which are part of the main educational program of technical universities, are demonstrated as one of the methods to increase the effectiveness of teaching foreign languages to students of non-linguistic specialties.
The aim of the study is to design an electronic terminographic product in the form of a terminological data bank (TDB) “Classification Parameters of Phraseological Units” with consistent implementation of infological and datological stages. The object in the study is the TDB, the focus is terms for designation of types of phraseological units (PU), which actively function in Ukrainian and Russian linguistics in the 2nd half of the 20th century to 2020. The methodology of the study is based on theoretical foundations of phraseology, terminology, applied linguistics, and computer terminography (linguistic database/data bank technology). The general and specific scientific methods are used: analysis and synthesis, ascent from the abstract to the concrete, classification, definitional analysis, identification, linguistic observation and description, selection, elements of a statistical method. The infological stage is related to the solution of a set of information problems: 1) definition of the principle of regularization of terms of phraseology, 2) selection of terms with the archeseme “phraseoclassification type” from scientific literature in order to form an alphabetical register, 3) establishment of the composition of the dedicated terminology subsystems, 4) description of terminology subsystems, 5) creation of an electronic catalog of terms for further computer processing of collected information. The main principle of systematization is the thesaurus: “terminological system – terminological microsystem – terminological subsystem – term”. The next step is to develop a datalogical scheme, which is a system of tables whose fields display information about the described terms in the form of data, for example: number, term, terminology subsystem, scientific source, definition, illustration, paradigm relations of the term. The TDB modeled in the work is a special terminology dictionary (monolingual (Ukrainian) in terms of the number of languages involved), which has 113 terms to denote types of PUs and covers 17 classifications of PUs. The TDB “Classification Parameters of Phraseological Units” is positioned as an electronic terminographic product created for storage of information, optimization and intensification of system research on fundamental issues of phraseology and phraseography, works devoted to problems of comprehensive analysis and parameterization of PUs. The prospect of the study is to create a Russian-language version of this terminographic product.
The aim of this paper is to retrieve the most relevant expansion words for expanding the initial query of the user in order to enhance the outcomes of web search results. Query expansion plays a major role in reformulating a user’s initial query to a one more pertinent to the user’s intended meaning. The reformulated query is then used to obtain more appropriate outcomes from a large amount of information on the web. The proposed semantic query expansion technique uses Wikipedia and WordNet as data sources. Wikipedia is taken as a base for all query expansions because it is one of the most diversified and relevant databases available on the web. To further improve the proposed query expansion technique,WordNet—a lexical database—is used as the as another data source because the synonyms (synsets) of the query term provided by it can be quite useful for query expansion. The proposed expansion technique successfully combines the two data sources to retrieve the most relevant expansion terms from the data sources in response to the user’s original query. The proposed work has been divided into four phases: (1) extraction of relevant words from Wikipedia (2) extraction of relevant words from WordNet (3) merging of the expansion terms obtained from Wikipedia and WordNet, and (4) query formulation by combining the expansion terms using Boolean operators. This reformulated query is then fired on the web to find the desired result. The Experimental result shows a significant improvement in information retrieval using query expansion.
Existing approaches for graph-based deep dependency parsing that force connexity in predicted graphs do no cover the structures observed in French treebanks. We propose a novel algorithm that covers the full set of possible structures. We evaluate our approach on the French corpora FTB and Sequoia and observe a trade-off between the validity of predicted structures and the quality of predictions