Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Social interactions enhance human memories, but little is known about how the neural mechanisms underlying episodic memories are modulated by rewarding outcomes in social interactions. To investigate this, fMRI data were recorded while healthy young adults encoded unfamiliar faces in either a competition or a control task. In the competition task, participants encoded opponents' faces in the rock-paper-scissors game, where trial-by-trial outcomes of Win, Draw, and Lose for participants were shown by facial expressions of opponents (Angry, Neutral, and Happy). In the control task, participants encoded faces by assessing facial expressions. After encoding, participants recognized faces previously learned. Behavioral data showed that emotional valence for opponents' Angry faces as the Win outcome was rated positively in the competition task, whereas the rating for Angry faces was rated negatively in the control task, and that Angry faces were remembered more accurately than Neutral or Happy faces in both tasks. fMRI data showed that activation in the medial orbitofrontal cortex (mOFC) paralleled the pattern of valence ratings, with greater activation for the Win than Draw or Lose conditions of the competition task, and the Angry condition of the control task. Moreover, functional connectivity between the mOFC and hippocampus was increased in Win compared to Angry, and the mOFC-hippocampus functional connectivity predicted individual differences in subsequent memory performance only in Win of the competition task, but not in any other conditions of the two tasks. These results demonstrate that the memory enhancement by context-dependent social rewards involves interactions between reward- and memory-related regions.
Focus of the CONcreTEXT task is conceptual concreteness: systems were solicited to compute a value expressing to what extent target concepts are concrete (i.e., more or less perceptually salient) within a given context of occurrence. To these ends, we have developed a new dataset which was annotated with concreteness ratings. Interestingly, these works extend information on conceptual concreteness available in existing (non contextual) norms derived from human judgments with new knowledge from recently developed neural architectures, in much the same multidisciplinary spirit whereby the CONcreTEXT task was organized.
This paper deals with the topic of lexical modality in Norwegian as a second language. Basing on data obtained from the ASK-corpus -The Norwegian Language Learner Corpus containing second language texts written in a language examination, the authors analyse the use of two modal particles, jo and nok, by three groups of Norwegian as a second language writers: English, Polish and German. The focus of the study is on analysing lexical patterns for co-occurrence of the modal particles. The patterns used by the learners are compared with the ones used by the native speakers of Norwegian and between the learner groups. The discrepancies found in the data are discussed within the broader framework of second language development and second language writing. The findings suggest that the second language writers' use of modal particles is influenced by several factors, such as general interlanguage tendencies, transfer from the learners' first languages and the perception of textual norms.
This paper explores the knowledge of linguistic structure learned by large artificial neural networks, trained via self-supervision, whereby the model simply tries to predict a masked word in a given context. Human language communication is via sequences of words, but language understanding requires constructing rich hierarchical structures that are never observed explicitly. The mechanisms for this have been a prime mystery of human language acquisition, while engineering work has mainly proceeded by supervised learning on treebanks of sentences hand labeled for this latent structure. However, we demonstrate that modern deep contextual language models learn major aspects of this structure, without any explicit supervision. We develop methods for identifying linguistic hierarchical structure emergent in artificial neural networks and demonstrate that components in these models focus on syntactic grammatical relationships and anaphoric coreference. Indeed, we show that a linear transformation of learned embeddings in these models captures parse tree distances to a surprising degree, allowing approximate reconstruction of the sentence tree structures normally assumed by linguists. These results help explain why these models have brought such large improvements across many language-understanding tasks.
The article provides an overview of the lexical and grammatical features as well as the sociopolitical environment of Marollien that originated in the 18th century as a dialect on the territory of Brussels. Marollien is essentially the Dutch language in its Brabantian dialect, strongly influenced by French. There are literary works, performances, and musicals written and staged in Marollien, as well as dictionaries and journals published in it. Historically, the Marollien dialect is a sociolect: it was generally used by Belgians coming to Brussels from Wallonia in search of a job and settling in one of the districts of Brussels — Marolles. A special emphasis is placed on lexical features of the dialect: gastronomic and everyday vocabulary are looked at and the examples of French loanwords and Southern Dutch language norm deviations are provided. Standard Dutch calques in French, when translating idioms in particular, are also identified. The differences between Dutch, French, and Marollien place names are illustrated. In the field of morphology and word formation, there is a regular mixture of Germanic and Romanic stems which is indicated. Examples of Marollien phonetic features are also provided. The article acknowledges frequent code switching in Marollien speech, which by and large resembles the phenomenon of linguistic interference. Due to the fact that Marollien is rapidly disappearing, the Brussels-Capital region is trying to support the dialect: various activities are being organized in order to propagate its use and enhance its prestige. Nevertheless, Marollien is not included in the well-known citizen initiative “Marnix Plan”, aimed at developing the methodology for the sequential study of several languages for all segments of the population in Brussels. This initiative is also discussed in the article.
The article is based on an understanding of the possible functional and semantic classification of lexical units of the Portuguese language according to a graded criterion, which corresponds to the operational perspective of the semantic application of the language. One of the forms of visualization of grading operations can be a grading scale, or a graduation scale, the possibility of which is supported by the existence of an intuitive perception of a certain sample, a certain point of reference, a certain norm, above and below which are certain zones of units that fall into the grading situation. The author notes that the grading operator as a minimal linguistic variable is not only a marker that specifies the degree of deviation from a certain ordinary level and provides a modification of the value (movement down or up the axiological scale), but also an element of ordering reasoning, expression of opinion, and personal attitude of the Portuguese speaker. The article analyzes operators that belong to the group of high-degree and ultimate-measure graduators. The analysis of the combinability of the operators considered by the author allowed us to distinguish two ways of grading limit features in the Portuguese language: ingerent and extensive. Extensive gain has more to do with the verb, in the amplification of which the orientation of the actants are expressed more explicitly. This allows you to select a special type of gain – actant gain. However, even when grading adjectives, some Portuguese ultimate-measure gradators or graduators are able to participate in extensive models, such as the quantifier pronoun todo, toda (all, entire, whole). In addition to differences in the method of modifying a trait (extensive or inherent) and in the modal part of the value, ultimate measure operators differ in the nature of the trait representation. Some of them represent a trait in statics, regardless of its previous development (absolutamente, inteiramente, totalmente), and others represent the ultimate measure of the trait as the result of its previous development and accumulation (completamente, todo, de todo).
The COVID-19 crisis resulted in a large proportion of the world's population having to employ social distancing measures and self-quarantine. Given that limiting social interaction impacts mental health, we assessed the effects of quarantine on emotive perception as a proxy of affective states. To this end, we conducted an online experiment whereby 112 participants provided affective ratings for a set of normative images and reported on their well-being during COVID-19 self-isolation. We found that current valence ratings were significantly lower than the original ones from 2015. This negative shift correlated with key aspects of the personal situation during the confinement, including working and living status, and subjective well-being. These findings indicate that quarantine impacts mood negatively, resulting in a negatively biased perception of emotive stimuli. Moreover, our online assessment method shows its validity for large-scale population studies on the impact of COVID-19 related mitigation methods and well-being.
Kinship is a fundamental and universal aspect of the structure of human society. The kinship category of 'grandparents' is socially salient, due to grandparents' investment in the care of the grandchildren as well as to older generations' control of wealth and cultural knowledge, but the evolutionary dynamics of grandparent terms has yet to be studied in a phylogenetically explicit context. Here, we present the first phylogenetic comparative study of grandparent terms by investigating 134 languages in Pama-Nyungan, an Australian family of hunter-gatherer languages. We infer that proto-Pama-Nyungan had, with high certainty, four separate terms for grandparents. This state then shifted into either a two-term system that distinguishes the genders of the grandparents or a three-term system that merges the 'parallel' grandparents, which could then transition into a different three-term system that merges the 'cross' grandparents. We find no support for the co-evolution of these systems with either community marriage organisation or post-marital residence. We find some evidence for the correlation of grandparent and grandchild terms, but no support for the correlation of grandparent and cross-cousin terms, suggesting that grandparents and grandchildren potentially form a single lexical category but that the entire kinship system does not necessarily change synchronously.
-Currently, the coverage of the studies of language levels between related languages is expanding in terms of historical continuity in linguistics, thus much attention is paid to the problem of finding complicity in related languages. It is because of a sign system, which saved the nation’s history, culture, cognition in a lexical richness of each language. Linguistic signs that have emerged as the norm and widespread among the people in the structure of any language are very common in the language of other people. This feature is very often found inside the Turkic languages themselves. This is a key factor that shows historical relations between Turkic peoples. Integral selection and study of pronouns, which constitute a large-scale part of the lexical fund, their comparison from the point of view of a separate linguistic phenomenon, are the definition of language and cognition of a related ethnic group and is also considered a spiritual and cultural source of accurate information about their relationship. The article discusses the similarities and differences of Turkic pronouns through the study by the method of historical comparison of pronouns in the Turkic languages.
This article discusses business papers XII-XIII century from the city of Augsburg, which is located in the south of Germany. The norm of the modern German language went through several stages of formation before acquiring a unified standard and becoming the so-called Standardsprache. The city of Augsburg belongs to the East Bavarian dialect region and is located on the border of Bavaria and Swabia. Analysis of the written language of documents of the XII-XIII century provided information on the interaction of the features of both dialects (Bavarian and Swabian). In this study, 5 documents related to various taxes were considered, which indicate that they were written in Augsburg, as well as 3 documents in the Augsburg monastery. It is important that for the documents considered there is no characteristic sequence in writing, that is, we are talking about the absence of a spelling norm. Confirmation of this fact is also given in the article with examples from the materials studied. The study showed the presence of similar characteristics in all studied, which indicates their undoubted linguistic kinship. Despite this, there are also features that are characteristic exclusively for the southwestern part of Germany and separately for the southeast. An analysis of the German southern dialects makes it possible to trace the development trend of the German language in its holy language in a period that is closely connected with the history of the German people. The processes of synergy between dialects within the framework of one language are considered, which draws attention to the beginning of the formation of the first national language, and subsequently the national one. The study revealed that Augsburg became a kind of conductor of the Bavarian dialect in the eastern part of the Swabian dialect. The isoglosses studied (phonetic, morphological, lexical) showed that these dialects can be combined linguistically as southern and considered a feature of the Germanic (Yerminon) range. Despite some linguistic differences, a relative unity of linguistic traditions is noted, indicating a sufficient proximity of the dialects of the southwestern and southeastern parts of Germany in the XII-XIII centuries.
The article deals with the concepts of "insulting potential" and "insult" and their representation in con-flict-ridden texts. The author tries to answer the questions which an expert faces when conducting a linguistic expertise and states the necessity to take into account pragmatic factors, as well as the extralinguistic situation as a whole, when analyzing the fact of insult. The urgency of the topic can be attributed to the need for an in-depth study of the interaction of various definitions of offensiveness, creating difficulties in qualifying the legal norm "insult"..
Cette étude analyse les erreurs lexicales dans la production écrite de lycéens maltais étudiant le français L2 aux niveaux B1 et B2, en se focalisant sur les erreurs attribuables à l’influence de la L1, qui, à Malte, comprend le maltais, l’anglais et l’italien. Des réflexions sont faites sur la difficulté d’admettre une idéologie translinguistique, tolérante de l’appui fourni par la L1 dans l’écriture en L2, dans le contexte d’un examen à un niveau avancé, avec ses normes de correction linguistique. Un corpus de copies d’examen aux niveaux Avancé et Intermédiaire est fouillé pour les possibilités de phénomènes de transfert, catégorisés en cinq types, émanant de difficultés orthographiques, de choix de mots, ou sémantiques, ces dernières provoquant l’utilisation des faux-amis. Les résultats sont comparés aux conclusions d’études faites dans les cadres maltais et international. La fréquence des contacts linguistiques dans le corpus est probablement attribuable tant à l’alternance codique, comportement omniprésent à Malte, qu’à la nature même de la rédaction en L2, activité forcément bilingue. Des calculs statistiques permettent des comparaisons des fréquences de contact aux niveaux Avancé et Intermédiaire, entre les copies mieux notées et les moins bien notées, comme entre les tâches plus exigeantes et les tâches plus simples.
This study aims to examine how humour can be used as a communication strategy in a crisis communication work with an objective of creating crisis awareness among the target audience and through this, contribute to the research field of Strategic communication and digital media. Research concerning humour as a strategy combined with risk communication is yet limited and therefore this paper has the ambition to contribute with new knowledge about whether humour as a strategy is appropriate and successful when transmitting a preparatory crisis message, that can be seen as a topic difficult to relate with for the target audience. The empirical material is limited to a digital advertising campaign. The campaign was launched in December 2019 by the Swedish Civil Contingencies Agency (MSB) on Swedish television and social media channels and consists of three videos from the campaign. Based on theories that concerns national risk- and crisis communication, humour as a strategy, national humour and social norms, a multimodal critical discourse analysis (MCDA) has been implemented on the empirical material to find out whether the producer’s lexical choices, representation of the characters and power relations can contribute with knowledge that regards if humour can work as a strategy, in a situation where a crisis doesn’t exists yet. The result shows that humour can work as a strategy if it is being used with caution and if the producer takes the specific context, culture and the target audience's earlier experiences of crisis into consideration, when adapting the preparatory message.
In this article, we present an automatic semantic role labeling system in Persian consisting of two modules: argument identification for specifying argument spans and argument classification for categorizing their semantic roles. Our modules have been trained on Persian Proposition Bank in which predicate-argument information is manually added as a layer on top of Persian Dependency Treebank with about 30,000 sentences. Therefore, our system was trained on 216,871 verbal predicates and 42,386 nonverbal ones consisting of 40,813 nouns and 1,573 adjectives with 33 semantic classes. As a supervised method, we used maximum entropy for building an argument identifier that results in human-level accuracy of 99% and support vector machine for an argument classifier with an F1 of 84. Regarding both verbal and nonverbal predicates with an expanded role set, we achieved reasonable results.
The paper analyzes the loanwords from the American English the Japanese language considering the diachronic perspective in relation to historical and socio-cultural processes that took place in Japan. The periodization of the waves of penetration of such borrowings into the Japanese language system considering socio-cultural shifts in Japanese society is offered. The first wave, which can be dated from the end of the XIX cent. till 1930s, consists of the first borrowings-Americanisms, the penetration of which into the Japanese language is connected with the first systematic contacts between Japan and the USA, as well as the humanitarian aid of the USA of Japan after the Great earthquake of 1923. The second wave can be dated from 1940s till 1980s; during these years in the context of post-war American occupation, Japan became obsessed with American mass culture and, consequently, spread its own mass culture created on the basis of an American one. In the Japanese language system, this stage is characterized by an avalanche-like enrichment of the gairaigo lexical layer by borrowings-Americanisms, followed by the "digestion" of foreign words and their deeper integration into the system of the national language through the creation of pseudo-English words called waseieigo, as well as the spread of abbreviation. In the field of linguistics, the second stage is characterized by the beginning of scientific understanding of the significance of borrowings-Americanisms in the Japanese language and the analysis of the destructive role of these units for the language culture. The third wave of penetration of American-English borrowings is believed to be related to the proliferation of the Internet, the main language of which is English; accordingly, this stage can be dated to the 1990s until now. The main feature of the last wave is the adaptation of borrowings to the needs and norms of the national language, resulting in the activation of hybrid word formation and the creation of mixed units consisting of either a Japanese root and a borrowed affix, or vice-versa, or shortened foreign and Japanese words (hybrid abbreviation).
Le mot est au centre de toutes les attentions dans l’entier de l’œuvre de Marivaux. Si ce questionnement sémantique est sans doute au cœur de toute entreprise littéraire, il revêt une acuité particulière pour l’auteur dont le style a été nommé « marivaudage », terme dont le sens premier en dit long sur le caractère exacerbé de cette thématique. En effet, Marivaux exhibe le doute lexical, il exhibe la polysémie, creuse la verticalité du sens comme pour en révéler des strates inouïes, pour en épuiser les possibles. Dans cette œuvre complexe, le mot ne tient pas de l’heureuse évidence mais est sans cesse soumis au soupçon: soupçon d’une manipulation, soupçon d’un sens second, soupçon d’un emploi galvaudé; un scepticisme dans la fiction qui constitue sans doute la marque de la quête aléthique de son auteur, car Marivaux a pensé le mot, en écrivain et en philosophe; une quête dont quelques textes théoriques gardent la trace. Cette thèse se propose donc d’observer le pourquoi et le comment du fonctionnement sémantique dans l’œuvre de Marivaux, d’interroger le questionnement permanent autour du mot _ un mot mis en question dans son sens, remis en question dans ses applications au sein d’une interaction _, de scruter les rouages du mécanisme lexical propre à cet auteur et ce, en observant le contexte de production des œuvres, puis le travail sur la répétition du mot et enfin le mot pris dans un réseau de résonance à différents niveaux au sein de la phrase, au sein du discours et au sein du monde et des normes communicationnels qui soutiennent tout échange.
Для преодоления интерференции в переводе используются переводческие трансформации. В данной статье рассматривается один из видов таких трансформаций, а именно, лексические, суть которых заключается в замене переводимой лексической единицы словом или словосочетанием, которое реализует сему данной единицы исходного языка (ИЯ). Приемы лексических трансформаций рассматриваются на примере научных текстов по математике. Упомянутый вид текстов не использовался ранее \nв качестве материала для анализа интерференции в переводе. Опытный переводчик постарается сделать так, чтобы текст перевода (ТП) соответствовал нормам языка перевода (ПЯ), и при этом сохранил коммуникативное задание ИТ (исходный текст) (то, ради чего был создан оригинал), и, следовательно, снизить или преодолеть влияние интерференции.= To overcome the interference in translation, we use translation transformations. This article discusses \none of the types of such transformations, namely, lexical ones, the essence of which is to replace the translated \nlexical unit with a word or phrase that implements this unit sema of the original language (OL). The techniques \nof lexical transformations are considered on the example of scientific texts in mathematics. The mentioned type \nof texts was not used earlier as a material for the analysis of interference in translation. A professional translator \nwill try to make the text of the translation (TT) comply with the norms of the translation language (TL), while \npreserving the communicative task of the source text (ST) (for which the original was created), and, therefore, \nreduce or overcome the influence of interference.
When a child is suspected to be the victim or sole witness of a crime, the manner in which information is gathered from the child becomes critical. A child forensic interview is the guided conversation that a legal expert conducts to elicit reliable information from a child. To help substantiate child testimony, it is important to discern characteristics of truthful and deceptive behavior in these interviews. The work presented uses various machine learning algorithms to identify differences in the speech of children when they are lying or being truthful, particularly when they have been asked by a confederate to deceive an interviewer. Results show that vocabulary and psycho-linguistic norms of a child's language use, in response to directed questions, provide substantial information to outperform human adults in detecting truthful statements.
The article deals with the features of linguo-pragmatical contents and ways of its expression in the text of the professional ethics code. It is described the main linguo-pragmatical categories of the ethical code and come to light modal meanings and means of their expression. The ethical code represents the specific genre establishing norms and rules of the office (professional) behavior based on the conventional system of moral ideals. The main objective of the ethical code creation is prevention of conflict situations and an illegal behavior. Standard code of ethics is considered as the coherent text representing a dichotomizing division of a speech product into the dynamic process of the language activity and result of this activity. The significant elements of pragmatics of Standard code of ethics are some categories of a modality and an appraisal as they form pragmatical contents of texts of the similar genres. Means of expression of the modality are the verbs with the meaning of need, the lexical and phraseological means being based on the modal meaning of approval / disapproval. It is allocated the values which became a basis for formation of a certain corporate picture of the world of public servants of the Russian Federation and municipal employees, for example, impartiality, conscientiousness, correctness, tolerance, and respect. It is come to the conclusion that the understanding of a lexical meaning of a word and its pragmatical opportunities are defined with an individual (corporate) picture of the world and cannot coincide with nationwide.
Focus of the CONcreTEXT task is conceptual concreteness: systems were solicited to compute a value expressing to what extent target concepts are concrete (i.e., more or less perceptually salient) within a given context of occurrence. To these ends, we have developed a new dataset which was annotated with concreteness ratings and used as gold standard in the evaluation of systems. Four teams participated in this first edition of the task, with a total of 15 runs submitted.Interestingly, these works extend information on conceptual concreteness available in existing (non contextual) norms derived from human judgments with new knowledge from recently developed neural architectures, in much the same multidisciplinary spirit whereby the CONcreTEXT task was organized.
Conceptual concreteness and categorical specificity are two continuous variables that allow distinguishing, for example, justice (low concreteness) from banana (high concreteness) and furniture (low specificity) from rocking chair (high specificity). The relation between these two variables is unclear, with some scholars suggesting that they might be highly correlated. In this study, we operationalize both variables and conduct a series of analyses on a sample of > 13,000 nouns, to investigate the relationship between them. Concreteness is operationalized by means of concreteness ratings, and specificity is operationalized as the relative position of the words in the WordNet taxonomy, which proxies this variable in the hypernym semantic relation. Findings from our studies show only a moderate correlation between concreteness and specificity. Moreover, the intersection of the two variables generates four groups of words that seem to denote qualitatively different types of concepts, which are, respectively, highly specific and highly concrete (typical concrete concepts denoting individual nouns), highly specific and highly abstract (among them many words denoting human-born creation and concepts within the social reality domains), highly generic and highly concrete (among which many mass nouns, or uncountable nouns), and highly generic and highly abstract (typical abstract concepts which are likely to be loaded with affective information, as suggested by previous literature). These results suggest that future studies should consider concreteness and specificity as two distinct dimensions of the general phenomenon called abstraction.
Анализируется дисциплинарная практика региональных адвокатских палат субъектов Российской Федерации. При этом основное внимание уделяется проблемам соблюдения адвокатами России норм русского литературного языка, связанных с речевым этикетом: далеко не всегда профессиональный защитник являет своей речью образец нравственности, интеллектуальности, лаконичности, легкости и изящества. Анализ деятельности адвокатских палат в отношении дисциплинарных проступков профессиональных защитников показывает, что при рассмотрении аналогичных ситуаций выносятся разные решения. Это относится прежде всего к замечанию и предупреждению. При выборе дисциплинарной комиссией одной из двух указанных мер взыскания сложной задачей оказывается определение степени тяжести совершенного проступка. Юрислингвистический анализ конкретных адвокатских высказываний, в которых имеет место нарушение норм профессионального речевого этикета и которые стали вследствие этого предметом дисциплинарного разбирательства, показывает, что на степень инвективности вербального выражения влияет множество факторов, как лингвистических, так и экстралингвистических. Выдвигается идея о необходимости внедрения единой шкалы инвективных слов (со стороны Федеральной палаты адвокатов) как одного из важных, входящих в комплекс лексических, грамматических, орфоэпических, изобразительно-выразительных средств языка, оснований разграничения мер дисциплинарной ответственности адвокатов за нарушение Кодекса профессиональной этики в части, касающейся проявления уважения к суду и лицам, участвующим в деле (ст. 12), а также обязанности адвоката сохранять честь и достоинство в любой ситуации (ст. 9). Помимо создания единой шкалы инвективности слов и выражений, для комплексного анализа спорных речевых высказываний адвокатов в отдельных случаях предлагается привлекать специалистов-лингвистов и психологов в целях определения конкретной меры дисциплинарной ответственности. The disciplinary practice of the regional law chambers of the entities of the Russian Federation is being analyzed. At the same time, the main focus is on the problems of the Russian lawyers’ compliance with the norms of the Russian literary language related to speech etiquette: it is not always a professional advocate who is a model of morality, intellectuality, brevity, lightness and grace. An analysis of the activities of the bar in relation to the disciplinary misconduct of professional defense lawyers shows that different decisions are made in similar situations. This applies primarily to observation and prevention. In selecting a disciplinary commission as one of these two penalties, it is difficult to determine the severity of the offence committed. The jurislingistic analysis of specific legal statements, in which there is a violation of the norms of professional speech etiquette and which have become the subject of disciplinary proceedings as a result, shows that the degree of invectiveness of verbal expression is influenced by many factors, both linguistic and extralinguistic. The idea is to introduce a single scale of invective words (on the part of the Federal Chamber of Advocates) as one of the important part of the complex of lexical, grammatical, orthopedic, visually expressive means of language, grounds for delineating the disciplinary measures of lawyers for violation of the Code of Professional Ethics in respect of the court and persons involved in the case (v. 12), and the duty of counsel to maintain honor and dignity in any situation (art. 9). In addition to creating a single scale of invective words and expressions, it is proposed in some cases to involve linguistic specialists and psychologists to determine a specific measure of disciplinary responsibility for a comprehensive analysis of the controversial speech statements of lawyers.
Grammatical models which represent the hierarchical structure of chord sequences have proven very useful in recent analyses of Jazz harmony. A critical resource for building and evaluating such models is a ground-truth database of syntax trees that encode hierarchical analyses of chord sequences. In this paper, we introduce the Jazz Harmony Treebank (JHT), a dataset of hierarchical analyses of complete Jazz standards. The analyses were created and checked by experts, based on lead sheets from the open iRealPro collection. The JHT is publicly available in JavaScript Object Notation (JSON), a human-understandable and machine-readable format for structured data. We additionally discuss statistical properties of the corpus and present a simple open-source web application for the graphical creation and editing of trees which was developed during the creation of the dataset.
When a new phenomenon or an advance in technology originates in society, it is natural that new terms appear to refer to the phenomenon, technology, new use, and so on. The impact of this new disease, COVID-19 is so strong that no one has been able to foresee how long they will have to live under conditions of isolation and social distance. COVID-19 has changed our lives drastically. A new normal is required, and it is affecting various areas of daily life. For this reason, new terms have appeared and certain expressions have acquired greater relevance due to their use in a generalized context due to the pandemic.There are few published studies on the impact of COVID-19 on the Spanish language. This research analyzes this effect of coronavirus pandemic on the lexical inventory of Spanish. Its main objective is to analyze the difference in the use of terms related to the coronavirus, depending on the country or region in the Spanish-speaking world. To collect the data, an online survey was conducted using a semi-closed questionnaire. With the help of volunteers, 346 questionnaires were collected for this study. The attitude of users towards the use of the new terms is examined, as well as the discriminatory use of language around the pandemic. It is concluded that the norm will end up accepting neologisms and variants that Spanish-speakers have innovated and put into circulation in this situation.
<h3>Introduction</h3><br> BOLT Egyptian Arabic-English Word Alignment -- Conversational Telephone Speech Training was developed by the Linguistic Data Consortium (LDC) and consists of 153,171 words of Egyptian Arabic and English parallel text enhanced with linguistic tags to indicate word relations. <br> The DARPA <a href="https://www.ldc.upenn.edu/collaborations/current-projects/bolt"> BOLT</a> (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference. <br> <h3>Data</h3><br> The source data in this release consists of transcripts of Egyptian Arabic conversational telephone speech (CTS) from LDC's CALLHOME and CALLFRIEND collections (<a href="../../../LDC97S45">LDC97S45</a>, <a href="../../../LDC97T19">LDC97T19</a>, <a href="../../../LDC2002S37">LDC2002S37</a>, <a href="../../../LDC2002T38">LDC2002T38</a>, <a href="../../../LDC96S49">LDC96S49</a>) that were translated into English by professional translation agencies and annotated for the word alignment task. <br> The BOLT word alignment task was built on treebank annotation. Specifically, Egyptian Arabic source tree tokens were automatically extracted from tree files in LDC's BOLT Egyptian Arabic Treebank. Those tree files had been tagged for part-of-speech and syntactically annotated. That data was then aligned and annotated for the word alignment task. <br> The data profile broken down by character tokens, tree tokens and segments appears below: <br> <table border="1" cellpadding="5"><br> <tbody><br> <tr><br> <td>Language</td><br> <td>Genre</td><br> <td>Files</td><br> <td>Words</td><br> <td>Tree-tokens</td><br> <td>Segments</td><br> </tr><br> <tr><br> <td>Egyptian Arabic</td><br> <td>CTS</td><br> <td>176</td><br> <td>153,171</td><br> <td>215,896</td><br> <td>20,010</td><br> </tr><br> </tbody><br> </table><br> <h3>Acknowledgement</h3><br> This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred. <br> <h3>Samples</h3><br> Please view the following samples: <br> <ul><br> <li><a href="desc/addenda/LDC2020T05.arz.tkn.txt">Egyptian-Arabic Token Sample</a></li><br> <li><a href="desc/addenda/LDC2020T05.eng.tkn.txt">English Token Sample</a></li><br> <li><a href="desc/addenda/LDC2020T05.wa.txt">Word Alignment Token</a></li><br> </ul><br> <h3>Updates</h3><br> None at this time. </br> Portions © 1996, 1997, 2002, 2012-2015, 2020 Trustees of the University of Pennsylvania
The article considers the features and benefits of using distance learning in the educational process. The content of the concept of “distance learning” and the views of scientists on its interpretation are presented. It reveals the teaching experience of the discipline “Ukrainian language (by professional direction)” by means of distance learning at the Ukrainian Language Department of I. Horbachevsky Ternopil National Medical University. The materials for checking the level of medical specialties students’ knowledge in the aforementioned discipline are offered for the following themes: “Language standard. Orthoepic norms, norms of emphasis”, “Lexical aspect of medical professional language. Phraseological units in professional language”, “Terminology in professional communication. Lexical-semantic relations in scientific terminology. Features of Ukrainian medical terminology”, “Dictionaries in professional communication. Types of dictionaries, their function and role in enhancing language culture”, “Morphological aspect of Medical Professional Language”, “Syntactic aspect of Medical Professional Language”, “Features of Ukrainian language etiquette. Communicative qualities of language culture. Doctor’s Language Etiquette”, “Public speech and its genres”. It emphasizes the usage of detailed instructions for practical tasks and for submitting samples to each task. The research demonstrates that the teacher must take into account the goal, the age of the students, the professional orientation of the tasks. It underscores that the given tasks have a practical orientation and cannot comprehensively assess the student’s knowledge level in the study of a specific topic, but they aim to diversify distance learning, make it interesting. Undoubtedly, the usage of computer-based distance learning provides a learning process, so it needs to be refined and developed. This type of training encourages the specialist to look for new forms and methods of teaching the discipline, and the student to work independently, to work with different sources of information. It should be noted that the development of distance learning in the Ukrainian education system is perspective.
Stylistics is a science born from the works of Charles Bally at the start of the 20th century. It centers on literary texts. To decipher them, stylistics focuses on methods and concepts from the language sciences. It is therefore effective in bringing to light the various subtleties of the African novel which continues to regenerate from the resources of African language and linguistic heritage. This regeneration results in an Africanization of the writing language. It follows that re-lexicalization is at the heart of verbal praxis. It is materialized by the embedding of oral ethno-texts in the novel. This creativity is a form of discursive subversion which is characterized by a polyphonic and transgeneric aesthetics. Le-fils-de-la-femme-mâle by Maurice Bandaman conforms to this standard. The typographical configuration shows that the enunciative foundations of the work rests on the structure of the traditional African tale. As a result, the novel resists the norms of traditional storytelling. Furthermore, the insertion of the African tale into the structure of the novel is not based on any stable rule. The discourse is based on an intertextual practice that upsets the architectonics of the classic novel. It breaks with the canon and imposes another reception modality which provides obvious enjoyment to the reader. This work highlights all of the enunciative, linguistic and discursive phenomena that contribute to the authenticity of this work.
This thesis explores the interaction between emotions and visual perception using large scale spatial environment as the medium of this interaction. Emotion has been documented to have an early effect on scene perception (Olofsson, Nordin, Sequeira, & Polich, 2008). Yet, most popularly-used scene stimuli, such as the IAPS or GAPED stimulus sets often depict salient objects embedded in naturalistic backgrounds, or “events” which contain rich social information, such as human faces or bodies. And thus, while previous studies are instrumental to our understanding of the role that social-emotion plays in visual perception, they do not isolate the effect of emotion from the social effects in order to address the specific role that emotion plays in scene recognition – defined here as the recognition of large-scale spatial environments. To address this question, we examined how early emotional valence and arousal impact scene processing, by conducting an Event-Related Potential (ERP) study using a well-controlled set of scene stimuli that reduced the social factor, by focusing on natural scenes which did not contain human faces or actors. The study comprised of two stages. First, we collected affective ratings of 440 natural scene images selected specifically so they will not contain human faces or bodies. Based on these ratings, we divided our scene stimuli into three distinct categories: pleasant, unpleasant, and neutral. In the second stage, we recorded ERPs from a separate group of participants as they viewed a subset of 270 scenes ranked highest in each of their respective categories. Scenes were presented for 200ms, back-masked using white noise, while participants performed an orthogonal fixation task. We found that emotional valence had significant impact on scene perception in which unpleasant scenes had higher P1, N1 and P2 peaks. However, we studied the relative contribution of emotional effect and low-level visual features using dominance analysis which can compare the relative importance of predictors in multiple regression. We found that the relative contribution of emotional effect and low-level visual features (operationalized by the GIST model, (Oliva & Torralba, 2006)) had complete dominance over emotional effects (both valence and arousal) on most early peaks and areas under the curve (AUC). We also found out that affective ratings were significantly influenced by the GIST intensities of the scenes in which scenes with high GIST intensities were more likely to be rated as unpleasant. We concluded that emotional impact in our stimulus set of natural scenes was mostly due to bottom-up effect on scene perception and that controlling for the low-level visual features (particularly the GIST intensity) would be an important step to confirm the affective impact on scene perception.
Cite the source of the dataset as: Ferraz Gerardi, Fabrício and Reichert, Stanislav (2020) TuLeD: Tupían lexical database. Version 0.8. Tübingen: Eberhard-Karls University
Purpose We examined four measures of lexical diversity in the narratives of children with typical language development (TLD) and developmental language disorder (DLD) that comprised the normative sample of the Edmonton Narrative Norms Instrument (Schneider et al., 2005). The purpose was to document the properties of each measure with respect to variations in utterance and sample length, developmental trends, and group differences. Method The sample consisted of 377 picture-elicited, story generation transcripts from children with TLD ( n = 300) and DLD ( n = 77) aged 4–9 years. We extracted the moving-average type–token ratio (MATTR) and the number of different words from the full sample, from samples equated for the number of utterances, and from samples equated for the total number of words. Results MATTR was the only measure to show no relationships to utterance or sample length. All measures showed significant positive growth with age and significant groupwise differences between children with TLD and DLD. However, the magnitude of age effects and differentiation between groups varied considerably across measures. Across measures, there were significant differences in the number of children with DLD who were identified with low lexical diversity relative to their same-age peers in the TLD group. Conclusion The results of this study support the view that different measures of lexical diversity may be appropriate for different clinical purposes. It is important for clinicians to understand how measures of lexical diversity function in order to make educated choices among measures and ensure appropriate interpretation.
The ability to process language data has become fundamental to the development of technologies in various areas of human life in the digital world. The development of digitally readable linguistic resources, methods, and tools is, therefore, also a key challenge for the contemporary Slovene language. This challenge has been recognized in the Slovene language community both at the professional and state level and has been the subject of many activities over the past ten years, which will be presented in this paper. The idea of a comprehensive dictionary database covering all levels of linguistic description in modern Slovene, from the morphological and lexical levels to the syntactic level, has already formulated within the framework of the European Social Fund’s Communication in Slovene (2008-2013) project; the Slovene Lexical Database was also created within the framework of this project. Two goals were pursued in designing the Slovene Lexical Database (SLD): creating linguistic descriptions of Slovene intended for human users that would also be useful for the machine processing of Slovene. Ever since the construction of the first Slovene corpus, it has become evident that there is a need for a description of modern Slovene based on real language data, and that it is necessary to understand the needs of language users to create useful language reference works. It also became apparent that only the digital medium enables the comprehensiveness of language description and that the design of the database must be adapted to it from the start. Also, the description must follow best practices as closely as possible in terms of formats and international standards, as this enables the inclusion of Slovene into a wider network of resources, such as Open Linked Data, babelNet and ELExIS. Due to time pressures and trends in lexicography, procedures to automate the extraction of linguistic data from corpora and the inclusion of crowdsourcing into the lexicographic process were taken into consideration. Following the essential idea of creating an all-inclusive digital dictionary database for Slovene, a few independent databases have been created over the past two years: the Collocations Dictionary of Modern Slovene, and the automatically generated Thesaurus of Modern Slovene, both of which also exist as independent online dictionary portals. One of the novelties that we put forward together with both dictionaries is the ‘responsive dictionary’ concept, which includes crowdsourcing methods. Ultimately, the Digital Dictionary Database provides all (other) levels of linguistic description: the morphological level with the Sloleks database upgrade, the phraseological level with the construction of a multi-word expressions lexicon, and the syntactic level with the formalization of Slovene verb valency patterns. Each of these databases contains its specific language data that will ultimately be included in the comprehensive Slovene Digital Dictionary Database, which will represent basic linguistic descriptions of Slovene both for the human and machine user.
This study investigated the interference of Bahasa Indonesia passive voice norm on English sentence. There are many studies that investigated the interference of native language on the learning of target language. Most of the studies talked about interference in the level of lexical, grammatical, phonetic, syntactical, and many more. However, the study about interference of a norm have never been discussed before. Thus, it is important to conduct this study to give some prove that norm of languages may interfere language learning. This study involved 50 students of Tour and Travel Business Department at Sekolah Tinggi Pariwisata (STP) AMPTA Yogyakarta. The data was collected by giving students 3 sentences in Bahasa Indonesia and they had to write them in English. The sentences that the students had produced were compared to the correct one. The finding shows that most of the students� sentences were interfered by the norm of passive voice in Bahasa Indonesia. It is due to the lack of students� understanding toward the concept of passive voice norms in both of Bahasa Indonesia and English. Thus, the teacher must give clear explanation about the norm of passive voice in both of languages.
The article shows the importance of the competent approach to learning as a priority of the newest direction of educational activity, and also emphasizes the activation of its distance form, which served as a prerequisite for the use of various network resources, including educational platforms and services. These two most important emphases are recognized as the most important ones in the educational process, which has recently moved from the classroom to the indirect computer space due to objective reasons. Modern educational transformations caused the dynamics in the Ukrainian-speaking system, which is primarily implemented by the growing number of uses of terms-loanwords mainly from English, which require not only a clear definition but also codification. Special attention is paid to concepts characterized by orthographic diversity, which, on the one hand, shakes the norms of modern Ukrainian literary language, and on the other hand does not fully ensure the acquisition of spelling competence and the formation of oral and written skills of students. It is noted that such lexical items or term phrases are used in informative-cognitive, informative and advertising texts, they are operated mainly by educational portals, websites and mass communication, in particular industrial ones. The analyzed nominations, taking into account the degree of adaptation, are combined into three groups: 1) names, both in Latin and Cyrillic; 2) two-component nominations, one component of which is in Cyrillic and the other one is in Latin; 3) lexical items that in the Ukrainian online space have not yet shown a graphic redesign. It is found that these concepts are adapted to the phonetic and grammatical systems of the Ukrainian language in case of writing these language units in Cyrillic. It is observed that a significant part of such nominations occurred as a result of transcription and showed swings in spelling. In this regard, some recommendations are given for their writing, which will help to protect the language standard and avoid inconsistencies with established norms.
Summary This paper compares the romanization of Gaul in the 1st century BC and the gallicization of the island of Martinique during 17th-century French colonial expansion, using criteria set out by Muf- wene's Founder Principle. The Founder Principle determines key ecological factors in the formation of creole vernaculars, such as the founding populations and their proportion to the whole, language varieties spoken, and the nature and evolution of the interactions of the founding populations (also referred to as “colonization styles”). Based on the comparison, it will be claimed that new languages arise when a language undergoes vehicularization and subsequently shifts from one speech community to another. In other words, linguistic genesis would be a complicated case of language contact, where not only one, but sev- eral dialects of both superstrate and substrate varieties are involved, in a historical context where the identity function of language, or the norm, is overriden by the need to communicate. Research also indicates that language varieties spoken at the time of the shift did not pertain to normative usage, but to popular varieties, dialects, or both, since the emerging vernaculars - in Gaul, as well as in Martinique - preserved some of their phonological and lexical particularities.
The chapter addresses the concept of linguistic norm in the tradition of classical grammar and rhetoric, paying special attention to activities concerning standardization processes in the Romance languages. Since a clear distinction between a prescriptive and a descriptive point of view is not given in "traditional grammar", the latter is manifested in the form of grammatical treatises which often also aimed at offering norms for "correct" language use. As a consequence thereof, our contribution will be concerned with aspects relating to the realm of the history of language sciences and, at least partially, to the history of rhetoric. The period taken into consideration ranges from Latin antiquity (Cicero, Quintilian) to the middle of the 17th century (Vaugelas). The topics to be discussed were selected with regard to the significance of the respective protagonists in the history of ideas in (Latin and) Romance language standardization.
New language phenomena are driven by social and political shifts at the global level. Even though the traditional literary norm is being destroyed, these linguistic innovations fulfil a language compensatory function. Internet communication and the new speech processes found in it provoke a keen research interest and are extensively explored by linguists. Major global changes in our life (cloud-based technologies, ecology, post-truth, the problem of generations, Big Data, etc.) were bound to transform communication itself. Therefore, we see changes in genres, functional styles, texts and our traditional ideas of various forms of the Russian national language usage. The Russian Internet (Runet) reveals language potential, fulfils the compensatory function of the language filling in all the elements missing so far and language shortcomings (neologisms denoting feminine gender-specific job titles, deviant verbal forms, new structures in comparative forms of adverbs and adjectives, etc.). The speech system of the Internet communication should be considered not as a double-sided one (oral and written) but as a conceptually new digital form of language use. In the democratic environment of pluralism, tolerance and the freedom of language use, lexical and lexical-grammatical innovations, “the new vernacular”, irregular grammar and lexical collocability, as well as the direct and conscious intention to break the norm of the literary language, should be justified and deemed a manifestation of the compensatory language function. Special attention is given to the acute problem of fundamental transformations in teaching practice.
Immersive 360º virtual reality (VR) movies can effectively evoke a wide range of different emotional experiences. To this end, they are increasingly deployed in entertainment, marketing and research. Because emotions influence decisions and behavior, it is important to assess the user’s affective appraisal of immersive 360º VR movies. Knowledge of this appraisal can serve to tune media content to achieve the desired emotional responses for a given purpose. To measure the affective appraisal of immersive VR movies, efficient immersive and validated instruments are required that minimally interfere with the VR experience itself. Here we investigated the convergent validity of a new efficient and intuitive graphical (emoji-based) affective self-report tool (the EmojiGrid) for the assessment of valence and arousal induced by videos representing 360º VEs (virtual environments). Thereto, 40 participants rated their emotional response (valence and arousal) to 62 videos from a validated public database of 360º VR movies using an EmojiGrid that was embedded in the VE, while we simultaneously assessed their autonomic physiological arousal through electrodermal activity. The mean affective ratings obtained with the EmojiGrid and those provided with the database (measured with an alternative and validated instrument) show excellent agreement for valence and good agreement for arousal. The mean arousal ratings obtained with the EmojiGrid also correlate strongly with autonomic physiological arousal. Thus, the EmojiGrid appears to be a valid and immersive affective self-report tool for measuring VE-induced emotions.
The purpose of this study was to find out the correlation between the level of academic background and the rules of Hangul orthography by examining the compliance of Korean spelling. To examine this, six universities were divided into three divisions by level into A, B, and C, and the status of marking by university was compared with the free bulletin board of Everytime, a college student community. As a result of the survey, A-grade K universities best observed the Hangul Hangul orthography followed by C-grade A universities. The university that did not keep the Hangul orthography well was a C-grade D university. By educational level, the grade with the lowest mislabeling rate is grade A. It can be seen that the status of compliance with the lexical norms is related to each level of education. However, it cannot be generalized because there are not many vocabulary in common from six universities, but there is some correlation between knowledge and actual notation.
In the last decade, intralingual translation has started to gain momentum amongst a number of translation academics. Nevertheless, some types of intralingual translation remain largely undiscovered, such as the process of abridgement in the production of simplified versions of classic literary works (i.e. graded readers). This article subjects three chapters of the abridged version of And Then There Were None by Agatha Christie to qualitative analysis using Descriptive Translation Studies theory. The aim is to contribute to bridging a research gap in Translation Studies by examining the norms and laws governing the process of abridgement. Translation norms and laws are detected by situating the source and the target texts in their respective socio-cultural backgrounds and by analysing translation shifts. Relevant shifts are identified by means of a check-list of features elaborated on the basis of theory on graded readers, which classifies them into lexical, structural and information shifts. The results of the analysis showcase the vast research potential of intralingual translation for language learning purposes. Keywords: Intralingual translation, abridgement, graded readers, DTS, translation norms, translation laws, Agatha Christie, And Then There Were None
assumptions about relevant dimensions. We employed the RC method to visualize mental representations of self and examined their relationships with traits related to self-image. For this purpose, 110 participants (70 women) performed a two-image forced choice RC task to generate a classification image of self (self-CI). Participants perceived their self-CIs as bearing a stronger resemblance to themselves than did CIs of others (filler-CIs). Valence ratings of participants who performed the RC task (RC sample) and of 30 independent raters both showed positive correlations with self-esteem, explicit self-evaluation, and extraversion. Moreover, valence ratings of independent raters were negatively correlated with social anxiety symptoms. On the other hand, valence ratings of the RC sample and independent raters were not correlated with depression symptoms, trait anxiety, or social desirability. The results imply that mental representations of self can be properly visualized by using the RC method.