Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Many language teachers use Information and Communications Technology (ICT) in their classrooms to create tasks, quizzes, or polls with general online learning platforms. Few teachers have experience, however, of incorporating online corpus tools in their teaching or assessment practices. This paper will explore how autonomous learning can be fostered by gradually introducing freely available lexical databases, online collocation dictionaries, pronunciation guides, concordancers, N-gram extractors, and other text analysis tools for vocabulary building, skills practice, or self-checking. Tasks used with English as a Foreign Language (EFL) undergraduates and teacher trainees on a Master’s Teaching English as a Foreign Language (MA TEFL) course will be presented. I will also explain why having some familiarity with linguistics research can enable teachers to use these applications more meaningfully.
In this study, we explored how contextual information about threat dynamics affected the electrophysiological correlates of face perception. Forty-six healthy native Swedish speakers read verbal descriptions signaling an immediate vs delayed intent to escalate or deescalate an interpersonal conflict. Each verbal description was followed by a face with an angry or neutral expression, for which participants rated valence and arousal. Affective ratings confirmed that the emotional intent expressed in the descriptions modulated emotional reactivity to the facial stimuli in the expected direction. The electrophysiological data showed that compared to neutral faces, angry faces resulted in enhanced early and late event-related potentials (VPP, P300 and LPP). Additionally, emotional intent and temporal immediacy modulated the VPP and P300 similarly across angry and neutral faces, suggesting that they influence early face perception independently of facial affect. By contrast, the LPP amplitude to faces revealed an interaction between facial expression and emotional intent. Deescalating descriptions eliminated the LPP differences between angry and neutral faces. Together, our results suggest that information about a person's intentions modulates the processing of facial expressions.
This dataset in CSV format contains all books from the web http://books.toscrape.com which has been got using web scraping method in November 2020. The CSV file has 12 columns called each of them like: title, image, rating, description, category, UPC, producttype, priceextax, priceincltax, tax, availability, numberreviews. The project was born as a practice for a subject of the Master of Science (MSc) of Data Science at the Universitat Oberta de Catalunya (UOC).
In this paper we shall analyze several examples of morphological and syntactic calques taken from the Russian language, used in oral and written communication by the Romanian-speaking population in the historical province Bessarabia, the present-day Republic of Moldova. Unlike semantic calques, the morphological loan translations we have identified in source texts are less numerous, because the morphological structure of the language is not so receptive to foreign influences as its lexical structure. In terms of the morphological loan translations, we have chosen to contrastively analyze the forms resulting from calquing the reflexive diathesisin Russian. The constructions with an obligatory reflexive are taken after the Russian language, when, in fact, the literary Romanian language norm often recommends the use of the active voice.The syntactic loan translations we have identified are much more numerous as compared to the morphological ones and, most of the times, they reside in imitating case relationships according to the Russian language pattern
Words should not be treated as if all they do is fulfil roles; they can also have relationships. In our everyday talk, we normally give the meanings of words in terms of their relationships. The approach for measuring semantic relatedness between words is determined by means of a set of words that are closely related to them, and these are normally referred to as the set of determiner words. This article examines the measures of semantic relatedness to which words are associated by means of the variety of types of semantic relationships – such as synonymy, hyponymy, function and association. This relationship is considered within the context of African Wordnet (AWN), which provides South African indigenous languages with a platform to access a machine-readable lexical database organised by meaning. This article argues that computing semantic relatedness using African Wordnet could be helpful in developing an understanding of the meaning of related words for use in different African languages.
We present our contribution to the EvaLatin shared task, which is the first\nevaluation campaign devoted to the evaluation of NLP tools for Latin. We\nsubmitted a system based on UDPipe 2.0, one of the winners of the CoNLL 2018\nShared Task, The 2018 Shared Task on Extrinsic Parser Evaluation and SIGMORPHON\n2019 Shared Task. Our system places first by a wide margin both in\nlemmatization and POS tagging in the open modality, where additional supervised\ndata is allowed, in which case we utilize all Universal Dependency Latin\ntreebanks. In the closed modality, where only the EvaLatin training data is\nallowed, our system achieves the best performance in lemmatization and in\nclassical subtask of POS tagging, while reaching second place in cross-genre\nand cross-time settings. In the ablation experiments, we also evaluate the\ninfluence of BERT and XLM-RoBERTa contextualized embeddings, and the treebank\nencodings of the different flavors of Latin treebanks.\n
With the rapid development of big data and deep learning, breakthroughs have been made in phonetic and textual research, the two fundamental attributes of language. Language is an essential medium of information exchange in teaching activity. The aim is to promote the transformation of the training mode and content of translation major and the application of the translation service industry in various fields. Based on previous research, the SCN-LSTM (Skip Convolutional Network and Long Short Term Memory) translation model of deep learning neural network is constructed by learning and training the real dataset and the public PTB (Penn Treebank Dataset). The feasibility of the model's performance, translation quality, and adaptability in practical teaching is analyzed to provide a theoretical basis for the research and application of the SCN-LSTM translation model in English teaching. The results show that the capability of the neural network for translation teaching is nearly one times higher than that of the traditional N-tuple translation model, and the fusion model performs much better than the single model, translation quality, and teaching effect. To be specific, the accuracy of the SCN-LSTM translation model based on deep learning neural network is 95.21%, the degree of translation confusion is reduced by 39.21% compared with that of the LSTM (Long Short Term Memory) model, and the adaptability is 0.4 times that of the N-tuple model. With the highest level of satisfaction in practical teaching evaluation, the SCN-LSTM translation model has achieved a favorable effect on the translation teaching of the English major. In summary, the performance and quality of the translation model are improved significantly by learning the language characteristics in translations by teachers and students, providing ideas for applying machine translation in professional translation teaching.
We propose the Graph2Graph Transformer architecture for conditioning on and predicting arbitrary graphs, and apply it to the challenging task of transition-based dependency parsing. After proposing two novel Transformer models of transition-based dependency parsing as strong baselines, we show that adding the proposed mechanisms for conditioning on and predicting graphs of Graph2Graph Transformer results in significant improvements, both with and without BERT pre-training. The novel baselines and their integration with Graph2Graph Transformer significantly outperform the state-of-the-art in traditional transition-based dependency parsing on both English Penn Treebank, and 13 languages of Universal Dependencies Treebanks.
The paper provides an overview of the treatment of normative labelled linguistic phenomena in a monolingual general explanatory dictionary with an informative and normative charac- ter, particularly in the first and second editions of the Dictionary of Standard Slovenian (SSKJ and SSKJ2). It focuses on the linguistic phenomena that are labelled in the first edition of the Dictionary of Standard Slovenian with an explicit normative usage note “incorrectly”. It compares the normative evaluation of these linguistic phenomena in the first edition with their normative evaluation in the second edition (2014). The purpose of the comparison is to find out how the normative evaluation of linguistic phenomena changed over time. By observing the lexicographical marking of linguistic phenomena, we can trace changes in linguistic norms and in the different views on labelling in dictionaries. A comparative review of the linguistic phenomena in SSKJ and SSKJ2 labelled with the normative usage label “incorrectly” shows that the normative evaluation of linguistic phenomena over time has become less stringent or, put differently, that linguistic phenomena in standard language have become more normatively acceptable. It also shows that the Slovenian monolingual general explanatory dictionary is moving away from explicit normative labelling of linguistic phenomena.
The article provides an overview of the lexical and grammatical features as well as the sociopolitical environment of Marollien that originated in the 18th century as a dialect on the territory of Brussels. Marollien is essentially the Dutch language in its Brabantian dialect, strongly influenced by French. There are literary works, performances, and musicals written and staged in Marollien, as well as dictionaries and journals published in it. Historically, the Marollien dialect is a sociolect: it was generally used by Belgians coming to Brussels from Wallonia in search of a job and settling in one of the districts of Brussels — Marolles. A special emphasis is placed on lexical features of the dialect: gastronomic and everyday vocabulary are looked at and the examples of French loanwords and Southern Dutch language norm deviations are provided. Standard Dutch calques in French, when translating idioms in particular, are also identified. The differences between Dutch, French, and Marollien place names are illustrated. In the field of morphology and word formation, there is a regular mixture of Germanic and Romanic stems which is indicated. Examples of Marollien phonetic features are also provided. The article acknowledges frequent code switching in Marollien speech, which by and large resembles the phenomenon of linguistic interference. Due to the fact that Marollien is rapidly disappearing, the Brussels-Capital region is trying to support the dialect: various activities are being organized in order to propagate its use and enhance its prestige. Nevertheless, Marollien is not included in the well-known citizen initiative “Marnix Plan”, aimed at developing the methodology for the sequential study of several languages for all segments of the population in Brussels. This initiative is also discussed in the article.
This study investigated EFL learners’ comprehension of scalar properties of three types of emotion verbs, namely, fear type, liking and disliking emotion verbs and compare their performance with instructors and native speakers of English. The participants were 38 non-native pre-service teachers from ELT department at a state university in Turkey, 11 ELT instructors at different universities and 10 native speakers from the USA and the UK. The data were collected through a scale construction task according to participants’ judgements on scalar emotion verbs in terms of their relative order on a linear scale. The results revealed that in terms of constructing consistent scales with with previously determined scales in literatutre, pre-service teachers performed poorly for fear-type and disliking emotion verbs, they were partly successful in constructing consistent scales for liking verbs. It was also found that similarly instructors performed poorly in constructing scales for fear-type and disliking verbs, but they were better than pre-service teachers. They were also successful in constructing scales for liking verbs. Native speakers were successful in fear-type and liking verbs; however, like non-native participants, they performed poorly in constructing consistent scales for disliking verbs. This means that there may be cross-cultural differences among participants’ judgement of emotion verbs on a linear scale in terms of their intensity. This study may provide valuable information for the studies on lexical resources (e.g., VerbNet, WordNet etc.) Previous studies (e.g. Fellbaum Mathieu,2014; Sheinman, Tokunaga, 2009) show a way to represent the scalar properties of emotion verbs in WordNet, and other possible extensions to additional verb families can cause a more subtle semantic analysis of emotion verbs in lexical databases with potential benefits for automatic inferencing, language pedagogy and translation. This study may contribute to semantic analysis of emotion verbs in lexical databases. It may also provide some implications for students, language teachers, and policy makers in terms of vocabulary learning and teaching.
English WordNet is an important synonym set to present the similarity of meanings between words. Synonym Set is built using Oxford Thesaurus which is accessed through lexico.com, which is a part of the lexical database that will be used. After using the extraction process through Oxford Thesaurus it will produce a synonym set with the same meaning between words. The difference between WordNet and ordinary dictionaries is that the word is interconnected with other words. One method employed for this approach is Robust Clustering Using Links method, which is similarity values and synonym sets that have been created to be used to build a lexical database. Therefore the main purpose of the development of the English WordNet is to produce an accurate synonym set using clustering techniques. The evaluation calculation will use the F-measure method and will use the gold standard for the calculation method. With the ROCK method, there is an increase in accuracy output from dataset input. Building the English wordnet is to improve words that can be used to help research and development of other language wordnets with role models using more accurate English wordnets. And the use of ROCK method there is an increase in the accuracy upon results of the development of English wordnet compared to the previous method, which is using hierarchical clustering. The outcome of this study resulted in improved accuracy so that the ROCK method is one of the good methods used in the development of the English wordnet.
This paper attempts to measure the similarity of frequently occurring modal auxiliaries in both L1 and L2 writings. The modal auxiliaries are known to be difficult areas of study for both L1 and L2, due to the overlapping in their use. A subset of TOEFL11 corpus and a subset of PTB(Penn Treebank) corpus were used to capture the similarities among modal auxiliaries within L1 and L2, respectively, and also across L1 and L2. Based on the hypothesis of distributional representation, similarities of modal auxiliaries were computed by applying cosine similarity and mutual information to every pair of words in the corpus, and then finally reducing dimensions with Singular Value Decomposition. The results show that the modals are found to be similar to other modals, and the distribution of modals in the reduced dimensions show that the modals in L1 and those in L2 exhibits different patters of similarities, while the distance among the modals in L1 is further away than the distance among those L2 modals.
The article deals with reproducing contaminated speech of literary characters in the Ukrainian translation s of J. Steinbeck’s novel "East of Eden" and the play "Hysteria" by Т. Johnson. Contaminated speech (a foreign accent) is a wide-spread stylistic means of characterizing literary characters. The specific nature of the accent remains under the dispute among researchers in translation studies. Some scholars regard the accent as a phonetic phenomenon only. However, our research is based on V. Vinogradov’s statement that accent covers the speech deviations on all language levels. In addition, the important aspect of investigating accents is understanding that an accent is a product of language interference. The literature review contains the works of theoreticians and practitioners of translation studies (A. Fedorov, S. Florin, V. Komissarov, O. Kopyl’na, T. Nekryach, O. Rebrii, S. Vlakhov, etc.). The research relevance is determined by anthropocentric approach to text interpretation, which allows to investigate literary characters as language personalities. The object of the research is the constituents of Spanish and Chinese accents of literary characters. The subject of the research comprises the translation tactics of reproducing the foreign accents. The purpose of the research is to reveal the specific traits and components of the foreign accent of literary characters and to single out the optimal translation tactics for its faithful reproduction in translation. To achieve the aim, the following objectives have been set: to examine the specific features of foreign accents of literary characters and to highlight the specific translation tactics of rendering foreign accents. Results of research. Creative writers resort to imitating foreign accents to create original characters and emphasize their foreign identity. The specific traits of a foreign accent can be represented in fiction in the form of deviations from the literary norm on the phonetic, grammatical, and lexical levels. Reproducing these specific traits can be regarded as a translation challenge difficulties and requires special skill s on the part of translator. According to V. Komissarov, the natural sounding of a foreigner’s speech should be the top priority in accent reproduction. The Ukrainian translations of T. Johnson’s play "Hysteria" and J. Steinbeck’s novel "East of Eden" demonstrate successful reproducing of the Spanish and Chinese accents. While creating the Spanish accent of his character T. Johnson resorted to incorrect use of words, grammatical mistakes, distortion of the phonetic and morphological structure of words. These features of the character’s speech make an integral component of his language personality, which is carefully conveyed by the Ukrainian translator. In order to preserve all motivated deviations from the literary norm, the translator uses the parallel translation. In the absence of appropriate equivalents the translator resorts to compensation by applying distortion of phonetic, morphological, and syntactic norms. All the translation means represented in this translation seem to be optimal for creating a stage dialogue. Salvador Dali’s Spanish sounds natural and easy for audial perception, which is the most important in drama translation. While translating the Chinese accent, Tetiana Nekryach resorts to the parallel translation of phonetic distortions of words. The translation tactics of compensation is applied for rendering the grammatical mistakes in the character’s speech. All translation means are used for embodiment the author’s intent and increasing the impact on readers. Thus, the Chinese accent are reproduced on all language levels, including phonetic, grammatical, and lexical. Originality. The study of approaches to conveying foreign accents of literary characters, considering the specific nature of the accent, represents the research originality. Thus, the phonetic, grammatical, and lexical components of the foreign accents were investigated, and the optimal translation tactics of their rendering were singled out. Conclusions. In conclusion, it should be mentioned that due to the differences in structures of the target and source languages, it is impossible to fully reproduce contaminated speech on the all language levels in translation. Therefore, the translator should find the functional analogues for faithful reproducing a foreign accent. Among the optimal translation tactics of accent reproduction, the parallel translation and compensation are singled out. The perspective of the research lies in further investigation of contaminated speech and comparison between the translation tactics applied in drama and prose translation.
Il presente contributo propone un primo confronto linguistico tra due testi franco-italiani e testi in francese, italiano e veneto antichi, avvalendosi dell’uso di corpora dotati di annotazione morfosintattica secondo i criteri adottati dal Penn Treebank Project. I testi franco-italiani investigati sono l’Entrée d’Espagne e l’Orlandino della Geste Francor. Lo studio esplora la sintassi delle frasi relative nei due testi franco-italiani, al fine di offrire una descrizione preliminare del loro sistema linguistico. L’analisi evidenzia la natura ibrida del sistema linguistico dei due testi franco-italiani, individuando tratti condivisi con le lingue romanze più prossime, nonché tratti peculiari del franco-italiano. This work represents a first attempt to compare the linguistic system of two Franco-Italian texts with texts in Old Italian, Old Venetan, and Old French, by using morpho-syntactically annotated corpora following the guidelines of the Penn Treebank Project.
The Index Thomisticus Treebank (IT-TB) is the syntactically annotated portion of the Index Thomisticus corpus (Busa, 1974-1980), which collects the opera omnia of Thomas Aquinas. The IT-TB includes the “analytical” (i.e. surface syntactic) annotation of the entire “Summa contra Gentiles” (4 books), as well as of the concordances of lemma “forma” from “Scriptum super libros sententiarum magistri Petri Lombardi” (entire) and from “Summa Theologiae” (part). In total, the IT-TB consists of more than 450,000 nodes in around 25,000 sentences. The annotation guidelines, inspired by those of the analytical layer of the Prague Dependency Treebank, are available here: http://static.perseus.tufts.edu/docs/guidelines.pdf. The IT-TB features also more than 2.000 sentences from “Summa contra Gentiles” annotated at the tectogrammatical (i.e. underlying syntactic) layer. The annotation guidelines are those of the Prague Dependency Treebank: https://ufal.mff.cuni.cz/pdt2.0/doc/manuals/en/t-layer/pdf/t-man-en.pdf.
CzeDLex 0.7 is the third development version of the Lexicon of Czech discourse connectives. The lexicon contains connectives partially automatically extracted from the Prague Discourse Treebank 2.0 (PDiT 2.0) and, as a supplementary resource, the Czech part of the Prague Czech–English Dependency Treebank with discourse annotation projected from the Penn Discourse Treebank 3.0. The most frequent entries in the lexicon (131 out of total 218 entries, covering more than 95% of discourse relations annotated in PDiT 2.0), have been manually checked, translated to English and supplemented with additional linguistic information.
Reading comprehension is the process of reader’s interaction with a written language and making meaning from the text. Studies show that there is a great difference between skilled and poor readers in terms of using the fundamental strategies of reading comprehension, i.e., cognitive and metacognitive strategies. Due to the lack of appropriate tool for measuring adults’ reading comprehension level in Iran and the lack of appropriate criteria for selecting the texts for such a tool, the aim of this study was to make a placement tool for evaluating adult Persian speakers’ reading comprehension level. It was amixed methods study and the aim was to answer three main questions about finding a text selection criterion for reading comprehension tests, using the selected criteria for making reading comprehension placement test and studying the validity and reliability of the devised test. Findings showed that text selection criteria should be fitted with patterns of international tests and the linguistic principles of text. The content validity of the text was approved by experts, after implementing their comments on the texts and questions. To ensure the test reliability and to do item analysis, the test was distributed among a sample of 60 MA students of University of Guilan at two stages. The reliability of the test was computed and Cronbach’s alpha were 0.84 and 0.82 respectively, showing the appropriate reliability of the tool. After normalization, this tool can be used to evaluate adult reading comprehension and in educational planning, it can be used for selecting the educational content. Introduction: Reading comprehension is the process of reader’s interaction with a written language and drawing meaning from the text. Generally, reading comprehension is a complex and multidimensional process which is done through two core processes. The first is decoding the symbols and recognizing the words, and the second is integrating the meaning of words in the context of the text (Gough & Tunmer, 1986; in: Atkinson, 2014). Learning reading comprehension is a long-term process; so it is at the end of the learning process that the adult reader can easily read different texts and draw the meaning from them. Studies show that there is a great difference between skilled and poor readers in terms of using the fundamental strategies of reading comprehension, i.e., cognitive and metacognitive strategies (Cain, Oakhill & Bryant, 2000). Weakness in the prerequisites of reading comprehension and failure in selecting the appropriate comprehension strategies are some of the important problems of students at different educational levels when reading different types of texts. Some international studies have been done on the reading comprehension such as PIRLS[1] and PISA[2] tests. During the recent years, the number of such studies has increased in Iran. Perhaps it could be due to the Iranian students’ low performance in the PIRLS test at different time intervals which shows their weakness in reading comprehension. Despite such weak results in international tests, and doing some related research in Iran, still there is no appropriate tool for determining the reading comprehension of Iranian people, especially adults. Living in the modern society needs learning and reading various texts. Despite the importance of this issue and the quantitative progress in the number of graduate students, there is no specific criterion to determine the educated adults’ level of reading comprehension. The development of higher education is a great scientific evolution that despite its positive effects has some shortcomings as well. One of the most important shortcomings is the lack of an appropriate placement tool for evaluating students’ reading comprehension in order to prepare suitable educational material. The aim of this study was to develop a placement tool for evaluating adult Persian speakers’ reading comprehension. The study followed three main objectives, i.e. finding text selection criteria and the related questions for reading comprehension tests, using the selected criteria for developing reading comprehension placement test and finally determining the validity and reliability of the designed test. Questions: There were three main questions in this study: 1. What are the text and question selection criteria for developing a reading comprehension placement test? 2. Which reading comprehension placement test could be designed for adults, implementing the above-mentioned criteria? 3. Does the designed test have validity and reliability? Method: It was amixed methods study.The qualitative part included finding the text and question selection criteria for developing the adult Persian speakers’ reading comprehension placement test and the steps of its development. Also, the quantitative part of the study included the pilot study of the mentioned test to determine its validity and reliability. The content validity of the test was checked by 5 experts. To examine its reliability, the test was distributed among two groups of MA students of the University of Guilan, who were selected using convenience sampling method (each group 30 students) with two months interval and the level of reliability was computed using Cronbach’s alpha. Along with calculation of the reliability of the test, the test items were analyzed in terms of item facility, item discrimination, and the distractors’ distribution. Findings: To select the texts, a combination ofcriteria introduced by the International Institute for the study of Reading Literacy for PIRLS test and Educational Testing Service (ETS) has been used. It was tried to match the selected texts in accordance with patterns of international tests and linguistic characteristics of the texts. These criteria included lexical and grammatical cohesion and also coherence of the texts. Taking into account all of the strategies underlying reading comprehension (i.e., inference-making, comprehension monitoring, text structure, etc.) the test questions were designed at 6 levels. These levels were selected based on the integration of Day and Park (2005) model and the design of TOFEL tests for reading comprehension placement tests. The content validity of the test was approved by 5 linguists, English language teaching, and Persian language teaching experts, after implementing their comments on the texts and questions. To ensure the test reliability, Cronbach’s alpha coefficient was used. First, the test was distributed among a sample of 30 MA students of University of Guilan and the level of Cronbach’s alpha was 0.82. Also, different levels of item analysis were conducted, including item facility, item discrimination, and the distractors’ distribution. To make sure of the reliability of the test, after revising the items, and with two months interval, the test was distributed among a new sample of 30 MA students of the University of Guilan. Again, the reliability of the test was computed and Cronbach’s alpha was 0.84, showing the appropriate reliability of the tool. Conclusion: This test could be used to assess adult Persian speakers’ level of reading comprehension and also its results could be used to select appropriate educational material. In the next step of this research, this tool should be distributed among a larger sample to determine its construct validity and also to compute its norms so that its results could be cited with more confidence. [1] The Progress in International Reading Literacy Study [2] The Program for International Student Assessment
The television ratings provide an effective way to analyze the popularity of TV programs and audiences’ watching habits. Most previous studies have analyzed the ratings from a single perspective. Few efforts have integrated analysis from different perspectives and explored the reasons for changes in ratings. In this paper, we design a visual analysis system called TVseer to analyze audience ratings from three perspectives: TV channels, TV programs, and audiences. The system can help users explore the factors that affect ratings, and assist them in decisions about program productions and schedules. There are six linked views in TVseer: the channel ratings view and program ratings view show ratings change information from the perspective of TV channels and programs respectively; the overlapping program competition view and the same-type program competition view indicate the competitive relationships among programs; the audience transfer view shows how audiences are moving among different channels; the audience group view displays audience groups based on their watching behavior. Besides, we also construct case studies and expert interviews to prove our system is useful and effective.
The emotions we experience shape our perception, and our emotion is shaped by our perceptions. Taste perception is also influenced by emotions. Positive and negative emotions alter sweetness, sourness, and bitterness perception. However, most previous studies mainly explored valence changes. The effect of arousal on taste perception is less studied. In this study, we asked volunteers to watch positive affect inducing videos with high arousal and low arousal. Our results showed a successful induction of high and low arousal levels as confirmed by self-report and electrophysiological signals. Moreover, self-report affective ratings did not show a significant effect on self-reported taste ratings. However, we found a negative correlation between smile occurrence and sweetness ratings. In addition, EDA scores were positively correlated with saltiness. This suggests that even if the self-reported affective state is not granular enough, looking at more fine-grained affective cues can inform ratings of taste.
In recent years, natural language processing (NLP) has become one of the most important areas with various applications in human's life. As the most fundamental task, the field of word embedding still requires more attention and research. Currently, existing works about word embedding are focusing on proposing novel embedding algorithms and dimension reduction techniques on well-trained word embeddings. In this paper, we propose to use a novel joint signal separation method - JIVE to jointly decompose various trained word embeddings into joint and individual components. Through this decomposition framework, we can easily investigate the similarity and difference among different word embeddings. We conducted extensive empirical study on word2vec, FastText and GLoVE trained on different corpus and with different dimensions. We compared the performance of different decomposed components based on sentiment analysis on Twitter and Stanford sentiment treebank. We found that by mapping different word embeddings into the joint component, sentiment performance can be greatly improved for the original word embeddings with lower performance. Moreover, we found that by concatenating different components together, the same model can achieve better performance. These findings provide great insights into the word embeddings and our work offer a new of generating word embeddings by fusing.
Instructions highlighting that backward conditional stimuli (CSs) stop unconditional stimuli (USs) result in their acquiring valence opposite to that of the US on explicit measures of valence. We assessed whether such instructions would influence startle blink modulation in the same way. Two groups were presented with concurrent forward and backward evaluative conditioning (CS-US-CS) using cartoon aliens as CSs, and pleasant, neutral, and unpleasant sounds as USs. Startle magnitude was measured during conditioning and valence ratings were assessed after conditioning. Participants in the "start-stop" instructions group (n = 41) were instructed to learn whether CSs started or stopped US presentations, while participants in the "observe" instructions group (n = 41) were told to pay attention to the stimuli as they would be asked questions about them after the experiment. In the start-stop instructions group backward CSs paired with positive USs were rated as less pleasant than backward CSs paired with neutral and negative USs (contrast effect), whereas ratings of backward CSs did not differ in the observe instructions group. Startle magnitude was larger during backward CSs paired with positive USs in comparison to CSs paired with neutral or negative USs in both instruction groups. Startle blink modulation was unaffected by instructions, suggesting that startle indexes the emotional state at the time of probe presentation rather than CS valence based on propositional information about the function of the CS.
In this paper, we aim to introduce the dependency annotation process of the largest and the only cross-linguistic Turkish dependency treebank which was translated from the original Penn Treebank corpus. Within the scope of this project, 16.400 sentences have been morphologically and semantically annotated, and the dependency relations were manually carried out by a team of linguists. It is hoped that this project will serve as a base for a successful dependency parser and a system which can automatically perform the bi-directional conversion between constituency and dependency trees.
The present study asks whether behaviors of another person can be intentionally forgotten, and whether forgetting affects how that person is evaluated. Participants read about negative and neutral behaviors of a fictional character and were then asked to forget or to keep remembering them. Afterwards, all participants learned of neutral behaviors associated with another character. After a short distractor or a 24-hour delay, implicit and explicit evaluations of both characters and memory for their behaviors were assessed. Implicit evaluations did not differ between the two characters and were insensitive to forget instructions. In comparison to the remember condition, participants who were told to forget recalled fewer of the first character's behaviors, and explicitly judged him as warmer and less dominant. Importantly, memory was negatively correlated with warmth and positively correlated with dominance ratings. The study shows that intentional forgetting can have long-lasting consequences for memory and explicit judgments of others.
Dependency parsing is a longstanding natural language processing task, with its outputs crucial to various downstream tasks. Recently, neural network based (NN-based) dependency parsing has achieved significant progress and obtained the state-of-the-art results. As we all know, NN-based approaches require massive amounts of labeled training data, which is very expensive because it requires human annotation by experts. Thus few industrial-oriented dependency parser tools are publicly available. In this report, we present Baidu Dependency Parser (DDParser), a new Chinese dependency parser trained on a large-scale manually labeled dataset called Baidu Chinese Treebank (DuCTB). DuCTB consists of about one million annotated sentences from multiple sources including search logs, Chinese newswire, various forum discourses, and conversation programs. DDParser is extended on the graph-based biaffine parser to accommodate to the characteristics of Chinese dataset. We conduct experiments on two test sets: the standard test set with the same distribution as the training set and the random test set sampled from other sources, and the labeled attachment scores (LAS) of them are 92.9% and 86.9% respectively. DDParser achieves the state-of-the-art results, and is released at https://github.com/baidu/DDParser.
This study aims to discuss the ethics of public communication in a digital age, based on the hadith perspective. It is based on the poor ethics of communication in society, which leads to the emerge of SARA, hoaxes, and politicization of religion. The article addresses prophetic traditions concerning ruwaybid}ah that could provide values on the norms of communication. Based on the lexical meaning, the word ruwaybid}ah means people who are weak towards noble matters. Muhammad’s description through the editorial “fools who meddled in community affairs” has a close relation to the cause of social inequality at that time. Through a socio-historical approach, it argues that anyone who does not have capabilities in a particular field and tries to impose his or her incapacity, they could belong to the ruwaybid}ah category. Further, this study attempts to formulate key important principles of today’s digital age communication based on the hadith.
This paper adds to the available resources for the under-resourced language Urdu by converting different types of existing treebanks for Urdu into a common format that is based on Universal Dependencies. We present comparative results for training two dependency parsers, the MaltParser and a transition-based BiLSTM parser on this new resource. The BiLSTM parser incorporates word embeddings which improve the parsing results significantly. The BiLSTM parser outperforms the MaltParser with a UAS of 89.6 and an LAS of 84.2 with respect to our standardized treebank resource.
Affective vocalisations such as screams and laughs can convey strong emotional content without verbal information. Previous research using morphed vocalisations (e.g. 25% fear/75% anger) has revealed categorical perception of emotion in voices, showing sudden shifts at emotion category boundaries. However, it is currently unknown how further modulation of vocalisations beyond the veridical emotion (e.g. 125% fear) affects perception. Caricatured facial expressions produce emotions that are perceived as more intense and distinctive, with faster recognition relative to the original and anti-caricatured (e.g. 75% fear) emotions, but a similar effect using vocal caricatures has not been previously examined. Furthermore, caricatures can play a key role in assessing how distinctiveness is identified, in particular by evaluating accounts of emotion perception with reference to prototypes (distance from the central stimulus) and exemplars (density of the stimulus space). Stimuli consisted of four emotions (anger, disgust, fear, and pleasure) morphed at 25% intervals between a neutral expression and each emotion from 25% to 125%, and between each pair of emotions. Emotion perception was assessed using emotion intensity ratings, valence and arousal ratings, speeded categorisation and paired similarity ratings. We report two key findings: 1) across tasks, there was a strongly linear effect of caricaturing, with caricatured emotions (125%) perceived as higher in emotion intensity and arousal, and recognised faster compared to the original emotion (100%) and anti-caricatures (25%-75%); 2) our results reveal evidence for a unique contribution of a prototype-based account in emotion recognition. We show for the first time that vocal caricature effects are comparable to those found previously with facial caricatures. The set of caricatured vocalisations provided open a promising line of research for investigating vocal affect perception and emotion processing deficits in clinical populations.
The current study looked at the impact of British regional accents on evaluations of eyewitness testimony in criminal trials. Ninety participants were randomly presented with one of three video recordings of eyewitness testimony manipulated to be representative of Received Pronunciation (RP), Multicultural London English (MLE) or Birmingham accents. The impact of the accent was measured through eyewitness (a) accuracy, (b) credibility, (c) deception, (d) prestige, and (e) trial outcome (defendant guilt and sentence). RP was rated more favourably than MLE on accuracy, credibility and prestige. Accuracy and prestige were significant with RP rated more highly than a Birmingham accent. RP appears to be viewed more favourably than the MLE and Birmingham accents although the witness’s accents did not affect ratings of defendant guilt. Taken together, these findings show a preference for eyewitnesses to have RP speech over some regional accents
В современных русских говорах, несмотря на воздействие на них норм литературного языка, межкультурную контактность и деформацию в результате воздействия средств массовой информации, постепенное исчезновение диалектов в условиях цивилизации медиа, сохраняется лексическое ядро, в котором особое место занимает лексика нравственно-религиозной сферы. Наличие этого лексического пласта выделяет русский язык среди других в отношении аксиологической акцентированности земного и небесного. В настоящей статье обобщены наблюдения автора, касающиеся состава, семантики, сложения лексических гнезд, составляющих макрополе народного православия, формальной стороны словопроизводства единиц соответствующей лексики на общерусском фоне с привлечением материала, собранного автором статьи в Тамбовской области. Выделены разновидности структурных типов номинаций, определены особенности и приоритеты словообразовательной креативности. Для достижения целей исследования применялись методы сопоставительного, лексического, словообразовательного, компонентного анализа. Despite the influence exerted on them by the literary norms of the Russian language, despite the intercultural contacts promoted by mass media, and despite gradual annihilation of dialects provoked by media-propelled civilization, modern Russian dialects preserve their pivotal lexemes which are mostly related to the sphere of morality and religion. This lexical stratum, the axiological significance of the celestial and the earthly realms, distinguishes the Russian language from other languages. The present article summarizes the author’s ideas related to word formation, to the composition and semantics of lexical clusters of the macrofield of public orthodoxy. The analysis involves the material collected by the author in the Tambov Region. It singles out structural types of nominations, defines the peculiarities and priorities of creative word-formation. To achieve the aim of the research, the author employs such methods as comparative analysis, lexical analysis, word-formation analysis, and componential analysis.
Ambiguous stimuli are useful for assessing emotional bias. For example, surprised faces could convey a positive or negative meaning, and the degree to which an individual interprets these expressions as positive or negative represents their “valence bias.” Currently, the most well-validated ambiguous stimuli for assessing valence bias include nonverbal signals (faces and scenes), overlooking an inherent ambiguity in verbal signals. This study identified 32 words with dual-valence ambiguity (i.e., relatively high intersubject variability in valence ratings and relatively slow response times) and length-matched clearly valenced words (16 positive, 16 negative). Preregistered analyses demonstrated that the words-based valence bias correlated with the bias for faces, r s (213) =.27, p <.001, and scenes, r s (204) =.46, p <.001. That is, the same people who interpret ambiguous faces/scenes as positive also interpret ambiguous words as positive. These findings provide a novel tool for measuring valence bias and greater generalizability, resulting in a more robust measure of this bias.
This research examined whether the semantic relationships between representational gestures and their lexical affiliates are evaluated similarly when lexical affiliates are conveyed via speech and text. In two studies, adult native English speakers rated the similarity of the meanings of representational gesture-word pairs presented via speech and text. Gesture-word pairs in each modality consisted of gestures and words matching in meaning (semantically-congruent pairs) as well as gestures and words mismatching in meaning (semantically-incongruent pairs). The results revealed that ratings differed by semantic congruency but not language modality. These findings provide the first evidence that semantic relationships between representational gestures and their lexical affiliates are evaluated similarly regardless of language modality. Moreover, this research provides an open normed database of semantically-congruent and semantically-incongruent gesture-word pairs in both text and speech that will be useful for future research investigating gesture-language integration.
A New Bilingual Phraseological Dictionary: A Polemical ReflectionThis article presents a critical analysis of the bilingual publication entitled A Lexicon of Polish and Ukrainian Active Phraseology (Leksykon aktywnej frazeologii polskiej i ukraińskiej / Leksykon pol′s′koï ta ukraïns′koï aktyvnoï frazeolohiï), compiled by Roman Tymoshuk, Wojciech Sosnowski, Maciej Jaskot and Yurii Ganoshenko. In the history of Ukrainian-Polish and Polish-Ukrainian phraseography of the twenty-first century, this is the second attempt at creating a bilingual phraseological dictionary, following the publication of A Concise Ukrainian-Polish Dictionary of Set Expressions: Lexical Equivalents, Phraseologisms, Proverbs and Sayings (Korotkyĭ ukraïns′ko-pol′s′kyĭ slovnyk ustalenykh vyraziv: Ekvivalenty slova, frazeolohizmy, prysliv′ia ta prykazky, Poznań and Kharkiv, 2017), compiled by Tetiana Kosmeda, Olena Homeniuk and Tetiana Osipova.The distinctive feature of the reviewed dictionary is that it contains phraseologisms which are most widely used in everyday speech. The compilers developed an original conception: (1) Polish-Ukrainian and Ukrainian-Polish parts differ in content, as they were compiled independently, yet most popular phraseologisms are included in both parts; (2) the most representative set expressions in active use in both languages were selected on the basis of questionnaires and mass media material; (3) entries include illustrative material; (4) it has an optimal size – about 1,000 phraseological units.On the other hand, the dictionary also has some drawbacks, such as: (1) it lacks key criteria for determining the status of the notion “active” phraseology; (2) it does not include slang phraseologisms which do not belong to literary language; (3) the meanings of phraseological units are described by means of simple syntactic structures which lack consistent criteria of clarity and comprehensibility of interpretation, and the dictionary does not cover all semantic potential and pragmatic information of the listed units; (4) excessively categorical interpretation of the notion zero equivalence; (5) not all entries contain stylistic labels, and those used are only of three types: slang, colloquial, vulgar; the label przyslowie/приказка (proverb) seems incorrect, as the compilers do not provide criteria of its separate status; (6) metalanguage of the dictionary is marked with some violations of orthographic and stylistic norms. Nevertheless, the dictionary has undoubtedly enriched the theory of phraseography, phraseographic practice and found its users. Polemiczna refleksja na temat nowego dwujęzycznego słownika frazeologicznegoNiniejszy artykuł przedstawia analizę krytyczną dwujęzycznego Leksykonu aktywnej frazeologii polskiej i ukraińskiej (Лексикон польської та української активної фразеології), autorstwa Romana Tymoshuka, Wojciecha Sosnowskiego, Macieja Jaskota i Yuriia Ganoshenki. Jest to drugi dwujęzyczny słownik frazeologiczny w historii ukraińsko-polskiej i polsko-ukraińskiej frazeografii XXI wieku, po Małym ukraińsko-polskim słowniku utrwalonych wyrażeń językowych: Ekwiwalenty, frazeologizmy, powiedzenia i przysłowia (Короткий українсько-польський словник усталених виразів: еквіваленти слова, фразеологізми, прислів’я та приказки, Poznań–Charków 2017) Tetiany Kosmedy, Oleny Homeniuk i Tetiany Osipovej.Cechą wyróżniającą recenzowany słownik jest jego zawartość – frazeologizmy powszechnie używane ca co dzień. Autorzy zastosowali oryginalną koncepcję, w myśl której: (1) pomimo że część polsko-ukraińska i ukraińsko-polska mają różną zawartość, powstawały bowiem niezależnie, każda z nich zawiera najbardziej popularne frazeologizmy; (2) najbardziej reprezentatywne wyrażenia wybrano na podstawie badań kwestionariuszowych i materiałów ze środków masowego przekazu; (3) hasła zawierają przykłady ilustrujące ich użycie; (4) słownik ma optymalną wielkość: zawiera około tysiąca jednostek frazeologicznych.Z drugiej zaś strony, słownik ma również pewne niedociągnięcia, wśród których należy wymienić: (1) brak określenia kluczowych elementów pojęcia „aktywna” frazeologia; (2) nieuwzględnienie wyrażeń slangowych nienależących do języka literackiego; (3) opis znaczeń frazeologizmów za pomocą prostych struktur syntaktycznych, którym brak spójnych kryteriów jasności i zrozumiałości interpretacyjnej, co sprawia, że słownik nie w pełni uwzględnia potencjał semantyczny i informację pragmatyczną przedstawianych wyrażeń; (4) zbyt kategoryczna interpretacja pojęcia brak ekwiwalencji; (5) nie wszystkie hasła są opatrzone kwalifikatorami stylistycznymi, a lista tych ostatnich ogranicza się jedynie do trzech: slang, potoczne, wulgarne; kwalifikator przysłowie/приказка wydaje się niepoprawny, jako że autorzy nie podają kryteriów, na podstawie których go wyodrębniono; (6) metajęzyk słownika cechuje się pewnymi naruszeniami norm ortograficznych i stylistycznych. Niemniej jednak słownik z pewnością wzbogacił teorię frazeografii i praktykę frazeograficzną, oraz spotkał się z zainteresowaniem ze strony użytkowników.
SloIE is a manually labelled dataset of Slovene idiomatic expressions. It contains 29,400 sentences with 75 different expressions that can occur with either a literal or an idiomatic meaning, with appropriate manual annotations for each token. The idiomatic expressions were selected from the Slovene Lexical Database (http://hdl.handle.net/11356/1030). We selected only expressions that can occur with both a literal and an idiomatic meaning. The sentences were extracted from the Gigafida corpus. For each sentence, the file first contains the text of the sentence prefixed by #. This is followed by a line of numbers indicating the positions of tokens that belong to the expression. The numbers also indicate the word order for expressions where the word order is flexible. They are ordered according to the dictionary form of the expression (e.g., the first number indicates the position where the first word of the expression - in its dictionary form - occurs). Each token is labelled with either 'DA', indicating tokens in an expression that have an idiomatic meaning, 'NE', indicating tokens in an expression that have a literal meaning, or '*', indicating tokens outside the expression. Additionally, 'NEJASEN ZGLED' indicates tokens where the annotators could not determine the meaning from the example sentence. Each token is also tagged with the dictionary form of the expression that is present in the sentence. Key reference: Škvorc, Tadej, Polona Gantar, and Marko Robnik-Šikonja. "MICE: Mining Idioms with Contextual Embeddings." arXiv preprint arXiv:2008.05759 (2020).
Abstract: This paper provides guidelines and examples for visualising lexical relations using Formal Concept Analysis. Relations in lexical databases often form trees, imperfect trees or poly-hierarchies which can be embedded into concept lattices. Manyto-many relations can be represented as concept lattices where the values from one domain are used as the formal objects and the values of the other domain as formal attributes. This paper further discusses algorithms for selecting meaningful subsets of lexical databases, the representation of complex relational structures in lexical databases and the use of lattices as basemaps for other lexical relations.
<strong>Linguatec Tolosa Treebank for Occitan</strong> Linguatec Tolosa Treebank is the first dependency treebank for Occitan, developed as part of the EFA 227/16 LINGUATEC Project, financed by the POCTEFA Interreg European funds. The current version of the treebank contains 13K tokens annotated for PoS tags, lemmas and syntactic dependencies. Linguistic annotation follows Universal Dependencies guidelines (https://universaldependencies.org/#language-u). A detailed corpus description is provided in the description file. A subset of texts was doubly annotated and these annotations were adjudicated in order to provide the final annotation. These texts are therefore the most suited to be used as test files in NLP experiments. The corpus files are stored in the ConLL-U format. Each sentence is preceded by a sentence ID and the original, non-tokenized text of the sentence. The annotation is provided in a column-based format defined as follows: 1. ID: Word index, integer starting at 1 for each new sentence; may be a range for multiword tokens.<br> 2. FORM: Word form or punctuation symbol.<br> 3. LEMMA: Lemma or stem of word form.<br> 4. UPOS: Universal part-of-speech tag.<br> 5. XPOS: Language-specific part-of-speech tag; underscore if not available.<br> 6. FEATS: List of morphological features from the universal feature inventory or from a defined language-specific extension; underscore if not available.<br> 7. HEAD: Head of the current word, which is either a value of ID or zero (0).<br> 8. DEPREL: Universal dependency relation to the HEAD<br> 9. DEPS: Enhanced dependency graph in the form of a list of head-deprel pairs.<br> 10. MISC: Any other annotation. The texts are distributed under the Creative Commons BY-NC-SA 4.0 license (https://creativecommons.org/licenses/by-nc-sa/4.0/deed.en).
An exploration of physiological correlates of subjective hedonic responses while eating food has practical and theoretical significance. Previous psychophysiological studies have suggested that some physiological measures, including facial electromyography (EMG), may correspond to hedonic responses while viewing food images or drinking liquids. However, whether consuming solid food could produce such subjective-physiological concordance remains untested. To investigate this issue, we assessed participants' subjective ratings of liking, wanting, valence, and arousal while they consumed gel-type food stimuli of various flavors and textures. We additionally measured their physiological signals, including facial EMG from the corrugator supercilii. The results showed that liking, wanting, and valence ratings were negatively correlated with corrugator supercilii EMG activity. Only the liking rating maintained a negative association with corrugator supercilii activity when the other ratings were partialed out. These data suggest that the subjective hedonic experience, specifically the liking state, during food consumption can be objectively assessed using facial EMG signals and may be influenced by such somatic signals.
Research has consistently demonstrated that faces manipulated to appear more masculine are perceived as more dominant. These studies, however, have used forced-choice paradigms, in which a pair of masculinized and feminized faces was presented side by side. These studies are susceptible to demand characteristics, because participants may be able to draw the conclusion that faces which appear more masculine should be rated as more dominant. To prevent this, we tested if dominance could be perceived when masculinized or feminized faces were presented individually for only 100 ms. We predicted higher dominance ratings to masculinized faces and better memory of them in a surprise recognition memory test. In the experiment, 96 men rated the physical dominance of 40 facial photographs (masculinized = 20, feminized = 20), which were randomly drawn from a larger set of faces. This was followed by a surprise recognition memory test. Half of the participants were assigned to a condition in which the contours of the facial photographs were set to an oval to control for sexual dimorphism in face shape. Overall, men assigned higher dominance ratings to masculinized faces, suggesting that they can appraise differences in facial sexual dimorphism following very brief exposure. This effect occurred regardless of whether the outline of the face was set to an oval, suggesting that masculinized internal facial features were sufficient to affect dominance ratings. However, participants' recognition memory did not differ for masculinized and feminized faces, which could be due to a floor effect.
The study analyzes the linguistic and stylistic features of comic texts in terms of their influence on translation peculiarities of short situational utterances. The relevance of this work is preconditioned by the fact that the issue of adequate translation and understanding of humor in English is becoming more and more relevant over time due to widespread prevalence of real interaction in the transcultural space. Short humorous texts can be found in different conversational situations, but most often in situations of real communication, which, in turn, allow establishing contact and maintaining productive relationships with representatives of other linguocultures. With the reference to complex methodology, the authors analyze short humorous utterances in the virtual discourse of English messengers. Based on the analysis of a vast representative material, the authors conclude that it is necessary to take into account not only the norms and rules of the literary language, but also the non-literary features of the language (for example, “the mobile language of youth”). It means the need to translate combinations of slang and graphic signs with their specific meaning in short humorous texts. It also entails the implementation of an adequate perlocutionary effect. The peculiarity of short humorous texts translation is largely related to the phenomenon of written pronunciation, as well as imitation of the oral register of communication. The authors’ conclusions prove that in order to comply with different variants of explication of these contaminated characteristics in the target text, the use of modulation transformations, differentiation and functional replacement, as well as compensation and zero translation - complex lexical transformations, is required. Another criterion influencing the translation peculiarity of these humorous texts is the optional and obligatory explication of the comic element. The contextual analysis of the texts confirmed that in more than 30% of cases, the author generated the comic component of the text unintentionally. When conveying random humor, it is important to preserve the situation of the original and the style of speech of the author of the original message, which entails the emulation of speech features used in the framework of self-presentation strategy.
Latvian Treebank (LVTB) is in development since 2010. Corpus is annotated manually according to SemTi-Kamols hybrid dependency-constituency grammar model. This version of the treebank contains data used for deriving LVTB version 2.5.
The Universal Dependencies treebanks 1 are a still-growing collection of treebanks for a wide range of languages, all annotated with a common inventory of dependency relations.Yet, the usages of the relations can be categorically different even for treebanks of the same language.We present a pilot study on identifying such inconsistencies in a language-independent way and conduct an experiment which illustrates that a proper handling of inconsistencies can improve parsing performance by several percentage points.
The article is devoted to the visual and speech genres of Old Tiflis, which are revealed in the visual and verbal texts of famous artists N. Pirosmani, V. Elibekyan and Armenian writer A. Ayvazyan. The carnival spirit of Old Tiflis influenced the visual and verbal signs of the gastronomic infrastructure. The article is written by a semiotic, typological method, as well as an interdiscursive approach (painting, fiction). In addition, Bakhtin’s theory has become a metalanguage for identifying speech genres of urban space. An empirical analysis of the material showed that Pirosmani's paintings were a vivid example of visual advertising and, in fact, were multichannel creolized texts. Only later, when the symbolic capital of Pirosmani increased, his paint- ings began to be perceived as glowing examples of primitivism. In other words, they were “connected” to global discourse, and their primary function was for- gotten. In the main function of the signage, the paintings were the products of demand of the gastronomic infrastructure of Old Tiflis: they depicted the frames “food”, “feast”, they were also visual menus of Tiflis gastropubs (dukhans). The utilitarian demand for signs generated the brilliant paintings of Pirosmani, and at the same time, he became the author of urban space. This phenomenon is also affected in the literary discourse (A. Ayvazyan), in which two methods of modeling urban space are manifested. Firstly, the city is mod- eled in consciousness, expressed at the level of the signified (maps, sketches), and then embodied in life. That is how many cities of historical Armenia were built in diachronic point of view. Secondly, the painter (Pirosmani, Grigor – the hero of the story of Ayvazyan) painted and created in the city space (interior, exterior). The motives of his work were his own inner experiences. The multilingualism of urban space created a demand for advertising signs in different languages. Since the “social bottom” did not possess the linguistic norm of the Russian language, in the eyes of native speakers the urban texts seemed ridiculous, perceived with humor. Similar texts are found in various artists of Old Tiflis (Pirosmani, Gudiashvili, Elibekyan ect.). The same thing appears in the art discourse (Ayvazyan), which gives examples of playing with names, with language (toasts of fun and burial), etc. In addition, an analysis of such a speech genre as toast revealed that it incorporates proverbs, sayings, and in the thematic plan, the following manifestations can be called: spell-wish, wish-curse, wish-criticism, etc. The presence of a similar speech genre and the revealing of the functioning of the toast showed that eating was comprehended by different wishes, and the thematic and syntagmatic analysis of toasts can become a tool for reconstructing the axiological system of Old Tiflis society (homeland, city, parents, children, sisters and brothers, uncles and aunts, etc.).
The normative space of legal discourse does not only contain norms of various types and functions - regulatory and constitutive, formal and non-formal - but also implicates general value-based benchmarks of both social behavior of society in general and the representatives of political and state power and their "way of thinking". The article considers the correlation between the etymology of the English word, used in the text of legal discourse, and the cognitive reconstruction of the meanings of the concept «reasonable authority» on the example of Family Procedure Rules used in England and Wales. Etymological memory of a word is viewed upon in this study as a «mobile determiner» of the context, able to modify the conception of the UK judiciary. The etymology of a word should reveal the deep semantics of lexical units, uncover inferential characteristics and show the connotations of the word.
Lexical normalization is the task of translating non-standard social media data to a standard form. Previous work has shown that this is beneficial for many downstream tasks in multiple languages. However, for Italian, there is no benchmark available for lexical normalization, despite the presence of many benchmarks for other tasks involving social media data. In this paper, we discuss the creation of a lexical normalization dataset for Italian. After two rounds of annotation, a Cohen’s kappa score of 78.64 is obtained. During this process, we also analyze the inter-annotator agreement for this task, which is only rarely done on datasets for lexical normalization,and when it is reported, the analysis usually remains shallow. Furthermore, we utilize this dataset to train a lexical normalization model and show that it can be used to improve dependency parsing of social media data. All annotated data and the code to reproduce the results are available at: http://bitbucket.org/robvanderg/normit.