Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
The United States with its presidents stepping into power from either the Democratic or Republican parties influences global affairs in one way or another. These two main political parties have long been struggling for power and the significance of tapping into the ideological inclinations of the two parties underscores scholars’ accountability toward raising the critical language awareness of the public which could be an initial step toward a change for the better. The presidential inaugural speeches, due to their programmatic and strategic nature, are of significance to researchers. This study employed van Dijk’s (2006b) socio-cognitive framework where he defined two levels of analyses for a political discourse including the micro-level and macro-level text analyses. The former included 25 discursive devices such as polarization, generalization, hyperbole, etc. The latter drew on the dichotomy of ‘positive self-representation’ and ‘negative other-representation’. In the present study, the linguistic features in 16 inaugural speeches delivered by American Democratic and Republican presidents from 1961 to 2017 were examined at both levels. The overall data analysis revealed that Democrats employed ‘norm expression’ and ‘presupposition’ significantly more than Republicans, while Republicans made more use of ‘categorization’, ‘lexicalization’, and ‘populism’. The macro-level comparison of the two parties indicated that both Democrats and Republicans resorted to using ‘positive self-representation’ significantly more than ‘negative other-representation’ while the deployment of ‘negative other-representation’ by Republicans was significantly more than that by Democrats. The findings of this study have some implications for English for political purposes, political studies, as well as attempts in discourse studies.
In this welcome study of the second-century Apocryphal Acts of Paul and Thecla (ATh for Acts of Thecla), J. D. McLarty is particularly interested in questions of emotion and identity. The exploration of themes such as gender, class, and citizenship is driven by narrative analysis and comparison with one of the five complete Greek novels, Chariton's Callirhoe. Revised from McLarty's 2011 PhD thesis, Thecla's Devotion leads the reader carefully through both ATh and Callirhoe, showing their striking emotive similarities and differences. McLarty argues that, despite being a work of Christian fiction, ATh provides a helpful example of how early ‘Christians in one part of the eastern Empire constructed an identity for themselves’ (p. 232). In Chapter 1, McLarty examines the context of ATh, discussing questions of composition, origin, and date. She accepts the view that the Thecla episode was originally included in the larger narrative of the Acts of Paul but was later disseminated independently with the development of the cult of St. Thecla (already present in the fourth century). It is likely that the text of ATh originated in ‘south central Asia Minor’ (p. 5) and is dated to the ‘mid-to-late second century’ (p. 7). Concerning whether the author of the text was male or female, McLarty concludes, ‘even if one wishes to argue for at least some female contribution to the narrative in the form of oral legend, these contributions have been absorbed into a masculine literary culture’ (p. 9). However, the readers and hearers of the text were likely mixed in both gender and class. Rather than a primarily oral development, McLarty sees the composition as reflecting a ‘predominantly literary milieu’ (p. 17), which both imitated and reworked traditions from the Acts of the Apostles. The world of ATh was a changing one, with a growing interest in the ethics of the individual and self-control. This focus is clearest in relation to discussions of marriage and continence contemporary to ATh. McLarty is, therefore, interested in the intersection of pagan ideas of self-control, reflected in Callirhoe, with a Christian view that contains a distinct ‘spiritual dimension to the control of the passions’ (p. 23). Story, in contrast to pure discourse, provides a unique opportunity to affect the emotions of the reader, but may also provide insight into the author's worldview. This study is broken into two parts. The first part analyses the plot of ATh in comparison to Callirhoe. In Chapter 2, McLarty outlines her methodology, focused especially on the concept of ‘ “affect” – the emotive atmosphere created by the construction of plot’ (p. 28). The next chapter provides a diachronic study through an extended discussion of the plots of Callirhoe (Book 1) and ATh. With both narratives, the points of interest are (1) the teleology of the plot, (2) the loci of tension, (3) and causation. While Callirhoe anticipates a safe return home, ATh ends with all of Thecla's ties to family and home broken, leaving her in ‘a plot space that is without definition’ (p. 89) as she goes about evangelizing. Plot tension presents a challenge to the reader's assumption that all will end well for the protagonists. This tension also provides Thecla with an opportunity to mature through perseverance, a process that culminates in self-baptism. While the concept of ‘Chance’ (τύχη) – sometimes personified as a deity – plays an important role in the causation of Callirhoe, it does not appear in ATh. Instead, God appears in the narrative through deus ex machina (in the theater of Iconium), in order to save the heroine. Human causation in ATh is largely confined to the lack of emotional control, with Paul being the main exception. Chapters 4 and 5 are devoted to a synchronic study, consisting of an analysis of time (Chapter 4), space, and place (Chapter 5). While the author of Callirhoe uses many technical literary features for emotive affect, ATh is often simplified and provides additional moral exhortation. McLarty discusses how the idea of place is subverted in ATh: Tombs and prisons become places of Christian devotion, and the stadium is used for baptism, rather than death. While McLarty claims that ‘Thecla is not portrayed as overturning male authority’ (p. 119), she does transgress boundaries of gender and class. Thecla is found denying her betrothed and, therefore, the civic value of marriage; she roams the streets unaccompanied and ultimately develops into a wandering ‘ “type” of Christian philosopher’ (p. 124). The second part of this study examines character and characterization as the locus of expressed emotion in the narratives. From here, ‘one can therefore glean useful information about the role of emotion in the ATh author's own community’ (p. 132). McLarty proposes a methodology that combines ancient and modern theories of character and compares ATh with Callirhoe as its ‘genre model’ (though she is clear that it is not the only model). Chapters 7 and 8 explore characterization in Callirhoe (Chapter 7) and ATh (Chapter 8), each divided by female and male characters in the narratives. Characters are described in terms of social class and emotional expression, which reveals that, in these narratives, emotion (produced from desire) is reserved mostly for the upper class. Callirhoe also shows the ‘potential destructiveness of uncontrolled masculine emotions’ (p. 166). At times the heroine is depicted as more masculine or ‘warlike’ (p. 167), while the hero is feminized. However, the conclusion of the narrative provides a return home, both physically and emotionally, with both characters fulfilling their culturally defined gender and societal roles. ATh, on the other hand, contrasts and subverts many of these expectations. Although Thecla is understood to be overtaken by passion for Paul, not unlike to the heroine of Callirhoe, it becomes clear to the reader that her true desire is for the gospel he preaches. McLarty concludes that Thecla's character remains that of the ‘perpetual παρθένος [unmarried girl]’ (p. 219). In spite of this, the narrative ends with Thecla reaching maturity through baptism and possessing ‘a masculine control of her emotions’ (p. 219). Again, the contrast with Callirhoe is apparent when the reader understands that Thecla does not return home but remains independent and isolated. This combination of celibacy and isolation subverts many of the second century assumptions about gender and class – namely, that young women (especially of high status) should marry and let their husbands protect them from falling victim to eros (desire). While Paul does not subvert gender norms as such, he does challenge many social categories. ATh portrays Paul as an opponent to the leading men of the city, who are fighting for Thecla's affection. However, ‘Paul does not take part in this contest, an indication that the honour system of the Christian community is different from that of the wider pagan society’ (p. 220). McLarty's final chapter provides a useful summary of emotion and identity in ATh, while also returning to some motifs that she hints at in the beginning of the study. In particular, she discusses the Christian narrative's interaction with pagan philosophy (especially Stoic and Cynic). ATh presents Paul as ‘like a philosopher in his mastery of the passions’ and as a wise man ‘countering the kinds of argument [against Christians] advanced by Celsus’ (p. 227). In contrast to these philosophies, the ‘mastery of the passions’ achieved by both Paul and Thecla comes ‘through Christ’ (p. 232). There is much to admire in McLarty's approach to ATh. Her detailed lexical analysis of both Callirhoe (Book 1) and ATh is sure to provide a useful guide for readers. McLarty also draws from a broad range of classical and Christian texts as comparanda to this apocryphal book. Although she makes a strong case for the predominantly Graeco–Roman background of ATh, one might wonder whether comparison with a Jewish novel like Joseph and Aseneth could further shed light on character and emotion. It is not clear that the bibliography has been fully updated from her 2011 thesis. For example, one will not find interaction with Susan Hylan's A Modest Apostle: Thecla and the History of Women in the Early Church (Oxford: OUP, 2015), who argues that Thecla does not depart from modesty, contrary to what McLarty claims (p. 170). Finally, it is possible that McLarty's emphasis on Thecla's isolation overshadows certain hints at her accumulation of followers. For this reason, McLarty does not mention the ‘band of young men and maidens’ (ATh p. 40), who accompany Thecla in her final return to Paul. This does not negate the overall point about Thecla's subversive character but may indicate some desire in the second century to follow the heroine's example. In the end, McLarty's captivating prose and persuasive arguments ensure that this work is an important contribution to the study of ATh.
Este trabalho objetiva analisar o vocabulario documentado pelo Projeto Atlas Linguistico do Brasil (ALiB), especificamente as ocorrencias em Mato Grosso do Sul das variantes lexicais que nomeiam o referente geralmente conhecido como boteco – QSL 202/ALiB (COMITE NACIONAL DO PROJETO ALiB, 2002, p. 37), area semântica Vida urbana, enunciadas por 28 informantes entrevistados pelos pesquisadores do Atlas Linguistico do Brasil (ALiB), habitantes de seis pontos de inquerito sul-mato-grossenses. Tratando-se de pesquisa descritiva com base em dados empiricos, o estudo tem como base teorica a Dialetologia e a Geolinguistica (FERREIRA; CARDOSO, 1994; ISQUERDO, 2003), bem como principios da Sociolinguistica (CAMACHO, 2001). Por meio da analise do corpus, os resultados finais obtidos mostraram o registro de 7 variantes lexicais para nomear o boteco, sendo que a de maior incidencia foi de boteco/botequinho/botiquim, com 44% das ocorrencias, seguida por bar, 30%, e na sequencia vem bolicho, com 10% dos registros. A maior variacao lexical ocorreu na capital, Campo Grande, com sete variantes, tambem unico ambiente em que se observou o item lexical emporio. Nota-se tambem que boteco/botequinho nao aparece como o de maior incidencia na capital e em Ponta Pora, locais em que se destaca bar (38% e 50% de produtividade, respectivamente). Ja uma menor variacao encontra-se na cidade de Paranaiba – com apenas duas unidades lexicas registradas – boteco/botequinho/botiquim e bar. Diante do exposto, pode-se concluir que, com os resultados obtidos e analisados, o estudo apresentado pode contribuir para as pesquisas dialetais, pois, a partir de dados geolinguisticos, analisou os nomes para boteco em Mato Grosso do Sul, ratificando a importância dos estudos lexicais para o conhecimento da norma linguistica de uma sociedade e sua intrinseca relacao com a historia e a cultura. ABSTRACT: This work aims to analyze the vocabulary documented by the Atlas Linguistic of Brazil Project (ALiB), specifically the occurrences in Mato Grosso do Sul of the lexical variants that name the referent generally known as boteco (bar) - QSL 202 / ALiB (NATIONAL COMMITTEE OF PROJECT ALiB, 2002, p. 37), semantic area Urban life, enunciated by 28 informants interviewed by researchers from the Linguistic Atlas of Brazil (ALiB), inhabitants of six survey points in that state. In the case of descriptive research based on empirical data, the study is theoretically based on Dialectology and Geolinguistics (FERREIRA; CARDOSO, 1994; ISQUERDO, 2003), as well as principles of Sociolinguistics (CAMACHO, 2001). Through the analysis of the corpus, the final results obtained showed the registration of 7 lexical variants to name the bar, with the highest incidence being boteco/botequinho/botiquim, with 44% of occurrences, followed by bar, 30%, and in the sequence comes bolicho, with 10% of the records. The largest lexical variation occurred in the capital, Campo Grande, with seven variants, also the only environment in which the lexical item emporio was observed. It should also be noted that boteco/botequinho does not appear as the one with the highest incidence in the capital and in Ponta Pora, places where bar stands out (38% and 50% of productivity, respectively). A smaller variation is found in the city of Paranaiba - with only two lexical units registered - boteco/botequinho/botiquim and bar. Given the above, it can be concluded that, with the results obtained and analyzed, the study presented can contribute to dialect research, since, using geolinguistic data, it analyzed the names for bar in Mato Grosso do Sul, ratifying the importance of lexical studies for the knowledge of the linguistic norm of a society and its intrinsic relationship with history and culture. KEYWORDS: Lexical norm; Mato Grosso do Sul; Atlas Linguistic of Brazil Project; bar.
With advancements in information and communication technology, there is a growing importance of natural language processing (NLP) in the development of language processing systems, such as machine translation, speech recognition, and linguistic database construction. Thus, this study explored methods to increase the accuracy of the NLP of Korean-Chinese corresponding grammatical structures, focused on the Korean word “있다”. The Korean-Chinese corresponding structures were analyzed with focus on the Korean word “있다” and its corresponding Chinese characters: “有,” “在” and “着”. This study translated related sentences using Google Translate, analyzed error types and organized grammatical structures. After that, erroneous sentences related to “있다” were translated using Google Translate, and the translated sentences were analyzed. Based on the findings, this study suggested methods to correct and adjust errors. The improvement methods drawn from the analysis of translated sentences can be summarized as follows. Improvement method 1: Provide a rule about spacing words in Korean Improvement method 2: Provide a rule about corresponding words e.g.) 서울=首尔,오빠=哥哥 Improvement method 3: Provide a rule about Korean-Chinese corresponding grammatical structures e.g.) Chinese: numeral + classifier + noun, Korean: noun + numeral + classifier It is hoped that this study will help establish a computerized system for Korean and Chinese.
Modern integration processes contribute to the formation a digital space and reinforce the necessity for communication between citizens of different countries in foreign languages. The number of people who are speaking more than one foreign language is growing. Currently, more than half of EU citizens are free to communicate in at least one foreign language. Therefore, in recent years, scientists have increased attention to solving objective problems related to the study of foreign languages, including the problem of linguistic interference. The purpose of the study is to determine the content and basic properties of "linguistic interference", to elucidate the features of its manifestation in increasing the level of bilingualism in the language environment. The author has considered the problems that arise in cross-cultural communication in the language of bilinguals as a result of the interaction and interpenetration of different language systems, disclosed the contents and the basic properties of the linguistic interference concept based on generalization and analysis of scientific approaches. Generally, the interference means the transfer of native language norms to another in the process of learning it. By other scientists’ definition, interference is a deviation from interacting languages norms including merging, mixing, contact, mutual influence, etc. Such a transfer of the one language norms to another can be done orally, in the process of speech, during translation, writing, etc. The interference occurrence problem has not only a linguistic basis, but also a psychological one, as linguistic communication is not only the exchange of speech structures, but also the psychological perception of speech patterns in mind. In this context, the features of manifestation of linguistic interference features in language environment bilingualism level conditions increasing have been clarified. The author also identified the linguistic interference types and the main reasons that cause it. In addition, cultural aspects of communication become important when the linguistic units understanding can be interpreted differently from the cultural context. It is determined that, on the one hand, interference leads to a linguistic norms violation due by mixing similar language units, and on the other hand, it acts as a specific factor in the development of speech.
E-Shops reviews have become a valuable source of opinions, feeling, and experience by existing users which marks the failure and success of products and services. A set of favorable reviews proves that the product or service has vended well and users are pleased with them, and opposing reviews effects adversely. These reviews provide worthwhile information that might be used by prospective users to find sentiments of previously preferred users before making any decision or transaction to buy right product or service from E-shops. Moreover, service providers, vendors or manufacturers also make use of these reviews to discover public opinion, as well as limitations of their products or services. But, the basic fact about underlying online reviews shows something else. These reviews might be spurious, posted or written with hidden purposes, they often involve helpful or favorable opinion in order to boost, promote, and publicize their products and services, or with pessimistic intention to harm opponent prestige and business as well. So, the exploration and investigation of such reviews before opinion mining is significant. In this paper, a methodology is proposed which includes reviews acquisition using the Tag path clustering approach about mobile devices of different make along with metadata from Flipkart, spurious reviews detection based on identical and nearly identical reviews using semantic similarity with review length. Further, an opinion mining process is carried out by using the lexical database (SentiWordNet) approach for the computation of sentimental degree and orientation of both spurious and legitimate reviews.
Social animals show reduced physiological responses to aversive events if a conspecific is physically present. Although humans are innately social, it is unclear whether the mere physical presence of another person is sufficient to reduce human autonomic responses to aversive events. In our study, participants experienced aversive and neutral sounds alone (alone treatment) or with an unknown person that was physically present without providing active support. The present person was a member of the participants' ethnical group (ingroup treatment) or a different ethnical group (outgroup treatment), inspired by studies that have found an impact of similarity on social modulation effects. We measured skin conductance responses (SCRs) and collected subjective similarity and affect ratings. The mere presence of an ingroup or outgroup person significantly reduced SCRs to the aversive sounds compared to the alone condition, in particular in participants with high situational anxiety. Moreover, the effect was stronger if participants perceived the ingroup or outgroup person as dissimilar to themselves. Our results indicate that the mere presence of another person was sufficient to diminish autonomic responses to aversive events in humans, and thus verify the translational validity of basic social modulation effects across different species.
Context: Parkinson’s disease (PD) is a neurodegenerative disease caused by degeneration of the dopaminesynthesizing cells of the mesostriatal-mesocortical neuronal pathway,which affects motor pathway in basal ganglia (BG). Neuropsychological studies showed that degeneration of dopamine neuroreceptor also affects nigrostriatal and mesocortical limbic system which is associated with emotional processing in PD. However, very few studies have identified deficit in selective attention in patients with PD patients except in patients with PD-MCI (PD-Mild Cognitive Impairment) or PD-D (PD-Dementia). Thus, the present study examined the effect of emotion on attentional processing in PD and matched control. Emotional flanker task was designed by using pictures selected from the International Affective Picture System (IAPS) based on their normative valence ratings. Results revealed that attentional processing of emotional images were slower in PD patients in comparison to matched healthy control.
This is an introduction to the proposed theme, in which the importance of sociolinguistic studies for the teaching, acquisition and learning of languages is emphasized. In addition, each text of the material is presented, starting with interviews with significant and current representatives of the variation sociolinguistics (Francisco Moreno Fernández and Juan Manuel Hernández Campoy) from the Hispanic and Anglo-Saxon spheres, respectively; then, it discusses the ten articles that deal with the theme from two perspectives: linguistic attitudes and beliefs of speakers and linguistic norms and policies. Finally, the reviews of two books related to the Special issue are commented: The Routledge handbook of Spanish as a heritage language, edited by Kim Potowsky, 2018, New York, Routledge publisher, and La trastienda de la enseñanza de lenguas extranjeras, by Francisco García Marcos, 2018, from the Interlingua collection of Editora Comares de Granada / Spain. The presentation is an invitation to readers to enjoy reading the Special issue.
We study the effect of rich supertag features in greedy transition-based dependency parsing. While previous studies have shown that sparse boolean features representing the 1-best supertag of a word can improve parsing accuracy, we show that we can get further improvements by adding a continuous vector representation of the entire supertag distribution for a word. In this way, we achieve the best results for greedy transition-based parsing with supertag features with $88.6\%$ LAS and $90.9\%$ UASon the English Penn Treebank converted to Stanford Dependencies.
Because of its focus on the past and on historical languages, the classics is a discipline that is particularly interested in translations and text alignment. Starting from a diachronic perspective, this contribution demonstrates how issues related to text alignment, present since antiquity, can be approached from a different angle and with entirely new opportunities thank to tools and methods developed in the field of digital humanities. By comparing examples from antiquity (e.g. Origen’s Hexapla from the third century CE) with modern projects based on treebanking and dependency grammar (e.g. the Ancient Greek and Latin Dependency Treebank [AGLDT] as part of the Perseus Digital Library from Tufts University), we shall present some new approaches and their potentials. In doing so, we shall also examine what status English has in these projects and how the different languages involved in each of them interact with English and/or with each other.
The goal of this special issue of Critical Multilingualism Studies “National Standards – Local Varieties: A Cross-Linguistic Discussion on Regional Variation in L2 Studies” is to incite a conversation on how topics such as linguistic norms and variation, dominant practices, ideologies, identities, and politics surrounding languages are discussed from a view outside of the dominant centers of linguistic norms.
Language users and learners are sensitive to distributional information in their environment, which enables them to extract regularities that occur in the language input that they are exposed to. This process is referred to as statistical learning. While the statistical learning phonotactic literature thoroughly investigates the learning of overall phonotactics in specific languages, little is known about cases where different phonological systems coexist within a single language. The Japanese lexicon is generally classified into four lexical strata according to the etymological status of each word (Itô & Mester, 1995, 1999, 2001). Although each stratum includes the internal phonological similarity in the Japanese language as a whole, there are also distinctive phonological properties. A recent study suggests that language users should be able to learn phonotactics of each sublexicon based on the same kind of statistical probabilities that computers analyse from language users’ accumulated lexicons (Morita, 2018). This thesis examines whether second-language (L2) learners can learn the loanword phonotactics/phonology of Japanese through experience of using and/or passive exposure to Japanese lexical stratification. Using two loanword phonological regularities (categorical and gradient rules) as a case study, two fully-crossed perceptual experiments involving English- speaking learners of Japanese, native speakers of Japanese, and English-speaking monolinguals are presented. The first experiment explores listeners’ phonotactic/phonological knowledge of nativised loanwords in Japanese using a well-formedness task which shows the adaptation of English final consonants in monosyllabic words. Listeners judge whether the pronunciation they hear is how the word would be pronounced if it was a Japanese word, rating how confident they are on a scale of 1-5. This study shows that L2 learners learn categorical rules, but not gradient patterns. This study also confirms that loanword phonotactics and overall phonotactics make separate contributions to perceived well-formedness. L2 learners access and make use of the sublexicon-specific probabilities of Japanese during the task. The second perceptual experiment is designed to support the findings in the first experiment, by testing for discrimination of non-native consonantal contrasts. Even under high memory demand, L2 learners show the ability to discriminate non-native consonantal contrasts (i.e., CVCV/CVCCV) effectively enough to support findings in the first experiment. These results suggest that L2 learners can implicitly detect the statistical structure of a language’s sublexicon phonology over the course of acquiring a natural language. However, while native speakers of Japanese learn a gradient rule, L2 learners of Japanese do not. A potential explanation for the differences in gradient rule learning is that the vocabulary size of the target language might play a crucial role. This remains an open question. In addition, the present work provides a basis for future investigation into whether L2 learners of Japanese, whose native language is other than English, are able to learn Japanese loanword phonotactics/phonology. L1 English-L2 Japanese speakers might gain advantage in perceiving the English input which inevitably overlaps with the phonological form of the host language.
Introduction. High-quality language education in technical universities requires its interdisciplinary relation to the content of highly specialised subjects corresponding to the training programmes aimed at instructing the future specialists. Educational materials in a foreign language are highly productive if they emphasise the terminology and professional vocabulary authentic to the current state of the scientific field. The aim of the study presented in the article was to assess the validity of the lexical material delivered in the course “English for Business Communication”, to determine the selection criteria for this vocabulary as well as the methods for its assimilation and practical application. Methodology and research methods. The applied corpus software enabled to obtain quantitative indicators of the distribution of foreign-language business vocabulary in the given training course. The lexical material being currently offered to students and the professional thesaurus identified via linguistic databases was compared with the use of comparative analysis and synthesis. Results and scientific novelty. The lexical units (terms, set expressions), which are the most active in the business sphere, were identified on the basis of its frequency. The authors established the correlation between them and educational vocabulary, both from the perspective of its integration into the course without block concentration throughout the course of university training, and from the perspective of the variety of methods used to practice this vocabulary. It is concluded that the applied educational material needs to be substantially adjusted. The vocabulary does not completely reflect the realities of the business communication sphere and the distribution of active vocational vocabulary regulated by methodological guidelines does not entirely contribute to its strong assimilation. According to the authors, the necessary changes to the approaches and methods for selecting and compiling lexical material and to the methodology for designing a foreign language course should be made on the basis of integrating pedagogical and linguistic knowledge, in particular, the methodology of teaching foreign languages and the corpus linguistics. Practical significance. The ways of integrating corpus programs in the process of developing the content of language disciplines, which are part of the main educational program of technical universities, are demonstrated as one of the methods to increase the effectiveness of teaching foreign languages to students of non-linguistic specialties.
Semantic Role Labelling (SRL) is the process of automatically finding the semantic roles of terms in a sentence. It is an essential task towards creating a machine-meaningful representation of textual information. One public linguistic resource commonly used for this task is the FrameNet Project. FrameNet is a human and machine-readable lexical database containing a considerable number of annotated sentences, those annotations link sentence fragments to semantic frames. However, while the annotations across all the documents covered in the dataset link to most of the frames, a large group of frames lack annotations in the documents pointing to them. In this paper, we present a data augmentation method for FrameNet documents that increases by over 13% the total number of annotations. Our approach relies on lexical, syntactic, and semantic aspects of the sentences to provide additional annotations. We evaluate the proposed augmentation method by comparing the performance of a state-of-the-art semantic-role-labelling system, trained using a dataset with and without augmentation.
In order to extract the semantic and grammatical information of sentences more effectively, this paper proposes a sentence sentiment classification method based on Self-supervised and Self-attention mechanism (SS-SAtt-BiLSTM). In this method, BiLSTM network is used to extract the feature of text context relationship, and self-supervised (SS) learning mode is introduced into the supervised sentence representation model. The sentence itself is used as the label data information of current words, and an improved self-attention mechanism (SA) is used to calculate the attention weight of each moment. The experimental results of MR and Stanford sentient treebank (sst-5) data sets show that this method reduces the dependence on tagged data, and the improved self-attention mechanism enables the model to learn more key features of sentences and improve the classification performance.
Noun phrases convey key information in communication and are of interest in NLP tasks. A base NP is defined as the headword and left-hand side modifiers of a noun phrase. In this thesis, we identify base NPs in Universal Dependencies treebanks in English and French using an RNN architecture.The data of this thesis consist of three multi-layered treebanks in which each sentence is annotated in both constituency and dependency formalisms. To build our training data, we find base NPs in the constituency layers and project them onto the dependency layer by labeling corresponding tokens. For input features, we devised 18 configurations of features available in UD annotation. We train RNN models with LSTM and GRU cells with different numbers of epochs on these configurations of features.Tested on monolingual and bilingual test sets, our models delivered satisfactory token-based F1 scores (92.70% on English, 94.87% on French, 94.29% on bilingual test set). The most predicative configuration of features is found out to be pos_dep_parent_child_morph, which covers 1) dependency relations between the current token, its syntactic head, its leftmost and rightmost syntactic dependents; 2) PoS tags of these tokens; and 3) morphological features of the current token.
As the number, size, and complexity of building construction projects increase, code compliance checking becomes more challenging because of the time-consuming, costly, and error-prone nature of a manual checking process. A fully automated code compliance checking would be desirable in facilitating a more efficient, cost effective, and human error-proof code checking. Such automation requires automated information extraction from building designs and building codes, and automated information transformation to a format that allows automated reasoning. Natural language processing (NLP) is an important technology to support such automated processing of building codes, because building codes are represented in natural language texts. Part-of-speech (POS) tagging, as an important basis of NLP tasks, must have a high performance to ensure the quality of the automated processing of building codes in such a compliance checking system. However, no systematic testing of existing POS taggers on domain specific building codes data have been performed. To address this gap, the authors analyzed the performance of seven state-of-the-at POS taggers on tagging building codes and compared their results to a manually-labeled gold standard. The authors aim to: (1) find the best performing tagger in terms of accuracy, and (2) identify common sources of errors. In providing the POS tags, the authors used the Penn Treebank tagset, which is a widely used tagset with a proper balance between conciseness and information richness. An average accuracy of 88.80% was found on the testing data. The Standford coreNLP tagger outperformed the other taggers in the experiment. Common sources of errors were identified to be: (1) word ambiguity, (2) rare words, and (3) unique meaning of common English words in the construction context. The found result of machine taggers on building codes calls for performance improvement, such as error-fixing transformational rules and machine taggers that are trained on building codes.
Cupping therapy has recently gained public attention and is widely used in many regions. Some patients are resistant to being treated with cupping therapy, as visually unpleasant marks on the skin may elicit negative reactions. This study aimed to identify the cognitive and emotional components of cupping therapy. Twenty-five healthy volunteers were presented with emotionally evocative visual stimuli representing fear, disgust, happiness, neutral emotion, and cupping, along with control images. Participants evaluated the valence and arousal level of each stimulus. Before the experiment, they completed the Fear of Pain Questionnaire-III. In two-dimensional affective space, emotional arousal increases as hedonic valence ratings become increasingly pleasant or unpleasant. Cupping therapy images were more unpleasant and more arousing than the control images. Cluster analysis showed that the response to cupping therapy images had emotional characteristics similar to those for fear images. Individuals with a greater fear of pain rated cupping therapy images as more unpleasant and more arousing. Psychophysical analysis showed that individuals experienced unpleasant and aroused emotional states in response to the cupping therapy images. Our findings suggest that cupping therapy might be associated with unpleasant-defensive motivation and motivational activation. Determining the emotional components of cupping therapy would help clinicians and researchers to understand the intrinsic effects of cupping therapy.
The Italian Sign Language (LIS) is the natural language used by the Italian Deaf community. This paper discusses the application of the Universal Dependencies (UD) format to the syntactic annotation of a LIS corpus. This investigation aims in particular at contributing to sign language research by addressing the challenges that the visual-manual modality of LIS creates generally in linguistic annotation and specifically in segmentation and syntactic analysis. We addressed two case studies from the storytelling domain first segmented on the ELAN platform, and second syntactically annotated using CoNLL-U format.
Delivery of best-practice care for posttraumatic stress disorder (PTSD) is a priority for clinicians working with active duty military personnel and veterans. The PTSD Clinicians Exchange, an Internet-based intervention, was designed to assist in disseminating clinically relevant information and resources that support delivery of key practices endorsed in the Veterans Administration (VA)-Department of Defense (DoD) Clinical Practice Guidelines (CPG) for the Management of Posttraumatic Stress. We conducted a randomized controlled trial to examine the effectiveness of the Clinicians Exchange intervention in increasing familiarity and perceived benefits of 26 CPG-related and emerging practices. The intervention consisted of ongoing access to an Internet resource featuring best-in-class resources for practices, self-management of burnout, and biweekly e-mail reminders highlighting selected practices. Mental health clinicians (N = 605) were recruited from three service sectors (VA, DoD, community); 32.7% of participants assigned to the Internet intervention accessed the site to view resources. Individuals who were offered the intervention increased their practice familiarity ratings significantly more than those assigned to a newsletter-only control condition, d = 0.27, p =.005. From baseline to 12-months, mean familiarity ratings of clinicians in the intervention group increased from 3.0 to 3.4 on scale of 1 (not at all) to 5 (extremely); mean ratings for the control group were 3.2 at both assessments. Clinicians generally viewed the CPG practices favorably, rating them as likely to benefit their clients. The results suggest that Internet-based resources may aid more comprehensive efforts to disseminate CPGs, but increasing clinician engagement will be important.
Dual language immersion (DLI) programs have emerged in the U.S. as effective ways to bring together language minority and language majority speakers in school settings with the goal of bilingualism and bi-literacy for all. However, the proliferation of these programs has raised concerns regarding issues of inequity and dissimilar power dynamics in these spaces (Cervantes-Soon, 2014 Cervantes-Soon, C. G. 2014. “A Critical Look at Dual Language Immersion in the New Latin@ Diaspora.” Bilingual Research Journal 37 (1): 64–82.[Taylor & Francis Online], [Google Scholar], “A Critical Look at Dual Language Immersion in the New Latin@ Diaspora.” Bilingual Research Journal 37 (1): 64–82; Flores, 2016, Do Black Lives matter in Bilingual Education [Web log post]. Accessed May 1, 2017. https://educationallinguist.wordpress.com/2016/09/11/do-black-lives-matter-in-bilingual-education/; Valdes, 1997, “Dual language immersion programs: A cautionary note concerning the education of language-minority students.” Harvard Educational Review 67: 391–430, 2018, “Analyzing the curricularization of language in two-way immersion education: Restating two cautionary notes.” Bilingual Research Journal). With this in mind, this study aims to shed light on the intricate social processes at work in DLI contexts. In particular, this paper examines first, how notions of language use, race, and ethnicity are socially constructed and intersect in DLI settings; and second, it explores how these ideas are discerned and re-shaped by young children into their own social and linguistic norms. Employing qualitative research methods, this year-long ethnographic case study uses the intersectional lens of raciolinguistics (Alim, Rickford & Ball, 2016 Alim, H. S., J. R. Rickford, and A. F. Ball, eds. 2016. Raciolinguistics: how Language Shapes our Ideas About Race. New York, NY: Oxford University Press.[Crossref], [Google Scholar], Raciolinguistics: how language shapes our ideas about race. New York, NY: Oxford University Press; Rosa & Flores, 2017, “Unsettling race and language: Toward a raciolinguistic perspective.” Language in Society 46 (5): 621–647), to examine the intricate cross-cutting dynamics at play in bilingual spaces. The exploration of these ideas helps to illuminate the ways in which language practices and interactions are shaped by social constructions from a very early age. Furthermore, it contributes to understandings of social perceptions and relations in multilingual/multicultural/multiethnic contemporary school settings.
Arabic diacritics play a significant role in distinguishing words with the same orthography but different meanings, pronunciations, and syntactic functions. The presence of Arabic diacritics can be useful in many natural language processing applications, such as text-to-speech tasks, machine translation, and part-of-speech tagging. This article discusses the use of bidirectional long short-term memory neural networks with conditional random fields for Arabic diacritization. This approach requires no morphological analyzers, dictionary, or feature engineering, but rather uses a sequence-to-sequence schema. The input is a sequence of characters that constitute the sentence, and the output consists of the corresponding diacritic(s) for each character in that sentence. The performance of the proposed approach was examined using four datasets with different sizes and genres, namely, the King Abdulaziz City for Science and Technology text-to-speech (KACST TTS) dataset, the Holy Quran, Sahih Al-Bukhary, and the Penn Arabic Treebank (ATB). For training, 60% of the sentences were randomly selected from each dataset, 20% were selected for validation, and 20% were selected for testing. The trained models achieved diacritic error rates of 3.41%, 1.34%, 1.57%, and 2.13% and word error rates of 14.46%, 4.92%, 5.65%, and 8.43% on the KACST TTS, Holy Quran, Sahih Al-Bukhary, and ATB datasets, respectively. Comparison of the proposed method with those used in other studies and existing systems revealed that its results are comparable to or better than those of the state-of-the-art methods.
本研究依据以谓词为核心的块依存语法构建块依存树库,在句内和句间寻找谓词所支配的组块,利用汉语中组块和组块间的依存关系补全缺省部分,明确谓词支配关系。目前共标注2199篇文本,涵盖百科、新闻两个领域,共约187万字语料。本文简述了块依存语法的原则,并对组块及其依存关系进行了定义。将详细介绍标注流程、标注一致率、数据分布等情况。基于现有的树库,本研究发现汉语中有约25%的小句是非自足的,约有88%的核心谓词可支配1~3个从属成分。
Even though Automatic Speech Recognition (ASR) systems significantly improved over the last decade, they still introduce a lot of errors when they transcribe voice to text. One of the most common reasons for these errors is phonetic confusion between similar-sounding expressions. As a result, ASR transcriptions often contain "quasi-oronyms", i.e., words or phrases that sound similar to the source ones, but that have completely different semantics (e.g., "win" instead of "when" or "accessible on defecting" instead of "accessible and affecting"). These errors significantly affect the performance of downstream Natural Language Understanding (NLU) models (e.g., intent classification, slot filling, etc.) and impair user experience. To make NLU models more robust to such errors, we propose novel phonetic-aware text representations. Specifically, we represent ASR transcriptions at the phoneme level, aiming to capture pronunciation similarities, which are typically neglected in word-level representations (e.g., word embeddings). To train and evaluate our phoneme representations, we generate noisy ASR transcriptions of four existing datasets - Stanford Sentiment Treebank, SQuAD, TREC Question Classification and Subjectivity Analysis - and show that common neural network architectures exploiting the proposed phoneme representations can effectively handle noisy transcriptions and significantly outperform state-of-the-art baselines. Finally, we confirm these results by testing our models on real utterances spoken to the Alexa virtual assistant.
We propose a transition-based approach that, by training a single model, can efficiently parse any input sentence with both constituent and dependency trees, supporting both continuous/projective and discontinuous/non-projective syntactic structures. To that end, we develop a Pointer Network architecture with two separate task-specific decoders and a common encoder, and follow a multitask learning strategy to jointly train them. The resulting quadratic system, not only becomes the first parser that can jointly produce both unrestricted constituent and dependency trees from a single model, but also proves that both syntactic formalisms can benefit from each other during training, achieving state-of-the-art accuracies in several widely-used benchmarks such as the continuous English and Chinese Penn Treebanks, as well as the discontinuous German NEGRA and TIGER datasets.
In this paper, we propose MCNN-ReMGU model based on multi-window convolution and residual-connected minimal gated unit (MGU) network for the natural language word prediction. First, the convolution kernels with different sizes are used to extract the local feature information of different graininess between the word sequences. Then, the extracted features are fed to the residual-connected MGU network. Finally, the prediction results are output by the SoftMax layer. Through the residual-connection processing of MGU network in the model, not only the problems of vanishing gradient and network degradation are effectively solved, but also the long-term dependence between word sequences is effectively extracted to predict the next word accurately. Meanwhile, the introduction of the convolution kernel in a convolutional neural network (CNN) enables the feature information between word sequences to be extracted more fully. The experimental results on the Penn Treebank and WikiText-2 datasets show that the proposed method has certain advantages in the word prediction task.
Recommender systems often involve multi-aspect factors. For example, when shopping for shoes online, consumers usually look through their images, ratings, and product's reviews before making their decisions. To learn multi-aspect factors, many context-aware models have been developed based on tensor factorizations. However, existing models assume multilinear structures in the tensor data, thus failing to capture nonlinear feature interactions. To fill this gap, we propose a novel nonlinear tensor machine, which combines deep neural networks and tensor algebra to capture nonlinear interactions among multi-aspect factors. We further consider adversarial learning to assist the training of our model. Extensive experiments demonstrate the effectiveness of the proposed model.
Both syntactic and semantic structures are key linguistic contextual clues, in which parsing the latter has been well shown beneficial from parsing the former. However, few works ever made an attempt to let semantic parsing help syntactic parsing. As linguistic representation formalisms, both syntax and semantics may be represented in either span (constituent/phrase) or dependency, on both of which joint learning was also seldom explored. In this paper, we propose a novel joint model of syntactic and semantic parsing on both span and dependency representations, which incorporates syntactic information effectively in the encoder of neural network and benefits from two representation formalisms in a uniform way. The experiments show that semantics and syntax can benefit each other by optimizing joint objectives. Our single model achieves new state-of-the-art or competitive results on both span and dependency semantic parsing on Propbank benchmarks and both dependency and constituent syntactic parsing on Penn Treebank.
Implicit relation classification on Penn Discourse TreeBank (PDTB) 2.0 is a common benchmark task for evaluating the understanding of discourse relations. However, the lack of consistency in preprocessing and evaluation poses challenges to fair comparison of results in the literature. In this work, we highlight these inconsistencies and propose an improved evaluation protocol. Paired with this protocol, we report strong baseline results from pretrained sentence encoders, which set the new state-of-the-art for PDTB 2.0. Furthermore, this work is the first to explore fine-grained relation classification on PDTB 3.0. We expect our work to serve as a point of comparison for future work, and also as an initiative to discuss models of larger context and possible data augmentations for downstream transferability.
Foreign language education primarily aims to cultivate learners’ competence to communicate in an additional language. However, the meaning of communication competence is not entirely transparent, especially given the current neoliberal valorization of communication in the knowledge economy. The meaning of communication can be scrutinized in two contradictory trends observed in language education: the exclusive focus on teaching English as a global language, signifying a homogenizing trend, and increased scholarly attention to the heterogeneity of linguistic forms and practices. This article examines how communication competence is differentially understood by policymakers and corporate workers in Japan. The authors examine a government report that evaluated the attainment of educational goals for coping with globalization and contrasting it with interview data drawn from another study on the communicative experiences of Japanese transnational workers in Asia. Political discourse analysis and content analysis reveal the paradoxical nature of what can be called neoliberal communication competence, which on the one hand conflates global communication with use of the four measurable skills in English to transmit information and, on the other hand, challenges linguistic norms, foregrounding plurilingualism and co‐constructed interactional competence. Transformation of policies and pedagogies can be pursued by appropriating neoliberal communication competence for achieving broader educational goals.
This dataset in CSV format contains all books from the web http://books.toscrape.com which has been got using web scraping method in November 2020. The CSV file has 12 columns called each of them like: title, image, rating, description, category, UPC, producttype, priceextax, priceincltax, tax, availability, numberreviews. The project was born as a practice for a subject of the Master of Science (MSc) of Data Science at the Universitat Oberta de Catalunya (UOC).
Universal Dependencies is an open community effort to create cross-linguistically consistent treebank annotation for many languages within a dependency-based lexicalist framework. The annotation consists in a linguistically motivated word segmentation; a morphological layer comprising lemmas, universal part-of-speech tags, and standardized morphological features; and a syntactic layer focusing on syntactic relations between predicates, arguments and modifiers. In this paper, we describe version 2 of the guidelines (UD v2), discuss the major changes from UD v1 to UD v2, and give an overview of the currently available treebanks for 90 languages.
Being a completely new communicative environment, the internet network generates new designations (including verbal) for several phenomena. Among them, there are blended nonce-words that are on the verge of norms violation. A number of these lexical units, which should be considered as a demonstration of linguistic personality creative potential, are related to the stock vocabulary of native twenty-first century Russian-language speakers. This phenomenon occurs various languages and different cultures. Therefore the article dwells upon blended neologisms, including native Russian and borrowed elements. The predominant role of situation and context in that pattern is obvious: those nonce-words have not been fixed in lexical system yet. Contaminants should be treated as an evidence of sustainable development of the whole language system, as a component of various discourses as well as integral part of everyday communication (via the internet network), and slang. Non-standard structure as well as combination of native and borrowed elements in blended nonce-words, provide them with strong emotive and expressive potential. Keywords: blended nonce-words, Internet communication, word building, borrowing
This article is devoted to the development of a linguistic model for describing the concept WOMAN. The material is women’s dialect discourse. The sources of the material are the Tomsk dialect corpus which includes materials of expeditions organized by dialectologists of Tomsk State University from 1946 to the present days on the territory of Middle Ob dialects spread. In the article we used modeling method based on the idea of the nominative field of a concept, as well as an interpretation technique relying on analysis of contexts, and a method of quantitative calculations used in relation to units that represent the concept. Lexical and phraseological units that make up the nominative field of the concept were revealed during the research. These units were divided into the following lexical-semantic groups: 1) the general nominations of a female person; 2) age and status in marriage; 4) status in the family hierarchy; 5) anatomical and biological characteristics; 6) character traits and behavior; 7) appearance characteristics; 8) profession and work processes. Elements of different layers of the concept are revealed in each lexical- semantic group. All of them give a general picture of ideas about women. So, the basis for identifying of gender conceptualizations and stereotypes is the presence of linguistic oppositions of male and female; the presence of a large number of lexical units that reflect the status of marriage (girl, bride, young woman, wife, mistress, old woman, widow, old girl, brooch and so on); lexical pairs that are opposed to each other on the basis of evaluation “positive” – “negative” (clean, clean – dirty, mistress – disheveled, etc.). A large number of words that negatively assess certain qualities and behavior of women (gossip girl, market woman, stramovka, etc.) indicate the high requirements imposed on the woman, the condemnation of deviations from social norms. The content of the concept of WOMAN depends on the specifics of rural existence, which is based on work, the presence of patriarchal gender stereotypes, social and historical events and processes. The significance of the research is determined by the possibility of using its results for development of a new interdisciplinary scientific field – gender dialectology that studies the gender characteristics of the dialect.
Abstract Lewis Carroll’s Alice’s Adventures in Wonderland ( 1865 ) and Through the Looking-Glass ( 1871 ), two linguistic treatises in disguise, create ingenious fantasy worlds where the rules of language and the conventions of communication are turned upside down. What is (semantically) illogical or (pragmatically) inappropriate confounds Alice, who struggles to make sense of nonsense and to keep the order of a polite, rational world in place. In her dialogues with anthropomorphic animals and objects, ambiguity and fallacy coexist with interactive manipulation, while her communicative expectations crumble and comic misunderstandings arise. This article looks into the construction of linguistic and pragmatic transgressions in Carroll’s acclaimed books with a view to unveiling their contribution to impoliteness. On the one hand, the paper analyses the structural mechanisms of wordplay vis-à-vis phonetic, morpho-syntactic and lexical ambiguity. On the other, it examines the pragmatic strategies whereby speech-act infelicities, conversational maxim violations, and bald-on-record clashes contribute to reversing the established conventions of (polite) social interaction. The premise guiding the analysis is that the pervasive existence of double meaning and incongruity in the Alice books underlies not only linguistic phenomena such as punning, neologism, and relexicalisation, but also interactive patterns, in which the expected norms of courteous conduct in social exchanges do not obtain. The antithetical and script-oppositional (hence, humorous) nature of this process defrauds outsider Alice – the victim, but at times the happy recipient, of the uncooperative challenges of this inverted, refracted, teasingly nonsensical world.
The purpose of this article is to present the characteristics of a group of anthroponyms in order to point out some directions for studies of Brazilian anthroponymy. This article is based on theoretical assumptions of Onomastics and on the interface between this field of study and Law. The anthroponyms analyzed are civil name, social name, ballot name and parliamentary name. Data were collected from the Superior Electoral Court, the Chamber of Deputies and court decisions from tribunals. Recent Brazilian anthroponymy studies demonstrate that research on personal names relating linguistic and legal aspects is still incipient. This article provides some suggestions that could bridge such a gap by analyzing lexical or grammatical aspects of data originating from legal norms or judicial decisions.
Käesolevas artiklis uuritakse kõrvutavalt originaaltekstiga Fjodor Dostojevski romaani „Vennad Karamazovid“ kahte eestikeelset tõlget, mille autoriteks on Aita Kurfeldt ja Virve Krimm. Analüüsi objektiks on jutustaja muutlik diskursus, mille eripära avaldub stiili ebaühtluses ehk muutlikkuses. Jutustaja „takerduval“ kõnel on romaanis oluline funktsioon, mis seisneb kaootilise, ebakindla kunstilise maailma loomises. Eestikeelseid tõlkeid vaadeldakse võrdluses lähtetekstiga mitmel mikrostilistilisel tasandil: kesksõnatarindid, sõnade ja sõnatüvede kordus, modaalsõnad, deminutiivid, fraseoloogilised üksused, grammatilistest normidest kõrvalekaldumine.
 
 The article studies two Estonian translations of Dostoyevsky’s novel The Brothers Karamazov by Aita Kurfeldt and Virve Krimm, comparing them to the source text. The tradition of translating Dostoyevsky’s works into Estonian has its beginning in the 20th century. It started with Johannes Aavik’s experimental translations and was continued by the classic of Estonian literature A. H. Tammsaare – in 1929, the first Estonian translation of Crime and Punishment appeared in the latter’s translation. In the 1930s, preparations began in Estonia to publish Dostoyevsky’s collected works in 15 volumes, and, as part of this initiative which involved several translators, the novel The Brothers Karamazov first appeared Estonian in Aita Kurfeldt’s translation (1939–1940). Kurfeldt’s translation was later edited and updated by Helle Tiisväli, and the new edition published by the Kupar publishing house in 2001. In the 21st century, the novel was translated for the second time and published by Varrak with an afterword by Peeter Torop in 2015–2016. The translator was Virve Krimm, a capable and talented translator who had already translated Dostoyevsky’s Demons as well as other books by classic Russian authors, e.g., Turgenev’s novel Home of the Gentry, his stories and prose poems; she had also been a co-translator of Tolstoy’s War and Peace. In Krimm’s obituary by the Translators’ Section of the Estonian Writers’ Union, her translation of The Brothers Karamazov was highly appreciated. Both translations were made during times free of the prescriptive norms of the Soviet regime. If ideological coercion in the narrower sense of the word (the authorities’ pressure on translators, editors and publishers) is considered, both translations can be regarded as expressions of the translators’ free choice – both were completed in free Estonia. 
 A conspicuous characteristic of Kurfeldt’s translation is her word-for-word reproduction of Dostoyevsky’s phrases or whole syntactic periods, preserving even the word order. The author of the later translation as well as the later editor of Kurfeldt’s translation have clearly tried to actively oppose Kurfeldt’s tendency towards literal translation. Still, the first translator’s “literal translation” cannot be claimed to be an indicator of dilettantism, as Kurfeldt’s attempts to copy Dostoyevsky’s syntax and even punctuation may be viewed as an essential effort to revive the narrator’s changeable, clumsy manner of speech in The Brothers Karamazov.
 This article analyses the narrator’s transmutable discourse, the peculiarity of which is expressed in the inconsistent or unstable style, as well as its translations into Estonian. In the novel, the narrator’s inconsistent speech has an essential function which consists in creating a chaotic, unstable artistic world. Studies of Dostoyevsky’s poetics have often drawn attention to the peculiarity of the narrator’s style and tone in his works. Mikhail Bakhtin noted that the narrator’s word constantly fluctuates between two extremes – the dryly informative, recording word and the word depicting the character. The researchers who have followed or developed Bakhtin’s theoretical conception have also noted that such inconsistent and hesitant narration style approaches, or actually is, non-literary language. As Aage A. Hansen-Löve has shown, the spontaneous or chaotic manner of narration was characteristic of the vanguard or initial period of Russian realism, but it was also preserved in the movement during its later years. Dostoyevsky modelled the type of the “non-professional” narrator as early as in the 1840s. The speech of this narrator is knowingly “non-literary”. The writer’s “carelessness with words” has also been described and analysed in literary studies as a deliberate device realised at different levels of the narrative: in composition (e.g. the stylistic inconsistency in chapter headings), syntax, lexical paradigm, structure of phraseological expressions and deviations from language norms.
 In this article, the Estonian translations are viewed in comparison with the source text on several microstylistic levels: participial constructions, repetition of words and word stems, modal words, diminutives, phraseological units, deviation from grammatical norms. The comparative analysis of the translations in the article does not attempt to characterise the translations in full, but only discusses the key tendencies in rendering the narrator’s unstable speech. The theoretical basis for the analysis derives from the virtual model of different translation types presented in Peeter Torop’s article “Tõlkeloo koostamise printsiibid” (“Principles of compiling translation history”, 1999).
 In conclusion, it appears that Kurfeldt’s translation is a text dominated by an orientation towards the expressive plane of the source text. The word order in sentences, punctuation marks, modal words and their positions in the text are rendered exactly. Still, the translation is inconsistent at the microstylistic level: the translator tries to replace functional repetitions occuring in the text with synonyms, changes participial constructions into subordinate clauses, and presents participles as verbs in the third person; in a number of cases Kurfeldt also omits words and phrases. The edited translation has undergone essential changes in its turn – the editor has striven for stylistically correct, fluent, “proper” speech which sometimes remains rather far from the original.
 Krimm’s translation has a considerably more complicated structure. Initially, it can be said that Krimm’s translation is oriented simultaneously towards the content plane of the source text, i.e. towards lexical and semantic precision, and sometimes also towards an equivalence with the rhythmic and intonational level of the expression plane of the original. Still, the precision of translating other levels of the expression plane of the original depends on the essentiality of the translated elements in the structure of the novel. Similarly to Kurfeldt, Krimm does not attempt to preserve diminutives, as these grammatical forms are not characteristic of the Estonian language. Thus, opting for an orientation mainly towards the expressive plane of the target text, Krimm continues many aspects of her personal tradition of translating Russian classics from the second half of the 20th century. Choosing the expression plane of the target text as a dominant was characteristic of many other Estonian translators in the Soviet period, as such a translation strategy compensated for the lack of political freedom. 
 The conclusions of the article concern only the recreation of the narrator’s uneven speech in the Estonian translations of The Brothers Karamazov by Kurfeldt and Krimm and, at this stage, do not expand to encompass other layers of the complicated structure of Dostoyevsky’s novel in the texts by the two translators. The article serves as the beginning of a study: further, both translations could be viewed in a broader ideological context, considering the dependence of concrete translation solutions on the translation norms of the 1930s, the normative requirements for literary translation in the 21st century, problems of editing of translations, as well as aspects related to political, literary, linguistic, intermedial and other translation-related contexts.
Abstract This study investigated the perceptions of high school EFL learners to the lexical instructional approach intervention in the contexts of learning vocabulary and grammar. Besides, an attempt was made to explore what difficulties the participants encountered during the experimentation. The data collected through the questionnaire were analyzed using a one-sample t-test, and the results showed that the estimated sample perception mean score was significantly higher than the hypothesized population perception mean score. This implies that EFL learners had positive perceptions towards the lexical instructional approach in the contexts of learning vocabulary and grammar. The data collected through interviews were analyzed qualitatively and the findings showed that students enjoyed and were interested in learning vocabulary and grammar through the lexical instructional approach. Students realized the importance of lexical chunks in learning vocabulary and grammar. In this regard, the interview results corroborated the results obtained from the questionnaire. Students encountered difficulties like lack of lexical awareness, lack of clear and adequate instructions on some activities, the lack of deliberate attention from some students during discussions, the lack of making some activities more interactive and engaging, and some classroom managerial problems. Finally, it was recommended that EFL teachers at high school should design their lexical approach-based activities systematically by considering their students’ interests, feelings, perceptions, levels, norms, cultures, and psychological setups.
Abstract The aims of this paper are to analyse differences in the degree of lexical variation (type/token ratio and hapax/token ratio) of reporting verbs in reporting clauses placed medially or in postposition in English, French and Czech fiction and to evaluate their consequences in translation, especially in regard to explicitation/implicitation. We expect that, in translations from a language with a low degree of lexical variation of reporting verbs into a language with a high degree of lexical variation, the frequency and the degree of explicitation will be higher than in translations involving languages less different with respect to lexical variation. The analysis, relying on data extracted from the InterCorp multilingual corpus, proposes a classification of reporting verbs based on the type and amount of information conveyed, which allows evaluating the degree of explicitation operated in translations. The results show that most shifts involve only the neutral reporting verb say/dire, replaced by a stylistically more specific synonym or by a verb explicitating information obvious from the context. This suggests that modifications of reporting verbs in translation are motivated primarily by respect for the stylistic norm of the target language and the degree of acceptability of the repetition of the neutral reporting verb.
The article studies the problem of interconnection and interpenetration of semiotic and gender discourses. The main linguistic level of gender representation is the lexical level in combination with the semantic one because the semantic parameters of the linguistic picture of the world are verbalized in the vocabulary. It is proved that with the help of feminine words the online newspaper Kolo solves a number of problems due to the institutional parameters of media discourse, which determine its context and represent four groups: social, cultural, ideological, communicative, semiotic. The general tendencies of the functioning of feminine words both in headings, and in texts are formulated: formation of public opinion (appeal to the formed stereotypes); subordination of plans of expression and content to the conditions of the mass communication environment (the question of interaction of language signs with signs of other types); establishing and maintaining trust in the sender of information by the recipient in order to ensure the reputation of mass media (compliance with the requirements of a gender-sensitive environment); preservation or violation of legal, cultural, ethical norms, worldview paradigm and picture of the world (correspondence and influence on the conceptual paradigm of the world). Gender discourse is presented in materials about the historical past, political, socio-economic, socio-domestic, cultural, sports, family life, criminal chronicle, health. Explicitly marked elements of gender form mainly the following lexical and semantic groups: names of persons by professional activity, occupation, names of persons by ethnicity, nationality, religion and belief, a territory of residence, position, rank, belonging to political groups, parties, character relationships with people of the opposite sex. Modern media realia make new demands on the online newspaper as: the publication of news in the form of a post, so it should be visible in the news feed. This task is realized by influencing through two channels of perception simultaneously: visual (image) and verbal (text). The course of all information processes takes place with the use of signs and sign systems. The visual content of the issues of the online newspaper Kolo includes illustrations – photos of men and women, as well as shared photos, gender accents are shifted to the female component. It is determined that the Ukrainian word-forming formant (-к-, -иц-, -ин-) in nouns of female names is the main feminine marker.
An attempt was made in the article to identify sources and ways to update the credit vocabulary in German at the end of XVII – mid XIX centuries. The article determines the structure of the correlated with the mentioned sphere German vocabulary during mentioned period. The article specifies the peculiarities of the analyzed terminology caused by both general trends in the development of the German language in the period of formation of the national norms of word usage (presence of a large number of synonymous and double-headed designations) and the specifics of folding of the lexicon of the financial and credit sphere (preservation of naming units of dying concepts, functioning of intermediate lexemes). The research allows us to state that at the current stage of credit terminology development the special dictionary is replenished in the course of lexical borrowing processes from such Western European languages as Italian and French. As for the intra-linguistic word-production of credit nominations, it took place mainly on the basis of word-producing resources in the course of semantic and morphemic derivation, with the obvious prevalence of addition.
The article analyses creative work of the German writers of the fourth migration wave. Stylistic, grammatical and lexical peculiarities of their works, including techniques of grotesque, usage of borrowings and neologisms, are identified. The paper aims to correlate the identified peculiarities with the norm of the German literary language. Originality of the study involves analysing syntactic and lexical peculiarities of the migrant writers’ novels.
The phenomenon of language game as a way of creating a journalistic image, also humorous, ironic, satirical, can be called one of the most actively developing traditions in modern Russian media. The aesthetics of postmodernism has increased interest in both the language game and the humour – an effective way to assess the events of reality, and to express the author's position. In the situation of language inflation, the audience, tired of the news, formed a request for a multifunctional text, involving co-creation, provided exactly by language game. Researchers define "the virus of irony" as a characteristic feature of modern media text, with the paradigm of irony constantly expanding: from light humor to destroying sarcasm and grotesque. The paper attempts to identify and systematize the functional types of language game as a way to create a comic effect. On the basis of the analysis of prominent journalists’ works the authors determined the techniques of language games, which are most actively used to create the humour – humorous, ironic, satirical – image. These are paronomasia, pun, occasionalisms, transformation of phraseological units, stylistic contrast. Language game as a deliberate violation of the norm is manifested at different levels of the text: grammatical, lexical-semantic, syntactic, and stylistic. The role of the humour in the headlines of modern publications of different directions is also characterized. Special attention is given to principles for the publicists’ appeal to laughter as a simple and sharp form of criticism that allows implementing a variety of communicative intentions: humour, outrage, and others.
Bilingual children show more variation in their language development than monolingual children, a fact that has been linked to their experience with their languages. Bilingual language experience also varies more than monolingual children's, both in terms of how much they hear the language spoken around them (exposure) and how much they speak the language themselves (production). This dissertation investigates the following aspects of the relationship between bilinguals’ language experience and development which are not well-understood: how children’s language production relates to their proficiency in that language, how children’s language exposure relates to receptive versus expressive and lexical versus grammatical skill, and how factors such as social context, cognates, working memory and indirect exposure contribute to bilingual proficiency. I investigate language experience and English proficiency in young school-aged bilinguals acquiring French and English in France. I use data from parental and child interviews to estimate English exposure – how much children regularly hear English – and two facets of English production – output, or how regularly children speak in English, and inter-speaker code-switching, which refers to how regularly children respond in French when spoken to in English. Those measures are then related to English proficiency scores from a picture-identification task, a picture-naming task, and a sentence repetition task targeting grammatical structures ranging in difficulty. The first objective of this study is to better understand bilingual children’s language production as it relates to their language proficiency. I find that how much children switch to speaking in French when addressed in English (inter-speaker code-switching) is closely related to all concurrent English proficiency scores and that this relationship is independent of and stronger than proficiency’s relationship with exposure. The more children switch to French when spoken to in English, the lower they score on all proficiency measures, receptive and expressive vocabulary, and sentence repetition, even when holding their level of English exposure constant. The second objective of this study is to investigate possible limits to the general pattern found in a large body of research on bilingual exposure, which is that lesser exposure leads to lesser skill in that language. First, language exposure may affect receptive skills less than expressive skills. Second, grammatical knowledge may also be less closely related to exposure than lexical knowledge. There are conflicting findings in the literature. My findings are consistent with a weak relationship between receptive skills and language exposure in bilingual children. Despite having lesser exposure to English (34% of their total language exposure), children in this study did not show a relation between variation in exposure and their English receptive vocabulary scores. In these children, the relationship between exposure and grammatical proficiency was similar to that with lexical proficiency. The third objective is to investigate additional contributors to bilingual proficiency. Previous research suggests that children’s socioeconomic status (SES), the status of the languages they speak, and the existence of cognates in their languages make contributions to bilingual children’s proficiency, and may in turn modulate the effect of diminished language exposure (e.g. Cobo-Lewis, Pearson, Eilers, & Umbel, 2002a; 2002b; Thordardottir, 2011). My results suggest that SES and high prestige of the languages being acquired may partially mitigate – though not eliminate – the effect of diminished exposure on bilinguals’ home language proficiency. Similar to findings for other bilingual children from mid- to high-SES backgrounds, these children showed age-related growth in English proficiency, and their English receptive skill differed minimally from monolingual norms. However, the effect of lesser exposure to English can be seen more clearly in their expressive skills, which were lower than monolingual norms and were predicted by variation in their English exposure. The effect of cognates in French and English was also investigated in terms of the advantage they conferred on my measures of lexical proficiency. This effect was significant in both receptive and expressive measures; thus, I conclude that the presence of cognates may also mitigate the effect of bilingual exposure. Finally, this investigation also examines additional individual factors that can influence language proficiency, but which have rarely been taken into account in studies of bilingual proficiency and both its relationship to exposure and production. Specifically, variation in children’s working memory and their exposure to language through overhearing adult conversation have both been linked to language learning in monolingual contexts but are not well understood in the context of bilingual development. In this study, verbal and visuospatial working memory were positively related to English proficiency scores. Indirect exposure from overheard English spoken between parents was not related to proficiency scores when holding direct English exposure from parents constant. However, indirect exposure was related to how much children produce English themselves to their parents, even while holding direct exposure constant, indicating that language use between parents may influence children’s language production with parents. This study contributes to our understanding of how bilingual language exposure and production relate to bilingual language proficiency in the following ways: first and most importantly, it adds to the small but growing literature that shows a strong link between bilingual children’s own production of a language and their lexical and grammatical skill in that language. It is also the first to my knowledge to find that a measure of children’s language production, inter-speaker code-switching, is negatively related not only to expressive but also to receptive lexical skill in the language that children switch from. Secondly, the finding that children’s English exposure is unrelated to their English receptive skill (but related to age, indicating continuing growth in these children) affirms exposure’s differential relationship with receptive versus expressive skills. It also documents a limited role for exposure in a new population (French-English bilinguals in France), supporting the role of cognates, socioeconomic status of children, and high social prestige of languages being acquired in mitigating the effect of bilingual exposure. Finally, in finding an independent contribution of working memory to lexical and grammatical skill in bilinguals, it highlights that these measures should be considered when investigating variation in bilingual proficiency.
The topic of this thesis is the computational methods for measurement of authorialstyle and algorithms of authorial attribution.The first aim of the thesis was an attempt at a quantifiable separation of various layers of authorial style (in the present case the lexical and grammatical layers) in order to estimate their influence on the results of a chosen method of authorial attribution. Within the scope of these studies I compared the distance, so called Burrows's Delta, between a pair of English novels by two chosen authors and automatically generated texts, whose statistical distributions of parts of speech were borrowed from one of the authors, while the vocabulary from the other one; additionally, in the computatrificial texts I left the sets of words of the first author if they belonged to a particular part of speech. Such procedure allowed to create a hybrid text, which was attributed to the first author, even though the majority of lexical items were that of the second author.The second aim was to identify the influences of the style and language of the original on the style of the translation. This part of research involved among others adapting Polish and English part of speech tag sets to form a common translatorial tag set. Beside making a couple of simple observations concerning the distributions and coocurrences of parts of speech in the two languages, I managed to determine some features of the selected translatorial corpus, which lie on the fringes of what seems a norm for Polish.The third aim was testing the accuracy of state of the art (unsupervised) clustering methods for automatic grouping of texts according to their author. The results show that the methods recognise authorship worse than the known supervised machine learning methods.In the thesis I made use of corpora totalling around 550 digitised English language novels and 100 Polish ones, as well as a parallel corpus of 39 novels of a single English author together with their translations by a single Polish translator. The research conducted involved utilising existing part of speech taggers (both for English and Polish), authorship attribution programmes, and programmes for graph clustering.
The article is concerned with the problem of correlation of the homogeneity and the co-ordination in French that is essential to differentiate a simple sentence with the similar verb predicates of a complex sentence. The urgency of such problems is based on the similarity of these syntactic constructions due to the co-ordination link existing in both constructions. This fact doesn’t allow the grammarians to arrive at a common view on the nature of the two constructions. The author proves the influence of the verb predicate syntactic links with the other parts of the sentence on classifying the structure as a simple or a complex sentence. In the paper there have been studied the similar verb predicates in the extended and unextended sentences. In the extended sentences the author focuses on the form and place of a complement, on the presence or absence of the adverbial modifier. The verb predicate grammar form itself influences the differentiating the two structures. Thus, it has been concluded that the main distinctive feature of predicate homogeneity is the grammatical marker. There have been detected the supplementary distinctive feature of predicate homogeneity is the semantic aspect, the lexical meaning in particular. The treated analysis of the empiric material shows the dependence of determining the two syntactic units on the stylistic norms and the rhetorical mode. The most important finding of the research is that, contrary some scientists’ opinion, there is no reason to abandon the term of the similar verb predicates in French.
Abstract This article is a synthesis of the major elements of a sociolinguistic theory presented by Jean Le Dû and Yves Le Berre in their recent book, Métamorphoses, Trente ans de sociolinguistique à Brest (1984–2014). Given that both authors come from native Breton-speaking families in Western Brittany and have experienced the language shift to French first-hand, they provide a unique, inside view of the process as well as the reasons Breton speakers opted in favour of French. The sociolinguistic concepts they have imagined provide highly useful tools that highlight the inseparable bond between language and the social, political and economic forces that govern our choices. More specifically, they point out that the “Breton language” is splintered into as many varieties as there are social and geographic entities in western Brittany. For this reason, it should not be viewed as a monolithic entity. Far from “reviving” or “saving” the language, the authors argue that the recent creation of a phonologically, grammatically and lexically unified Breton norm is often so distant from the vernacular language that it has provoked a new form of diglossia which failed to reverse the break in the transmission of the natural language. The book provides tremendous insight into the complex issues which lead people to shift to another language. Language planners and scholars working on similar endangered language situations and who want to understand the mechanisms at work (and thus hopefully have some success in their endeavours) would do well to take heed of their experience.