Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
In this article describes the results of research on students of UIN Raden Fatah Palembang who can distinguish standard and non-standard languages, it seems that there are still not being able to distinguish these languages, so the purpose of the research is the ability and understanding in distinguishing standard and non-standard languages according to the language,research findings reveal that 85% of students can distinguish standard languages. According to linguistic rules and 15% of students who cannot distinguish non-standard languages. There are many ways to obtain data, including through observation and interviews. according to research results some people still can't understand why this is so because standard Indonesian is needed to make it easier and function as the national language, the language of unity, the pride language of the Indonesian language. In the world of education, Indonesian language subjects are considered to be able to help students to be able to understand and understand the rules of language, therefore students are expected to be able to distinguish them based on the context and purpose of their use not violating Indonesian linguistic norms. the ability of students to recognize and master the appropriate language in accordance with the relevant Indonesian rules. This can help students to be able to distinguish and master vocabulary.
Adaptive gradient-based optimizers, notably Adam, have left their mark in training large-scale deep learning models, offering fast convergence and robustness to hyperparameter settings. However, they often struggle with generalization, attributed to their tendency to converge to sharp minima in the loss landscape. To address this, we propose a new memory-augmented version of Adam that encourages exploration towards flatter minima by incorporating a buffer of critical momentum terms during training. This buffer prompts the optimizer to overshoot beyond narrow minima, promoting exploration. Through comprehensive analysis in simple settings, we illustrate the efficacy of our approach in increasing exploration and bias towards flatter minima. We empirically demonstrate that it can improve model performance for image classification on ImageNet and CIFAR10/100, language modelling on Penn Treebank, and online learning tasks on TinyImageNet and 5-dataset. Our code is available at \url{https://github.com/chandar-lab/CMOptimizer}.
The Dagstuhl Seminar 23191 entitled "Universals of Linguistic Idiosyncrasy in Multilingual Computational Linguistics" took place May 7-12, 2023. Its main objectives were to deepen the understanding of language universals and linguistic idiosyncrasy, to harness idiosyncrasy in treebanking frameworks in computationally tractable ways, and to promote a higher degree of convergence in universalism-driven initiatives to natural language morphology, syntax and semantics. Most of the seminar was devoted to working group discussions, covering topics such as: representations below and beyond word boundaries; annotation of particular kinds of constructions; semantic representations, in particular for multiword expressions; finding idiosyncrasy in corpora; large language models; and methodological issues, community interactions and cross-community initiatives. Thanks to the collaboration of linguistic typologists, NLP experts and experts in different annotation frameworks, significant progress was made towards the theoretical, practical and networking objectives of the seminar.
Our understanding of a scene, although typically described only within the visual domain, can be influenced by other modalities. Here, we examine the link between visual and auditory cognitive processing, or cross-modal processing, using an affective priming paradigm. Current theories do not typically incorporate scene affect into models of scene understanding, yet we explore how the affect associated with a scene can be modulated by music. In the current experiment, participants (N = 39) rated both how much they enjoyed musical excerpts and images of everyday, neutral scenes. A novel musical stimulus dataset was created for the present study to ensure the musical examples did not carry semantic associations and that the observed effects would be attributed to their affectual influence. This dataset included sixty-four original miniature piano compositions composed with features controlled along six binaries. Participants listened to a brief musical excerpt, reported an affect rating from really dislike to really like, then viewed and rated a neutral scene (e.g., a dining room). A significant difference between scene affect ratings after participants heard music they disliked and liked was found on both the participant and individual scene levels. These results imply auditory processing plays a role in scene understanding. Not only were participants’ scene ratings modulated by the affect of the musical stimuli, but the same scene was rated more positively or negatively depending on the affect of the preceding musical example. Crossmodal processing occurs between music and scene perception, and our results demonstrate how one can affect the perception of the other.
We introduce an encoding for parsing as sequence labeling that can represent any projective dependency tree as a sequence of 4-bit labels, one per word. The bits in each word's label represent (1) whether it is a right or left dependent, (2) whether it is the outermost (left/right) dependent of its parent, (3) whether it has any left children and (4) whether it has any right children. We show that this provides an injective mapping from trees to labels that can be encoded and decoded in linear time. We then define a 7-bit extension that represents an extra plane of arcs, extending the coverage to almost full non-projectivity (over 99.9% empirical arc coverage). Results on a set of diverse treebanks show that our 7-bit encoding obtains substantial accuracy gains over the previously best-performing sequence labeling encodings.
Abstract This chapter examines issues that arise in speech communities in response to changes in the language. Some usages are more appropriate for formal contexts than for informal ones, while usages out of place in formal contexts may be perfectly fine in informal ones. Schooling tends to accord formal styles a privileged status over others, while a great deal of change comes from the language spoken more casually. While the singular noun English encourages us to view the language as a single, unified speech form, this has never been so. A standard language is a set of linguistic norms established by some generally accepted political or social authority. But different speech communities can have different standards, leading to greater variety across English dialects. Furthermore, each dialect by itself encompasses a variety of “standards,” depending on whether we are speaking or writing and to whom. Writing is generally more conservative, while oral language changes far more readily—to the dismay of those who prefer to leave things as they are.
When processing written German language, it is helpful, to use the base form (or: lemma) of possibly inflected words, such as verbs, nouns or named entities. However, for German text from the (bio)medical domain, e.g., discharge letters, or entries stored in electronic medical or health records (EMR, EHR), difficulties exist in finding the correct lemma, as, for instance, the medical language has roots in Latin or Greek. In such cases, stemming techniques might provide inaccurate results for text written in German. This study demonstrates a Machine Learning approach for training Apache OpenNLP-based lemmatizer models from publicly available German treebanks. The resulting four "DE-Lemma" models were evaluated against a sample of (bio)medical nouns, randomly selected from real-world discharge letters. The most promising DE-Lemma model achieved an accuracy of 88.0% (F1 =.936).
Recent research on shallow discourse parsing has given renewed attention to the role of discourse relation signals, in particular explicit connectives and so-called alternative lexicalizations. In our work, we first develop new models for extracting signals and classifying their senses, both for explicit connectives and alternative lexicalizations, based on the Penn Discourse Treebank v3 corpus. Thereafter, we apply these models to various raw corpora, and we introduce ‘discourse sense flows’, a new way of modeling the rhetorical style of a document by the linear order of coherence relations, as captured by the PDTB senses. The corpora span several genres and domains, and we undertake comparative analyses of the sense flows, as well as experiments on automatic genre/domain discrimination using discourse sense flow patterns as features. We find that n-gram patterns are indeed stronger predictors than simple sense (unigram) distributions.
Music-evoked autobiographical memories (MEAMs) are typically elicited by music that listeners have heard before. While studies that have directly manipulated music familiarity show that familiar music evokes more MEAMs than music listeners have not heard before, music that is unfamiliar to the listener can also sporadically cue autobiographical memory. Here we examined whether music that sounds familiar even without previous exposure can produce spontaneous MEAMs. Cognitively healthy older adults (N=75, ages 65-80 years) listened to music clips that were chosen by researchers to be either familiar or unfamiliar (i.e., varying by prior exposure). Participants then disclosed whether the clip elicited a MEAM and later provided self-reported familiarity ratings for each. Self-reported familiarity was positively associated with the occurrence of MEAMs in response to familiar, but not the unfamiliar, music. The likelihood of reporting MEAMs for music released during youth (i.e., the “reminiscence bump”) relative to young adulthood (20-25 years) included both music released during participants’ adolescence (14-18 years) and middle childhood (5-9 years) once self-reported familiarity was accounted for. These developmental effects could not be accounted for by music-evoked affect. Overall, our results suggest that the phenomenon of MEAMs hinges upon both perceptions of familiarity and prior exposure.
Individuals with depression experience more negative imagery and less vivid positive imagery, and the late positive potential (LPP) is considered as a viable biomarker for negative attentional and memory biases in depression; however, the LPP response to emotional imagery in depressed individuals remains unclear. This study aims to investigate neural response to emotional imagery in depressed individuals. ERPs were recorded from 40 depressed participants and 44 healthy controls during the encoding-imagery task. Depressed participants scored significantly lower in the valence rating of sad and neutral imagery compared to healthy participants. Importantly, the LPP amplitudes to sad imagery in depressed participants were significantly larger than healthy controls, particularly in the middle (800-1,400 ms) and late time windows(1,400-2,000 ms). Furthermore, depressed individuals exhibited significantly higher LPP amplitudes for sad imagery compared to happy imagery, whereas healthy participants showed the opposite pattern. The present study provides evidence that depressed individuals display abnormal electrophysiological reactivity to sad imagery, which offers a new perspective for understanding the mechanisms underlying depression.
Dialogue-level dependency parsing has received insufficient attention, especially for Chinese. To this end, we draw on ideas from syntactic dependency and rhetorical structure theory (RST), developing a high-quality human-annotated corpus, which contains 850 dialogues and 199,803 dependencies. Considering that such tasks suffer from high annotation costs, we investigate zero-shot and few-shot scenarios. Based on an existing syntactic treebank, we adopt a signal-based method to transform seen syntactic dependencies into unseen ones between elementary discourse units (EDUs), where the signals are detected by masked language modeling. Besides, we apply single-view and multi-view data selection to access reliable pseudo-labeled instances. Experimental results show the effectiveness of these baselines. Moreover, we discuss several crucial points about our dataset and approach.
People's perceptions are influenced by several sentences. The reader is better able to understand the statement through various entities thanks to these perceptions. Named Entity Recognition (NER) is the term used for this technique in NLP. Confusion about whether a word represents the name of a person, place, or organization, or whether an integer represents a date, time, or amount of money, is one of the main issues in these NLP processes. One of the difficult jobs that previously required a great deal of expertise in the area of feature engineering along with lexical databases to attain good performance is named entity recognition. Because of this, we create Scikit-learn as well as Keras algorithms for NER algorithms to accurately label the entities and divide them into the most likely and unlikely transitions. Finally, we implement Biological NER Labelling on protein and gene sequences to conclude our work.
This paper describes Hungarian infinitive constructions using a data-driven approach.It aims to study these constructions along formal-distributional features, utilizing authentic corpus data and frequency counts.Our findings derive from an open-source dataset extracted from a 776.9-million-word treebank of Hungarian.The resource contains more than 9 million instances of infinitive constructions, annotated for a wide range of linguistic features and metadata.We discuss the following topics: inflected infinitives, verb clusters consisting of multiple infinitives, auxiliarylike lexical items, separable preverbs, the attested word orders of infinitives, finite verbs and their respective preverbs, and finally, detailed distributions of preverbs in infinitive constructions.Our study reveals some trends which would have remained unseen without a quantitative approach.
Abstract The article maps the linguistic views on the common language users as non-experts in the field of linguistics and their ability to think reflectively and rationally about issues related to language. An overview of attitudes towards non-linguists is presented against the background of the development of linguistics, ranging from a structural understanding of language with an emphasis on standard language cultivation and linguistic prescription, through a sociolinguistic approach that emphasizes the role of the language user as a creator of the linguistic norm and its variation, to the view of folk linguistics and citizen linguistics, which examine how ordinary people in various forms of public communication present their opinions, beliefs, as well as their myths and ideologies about language. At the same time, the paper argues for the view that some folk knowledge and beliefs about language are not only incorrect or inaccurate, but, on the contrary, that they provide valuable information for linguistics about the background of language behaviour and language change. The material was drawn from the databases of Slovak linguistic journals and the specialized corpus of the journal Slovenská reč.
This study is a values-driven approach to figures of speech, depicting language and its standardisation. We explore a discourse about the modernisation of linguistic norms that took place in Estonian public media in 2020–2022 and reached the point of being labelled a crisis. The debate took place mostly in the form of opinion-writing texts, expressing the writers’ subjective perspectives. During the discussions, two parties with different outlooks on language and language planning issues emerged, representing the dichotomy of liberal and conservative value models. The focus of the study is on the interplay between values and patterns of figurative thought, as metaphors were extensively used to strengthen the arguments of both sides. The analysis, based on the theoretical-methodological means of the Conceptual Metaphor Theory, Figurative Framing, Metaphor Scenario Analysis, Systemic Functional Linguistics and Critical Discourse Analysis, revealed that the opposing parties favoured certain metaphors when depicting language. As a side issue, we also address the dynamics of power relations through the language crisis discourse.
Abstract In the contribution, we provide a theory-based and corpus-verified description of expressions for measure in Czech. We demonstrate that the measure expressions may modify quantity of entities ( approximately ten boys ), internal characteristics of events ( he works a lot ), properties ( very big ) and relations ( completely without sound ). We distinguish between the measure expressions that are an answer to the question To what extent? (Extent-modifiers) and expressions that modify an answer to the question How many? (Quantity-modifiers). The Extent-modifiers are formally, structurally and semantically more diverse than the Quantity-modifiers. For the Quantity-modifiers a list of forms and functions is provided. Theoretical knowledge stemming from the analysis will subsequently be used to improve the annotation in the Prague Dependency Treebanks. It can be also useful for other semantically-oriented descriptions of language.
Methods of defining ontologies, word disambiguation methods, computer systems, and articles of manufacture are described according to some aspects. In one aspect, a word disambiguation method includes accessing textual content to be disambiguated, wherein the textual content comprises a plurality of words individually comprising a plurality of word senses, for an individual word of the textual content, identifying one of the word senses of the word as indicative of the meaning of the word in the textual content, for the individual word, selecting one of a plurality of event classes of a lexical database ontology using the identified word sense of the individual word, and for the individual word, associating the selected one of the event classes with the textual content to provide disambiguation of a meaning of the individual word in the textual content.
This research aims to examine the strategies of inclusion and exclusion used by Republika.com in relation to the news coverage titled "Beredar Pesan Singkat Anies Ditolak Datang Isi Seminar di UGM, Pihak Rektorat Membantah" using Theo Van Leeuwen's critical discourse analysis. The findings of this research will be utilized to suggest teaching resources. This study is qualitative and adopts a descriptive approach. The news coverage "Beredar Pesan Singkat Anies Ditolak Datang Isi Seminar di UGM, Pihak Rektorat Membantah" is the object of the research, with Republika.com as the focus of the study. In this investigation, it was discovered that Republika.com's reporting employs tactics of exclusion and inclusion. The results of Theo Van Leeuwen's critical discourse analysis can contribute to the knowledge of the public and readers. The findings of this research can be applied to theory development, aiming to understand the structure and linguistic norms of news text, both in oral and written forms. Keywords: Theo Van Leeuwen's Discourse Analysis, News, Exclusion, Inclusion
Abstract Introduction Increased slow-wave activity (SWA) following sleep loss is a well-established physiological marker of increased sleep pressure. Recent research found SWA changes in response to sleep loss partially accounted for next-day decrements in positive affect (Finan et al., 2015; 2017). To our knowledge, however, this line of inquiry has been limited to adults. Recently, we found greater SWA during a night of healthy sleep predicted greater next-day positive emotion among pre-pubertal children (Rech et al., 2022). Here, we explored whether SWA during the second of two nights of partial sleep restriction predicted next-day emotional responses in the same cohort. Methods N=24 healthy, unmedicated children (7 to 11 years) completed a baseline assessment and one night of at-home PSG monitoring (10 hours in bed). Ten days later, children completed one night of at-home sleep restriction (7-hour sleep opportunity) followed by one night of in-lab sleep restriction with PSG monitoring (6-hour sleep opportunity). The next day, children completed an in-lab emotional assessment where they provided arousal and valence ratings in response to positive affective images from the International Affective Picture System (IAPS). Results Controlling for age and total sleep time during the first sleep restriction night, hierarchical multiple regression analyses examined if N3 SWA in frontal, central, and occipital regions during the second sleep restriction night predicted IAPS ratings. In separate models, both frontal, F(1,20)=8.535, p=.008, and occipital SWA density, F(1,20)=12.313, p=.002, predicted valence ratings in response to positive images (where lower scores indicated more positive valence). Conclusion While preliminary, findings indicate relationships between “rebound” SWA after sleep restriction and next-day emotional functioning in pre-pubertal children. Findings suggest that, prior to the pubertal transition, trait-based differences in homeostatic sleep pressure and recovery from sleep loss might alter emotional responding, potentially influencing risk for later mood disturbances. Support (if any) This research was supported by NIMH grant #R21MH099351 awarded to C.A.A.
Sociolects, as social varieties of language, should be classified and described on the basis of both linguistic and sociological criteria. In practice, however, most often only either the first one (e.g. according to Antoni Furdal and Danuta Buttler) or the second one is used (e.g. according to Aleksander Wilkoń and Tomasz Piekot). Both these criteria are less frequently combined (Stanisław Grabias). The article proposes those aspects of the sociological and linguistic functioning of language varieties that should be used by sociolinguistics especially in the characterisation of sociolects, but also in their typology. The sociological criteria include: the type of contacts in the group (including contacts made via the Internet), the degree of group formalisation, the durability of the group and the type of bond that binds the group together. Among the linguistic criteria the following were distinguished: nomination and expressiveness, the method of acquiring a sociolect and the degree of its codification and the rank of the linguistic norm. The article also clarifies the concept of the distinguishing features of the sociolectic vocabulary.
This paper confirms that, in English binary coordinations, left conjuncts tend to be shorter than right conjuncts, regardless of the position of the governor of the coordination. We demonstrate that this tendency becomes stronger when length differences are greater, but only when the governor is on the left or absent, not when it is on the right. We explain this effect via Dependency Length Minimization and we show that this explanation provides support for symmetrical dependency structures of coordination (where coordination is multi-headed by all conjuncts, as in Word Grammar or in enhanced Universal Dependencies, or where it single-headed by the conjunction, as in the Prague Dependency Treebank), as opposed to asymmetrical structures (where coordination is headed by the first conjunct, as in the Meaning–Text Theory or in basic Universal Dependencies).
Focusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet.In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g.elements of wordnet taxonomy, quantifier phrases, certain collocations).In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches.We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one).Language model also proves to be better than a feature-based logistic regression model.
Many primates produce copulation calls, but we have surprisingly little data on what human sex sounds like. I present 34 hours of audio recordings from 2239 authentic sexual episodes shared online, each with one vocalizer (1950 female, 289 male). Both acoustic features and arousal ratings from a perceptual experiment follow an inverted-U curve, revealing the likely time of orgasm. Sexual vocalizations become longer, louder, more high-pitched, voiced, and unpredictable at orgasm in both men and women. Men are not less vocal overall, but women start moaning at an earlier stage; speech or even minimally verbalized exclamations are uncommon. While excessive vocalizing sounds inauthentic to listeners, vocal bursts at peak arousal are ubiquitous and less verbalized than in the build-up phase, suggesting limited volitional control. Human sexual vocalizations likely include both consciously controlled and spontaneous moans of pleasure, perhaps best understood as sounds of liking rather than signals specific to copulation.
The goal of visual word sense disambiguation is to find the image that best matches the provided description of the word's meaning. It is a challenging problem, requiring approaches that combine language and image understanding. In this paper, we present our submission to SemEval 2023 visual word sense disambiguation shared task. The proposed system integrates multimodal embeddings, learning to rank methods, and knowledge-based approaches. We build a classifier based on the CLIP model, whose results are enriched with additional information retrieved from Wikipedia and lexical databases. Our solution was ranked third in the multilingual task and won in the Persian track, one of the three language subtasks.
In this paper, I will explore the concept of 'yakuwarigo' in Japanese language and present text analysis of three female characters, Eboshi, Rin and Yubaba from two anime movies of Hayao Miyazaki (Princess Mononoke, Spirited Away). My aim is to demonstrate how the well-known director employs role language, particularly masculine language, to empower his female characters to take on prominent roles in a society where men traditionally hold dominance. In terms of film analysis, Miyazaki places his female characters in the public sphere, making them active participants in the storyline while consistently defying traditional Japanese feminine conventions. This study is closely tied to the field of gender linguistics and linguistic ideology from an analytical perspective, aiming to illustrate how Miyazaki's female characters diverge from linguistic norms in their dialogues.
The article deals with the role of the contemporary Bulgarian linguist in the complex processes of codification of literary-linguistic norms. The text is linked to the 75th anniversary of the eminent Bulgarian linguist Prof. Vladko Murdarov. The object of the analytical observations in the text are problems related to the codification of literary-language norms, to the curricula and textbooks on Bulgarian language in secondary school; to the notion of language policy, as well as to the issues of the public image of the Bulgarian language in the media space. A brief overview is given of important historical processes that have shaped the development of the Bulgarian language and its science. The author summarizes the important features that every contemporary linguist should possess - depth of linguistic knowledge in synchronous and diachronic terms, a sense of the dynamics of linguistic processes in order to be able to objectively reflect the complex processes of language development, linking them to the complex and dynamic processes of social development.
The paper examines team building in multicultural and multiethnic work environments – particularly in NGOs – by linking classic accounts of effective teams with sociolinguistic and socio-emotional perspectives. It conceptualises workplace teams as speech communities in which members share, negotiate, and sometimes contest linguistic norms, expectations, and emotional display rules. Drawing on management and organisational studies, the text outlines key structural conditions for effective teams, including clear goals, appropriate leadership, resource allocation, and mechanisms for accountability and cooperation. These insights are then integrated with research on emotions, culture, and communication, which highlights the role of emotional attachment, trust, and creative problem-solving in sustaining team cohesion. The paper argues that effective team building in such contexts requires not only formal structures but also deliberate cultivation of shared communicative practices and affective climates that support participation, innovation, and mutual understanding across cultural and linguistic differences.
This chapter will briefly examine the two projects for historical dictionaries that the Royal Spanish Academy (RAE) carried out in the twentieth century. It will also consider the publication, during the twentieth and twenty-first centuries, of the historical dictionaries restricted to a specific geographical area (specifically those from Costa Rica, the Canary Islands and Venezuela). We then present the Nuevo diccionario histórico del español (NDHE) project. This native digital dictionary, still under development, was conceived as a lexical database to be exploited as a historical dictionary. Finally, we offer some reflections on the future of the diachronic lexicography of the Spanish language.
Numerous cybercriminals are active in the online realm, carrying out cyber-crimes according to predetermined and preplanned agendas. Cyberbullying, which was formerly limited to physical limits, has now expanded online as a result of technology advancements. One type of cyberbullying is denigration or insult. The cyberbullying cases are in exponential rise in social media as per the reports of Computer Emergency Team by Sri Lanka. Insulting words are changeable in dynamic and the same terminology may have numerous meanings depending on the context. Bullying cannot be defined just because a statement comprises such a term. As a result, when classifying comments, standard keyword detecting approaches are insufficient. Other languages also may have dealt with this issue by utilizing lexical databases like WordNet, which might give synonyms as well as homonyms for words. Because no adequate lexical database mainly for the English language has been built, recognizing a word like bullying is difficult. As a result, employed rules to solve the problem. Facebook comments containing profanity were gathered, outliers were eliminated, and the remaining messages were pre-processed. Five feature extraction rules were employed to assess insult in the text. Following that, used the Support Vector Machine (SVM) technique. Using an F1-score of 85%, the findings demonstrate that when compared to existing works, SVM performs better. The focus on English language cyberbully identification, which has never been addressed earlier, distinguishes this study.
As the most widely studied and spoken language worldwide, English is a medium of communication between speakers of varied language backgrounds. Global English users need not converge on any one variety; rather, they need adaptive, flexible skills that support communication with speakers of diverse World English (WE) varieties (Kirkpartick 2007). Nevertheless, many English learners and teachers around the world continue to adhere to ‘native speaker’1 models that prioritize language varieties from English-dominant nations, such as the United States or the United Kingdom (Tseng 2019). We have encountered this perspective first-hand in our work in Indonesia. Hanung has taught English and trained teachers at an Islamic University in Central Java for over twenty years. Tabitha has nearly twenty years of experience in ELT, including three years as a visiting instructor at Hanung’s institution. Among the students and teachers we have worked with in Central Java, the dominant language learning model is the ‘native English speaker’. For instance, on a survey given to our first-year English-major students, 81 percent reported wanting to ‘sound like a native speaker’. To challenge this tendency, Simanjuntak and Lien (2021) encourage the use of materials from global and local language communities. Using diverse listening materials can help students understand a wide variety of language varieties and accents and adapt to dialects they have never encountered before. We found podcasts, free audio recordings which are automatically delivered to users’ devices, to be a particularly rich and convenient source of listening materials featuring WE speakers. In the sections below, we describe our teaching context and use of podcasts. We close by discussing how our approach could transfer to other contexts where students and teachers continue to aspire to ‘native speaker’ models, and offer recommendations for teachers interested in shifting to a WE approach and exposing students to varied linguistic norms through podcasts. We collaborated to design lessons for ‘Listening for General Communication’, a course for first-year English majors. The students met weekly from September to December, 2021. Because of COVID-19, the course was taught asynchronously online, with the exception of in-person sessions in weeks 13 and 14. Podcasts were well suited to asynchronous online teaching because their easy accessibility facilitated students’ independent listening. Teachers can create their own podcasts (for a discussion of teacher-created podcasts, see Ingham 2022), but we found several high-quality podcasts online which featured WE speakers. Class activities revolved around seven episodes of the 22.33 podcast, which features participants in exchange programmes sponsored by the US Department of State. We used this podcast because it features speakers with diverse accents, language use, and perspectives, and focuses on intercultural encounters. With the wide variety of podcasts available online, teachers can find podcasts featuring diverse speakers which match their course objectives, contexts, and aims. In Indonesia, English is compulsory at the secondary and tertiary levels, so our students had studied English for approximately six years. Many students at Islamic universities such as ours come from secondary schools whose curriculum focuses primarily on the teaching of Islam. English instruction at these schools focuses on receptive skills, to allow students to access texts about Islam and Muslims written for international audiences. Students typically also study Arabic, but have stronger English skills due to English’s status as the first foreign language in Indonesia, and its use in mass media. The language model used in that media is most commonly standard American English. Students in our English Department take courses in language, linguistics, and teaching methods, and most enter with language skills at the B1 or B2 level. After graduation, most students gain employment as English teachers. At the beginning of the semester, we asked students to reflect upon their future use of English. Although most hoped their own language use would come to match ‘native speaker’ norms, students also acknowledged that they would probably use English to communicate with people using a variety of language norms. They expressed interest in listening to WE speakers and learning about people from around the world. We introduced the podcasts and explained that they featured WE speakers, thereby exposing students to diverse language varieties. In the first asynchronous session for each podcast, students completed pre-listening activities. Podcasts featuring WE speakers typically discussed contexts unfamiliar to our students, so these activities were essential to support comprehension. When students had some relevant prior knowledge, we activated that knowledge through reflection, opinion sharing, and hypothetical situations. When significant content was new to students, short readings, visuals, and online research helped build background knowledge. Taking the Observing Ramadan podcast as an example, students shared their own memories of observing Ramadan, considered which aspects of Ramadan might be surprising to a non-Muslim, and looked up new vocabulary, such as ‘muscle through’ and ‘famished’. Students were expected to complete while-listening activities between the two sessions. A benefit of podcasts is that students can adjust the audio speed and listen to the podcast repeatedly, if needed. To further support students’ comprehension of the WE speakers, we offered instructional scaffolds. For some lessons, students received a graphic organizer to track the various voices and topics. In other lessons, questions were divided into sections corresponding to timestamps in the podcast, helping students know when to listen intensively. For Observing Ramadan, students completed a chart with the names and nationalities of the six speakers by entering each speakers’ perspective on challenges related to fasting and non-Muslims’ awareness of Ramadan. In the second asynchronous session, students completed post-listening activities to build on and apply what they had learnt from listening. These discussions aimed to help students make personal connections and identify universal human values across the WE speakers’ varied contexts. For example, the application questions for Observing Ramadan asked students to reflect on what non-Muslims could learn from experiencing Ramadan. Students responded favourably to the use of podcasts featuring WE speakers, and felt their language skills had improved. One student explained, ‘My understanding of vocabulary and pronunciation increased, without me realizing it.’ On a survey the end of the semester, 89 percent of students said they felt their vocabulary had increased, and 76 percent reported that they appreciated the opportunity to hear good pronunciation. Students were also positively inclined toward the WE language models, which included speakers from Bangladesh, Canada, Ghana, Jordan, Lithuania, Saudi Arabia, the United States, and Yemen. Students said the speakers’ accents and vocabulary were not as difficult to understand as they had anticipated. In fact, some preferred listening to WE speakers, as the following comment reveals: ‘For me personally, listening to speakers who use English as their additional language is more easier than native speaker because we are both learning, so we have some similarities in pronunciation, or use general words.’ Students appreciated the slower speech and simpler vocabulary used by WE speakers. Students especially appreciated podcasts focused on familiar topics. Favourite episodes included Observing Ramadan, an episode about a Muslim prison chaplain in Canada, and one about a journalist who studied in Indonesia. A student explained that she preferred these episodes because ‘the experiences of the interviewees relate to our lives’. We found that the best-received podcasts were those with topics connected to students’ experiences. We see great potential in the use of podcasts featuring WE speakers in other contexts where the ‘native speaker’ model continues to hold power. We have several recommendations based on our experiences. First, we recommend orienting students to the purpose of listening to diverse WE speakers. As mentioned above, most of our students initially wished to sound like a ‘native speaker’, so it was important to offer a rationale for our seemingly contradictory WE paradigm. We believe that helping students consider the importance of engaging with a wide variety of WE speakers resulted in higher motivation and interest. Second, we suggest finding materials that include content related to students’ lives. Although our goal was to expose students to new language varieties and perspectives, we found that students were more interested when at least some of the content was familiar. They needed to be able to make personal connections with the material. When those connections do not come naturally, teachers should use teaching strategies such as guided reflection, discussion, and hypothetical scenarios to help students see how the speakers’ experiences are similar to or different from their own. Lastly, we encourage the use of materials which not only include diverse language varieties but also offer exposure to new cultures. Podcasts from diverse cultural contexts support the development of intercultural competence in addition to language competence. Given its focus on intercultural exchange, the 22.33 podcast was a rich source of such content. Other podcasts with diverse global speakers and intercultural themes include: Global Voices, Voices of Exchange, Rough Translation, The Europeans, and Sound Africa. As other practitioners use these and other podcasts to expose students to WE speakers, we hope they will also share their experiences. Last version received November 2022 Hanung Triyoko is the Head of the Language Development Unit at UIN Salatiga, Indonesia. He is an enthusiast for mutual relationships among educators across the globe and for encouraging his students and colleagues to have everyone’s unique contribution to use language as means of peace and better understanding of life and humankind. He is known as a master trainer and developer of a massive open online course (MOOC) for various levels of English teachers in Indonesia sponsored by RELO of the US Embassy. Email: [email protected] Tabitha Kidwell is a Professorial Lecturer in the TESOL programme at American University, and was previously a visiting lecturer at UIN Salatiga, Indonesia. She teaches academic writing, applied linguistics, and TESOL methods courses, and has conducted professional development for language teachers around the world. Her research focuses on innovative pedagogies, intercultural teaching approaches, and language teacher education. Email: [email protected]
Political correctness, seen as a form of linguistic interventionism, and its derivate, political incorrectness, are frequently invoked to take a stand on language. These two expressions are considered as notions whose use is increasing in the media and whose semantic content remains vague. This study focuses on a corpus of French and German metalanguage polemics on Twitter: these are positions on language taken by a speaker in the name of political (in)correctness that uses metalanguage and/or metalinguistic markers. Political correctness is considered here as a formula in order to show how its use in online exchanges makes it possible to organize the relationships between participants and contributes towards shaping linguistic norms. The analysis distinguishes, in a polemical and argumentative context, between the use of the formula to disqualify the other and as a decommitment marker that uses humor to defuse the aggressive character of utterances.
This chapter investigates adolescents’ agencies and identity constructitions within the family by zooming in on their playful and metalinguisticc language practices in family interactions. The chapter provides an overview of research on playful talk among adolescents and in multilingual families, as well as on discourse analytic approaches to parent-adolescent talk. By employing a sociolinguistic interactional analysis of family interactions, the chapter exemplifies how adolescents use playful language practices, such as teasing and mock speech, to socially position themselves and their parents. The examples show how the adolescents subvert generational hierarchies, accentuating generational differences as well as challenging conventional ideas of knowledgeable adults fostering the young. The chapter also demonstrates how interactional analyses of everyday interactions contribute to our understanding of how social and linguistic norms are negotiated in (multilingual) families.
This paper discusses the evolution of documentary culture in early medieval Tuscia by quantitatively examining the Latin spelling of charter scribes in relation to the following factors: time, the distinction between the formulaic and non‐formulaic parts of the document, the scribe’s domicile, the scribe’s professional status, and the document type. The paper asks what the spelling of charters tells us about administrative and socio‐cultural changes in charter production and in scribal education. The research data is 997 charters from the Late Latin Charter Treebank, and the approach that of philological corpus linguistics.
Foregrounding denotes a linguistic phenomenon which allows a text to stand out against the backdrop of the text that conforms with the linguistic norms. The article sheds the light on its origin, the tendency of investigations of the theory of foregrounding as well as a wide range of scientific approaches deployed in an attempt to investigate the notion “foregrounding” and enlarges on what part it plays in generating meaning in a text. In the course of the research the following conclusions were drawn: a text that comprise foregrounded language means stimulates mental work and is perceived to be intricate and thought-provoking by a reader. Foregrounding gives a rise to the creation of a special perception of the object, the creation of a ‘vision’ of it, and not ‘recognition’”. Foregrounding in a language is normally realized through the usage of stylistic devices as well as consistent and systematic nature of actualization. It enables a writer to fulfill writer’s specific aims and motivations through the text.
In literary stylistics, "deviation" is a crucial technique in poetic creation, where breaking linguistic norms can produce a foregrounding effect, thereby highlighting a poem's theme. Barry Cole's "Reported Missing" is a narrative poem abundant with deviations. It tells the story of an unsuccessful dialogue between a police officer and a man searching for his missing lover. Since the language used by the two men is disparate, their conversation ends in vain. The woman remains missing as a result of their communication breakdown. This paper aims to uncover the sources of the failure by examining the deviations in this unsuccessful dialogue, including deviation in the domain, deviation in the medium of transmission, and deviation in the tenor of discourse. These linguistic deviations represent the dramatic conflicts between the objective and subjective world, sparking reflection on the subjective beauty of one's emotions and the harsh objective reality.
This qualitative study examines the cultural and social meanings that social media users ascribe to deviations from the norms of the standard Russian language. It uses critical discourse analysis to explore trends and widely reproduced online ‘commonplaces’ in Russian-language internet users’ expressions of their attitude toward breaches of linguistic norms. It focuses on users’ arguments about whether and to what degree linguistic errors are tolerable and on how they justify their criticisms. It also discusses how Runet users assess others’ language as correct or incorrect, and how these assessments have changed over time, finding that since the mid-2010s there has been a hardening of attitudes toward linguistic deviancy, which may be explained by the increase in user numbers and consequent mainstreaming of the Internet in Russia.
Word embeddings trained on the lemmatised TOROT Treebank, using Word2Vec and the following parameters: sg = True min_count = <1,3,5> window = <3,5> vector_size = <100,200,300> epochs = 5 One model was trained for each combination of the parameters enclosed in angled brackets (< >). The release contains both the full models (.model) and the plain vector files (_vectors.txt). The models are named according to the parameters they were trained with. Note that these are the result of very preliminary experiments and no systematic evaluation of their quality was carried out, so use with caution.
Abstract This chapter provides basic data about Language Technology for the Czech language. After a brief introduction with general facts about the language (history, linguistic features, writing system, dialects), we touch upon Czech in the digital sphere. The main achievements in the field of NLP are presented: important datasets (corpora, treebanks, lexicons etc.) and tools (morphological analyzers, taggers, automatic translators, voice recognisers and generators, keyword extracters etc).
Foregrounding is a linguistic phenomenon which allows a text to stand out against the backdrop of the text that conforms with the linguistic norms. The article sheds the light on the origin of the concept of foregrounding, the tendency of investigating its theory, as well as a wide range of scientific approaches. In the course of the research the following conclusions were drawn: a text that comprise foregrounded language means stimulates mental work and is perceived to be intricate and thought-provoking by a reader. Foregrounding gives a rise to the creation of a special perception of the object. Foregrounding in a language is normally realized through the usage of stylistic devices as well as consistent and systematic nature of actualization. It enables a writer to fulfill the writer’s specific aims and motivations through the text.
This chapter deals with the morphosyntactic analysis of languages and the compilation of morphosyntactic data in digital format, with special attention to Spanish. The basic concepts of formal grammar and its typology are presented, as well as the concept of the analyser as a tool that applies formal grammars to text analysis. The computational treatment of morphosyntax includes several preprocesses, such as tokenization and the treatment of different kinds of lexical units. We distinguish between partial and full analysis, as well as different approaches to syntactic annotation: constituent analysis and the analysis of dependencies. We provide a fairly detailed presentation of the methodology used for the annotation of treebanks. We describe and analyse the basic content of the main Spanish treebanks. The approach to morphosyntax from a computational perspective also involves addressing the complexity of the analysis of real language corpora, usually extracted from the Internet, in which the language is less formal and more spontaneous. The concepts of the sentence, constituents and syntactic and semantic coherence often do not follow prescriptive rules. Finally, future lines of research are set out, concretely the definition of syntactic tagsets valid for the largest possible number of languages and the treatment of non-normative corpora and specific syntactic structures.
In this repository, we release a series of vector space models of Ancient Greek, trained following different architectures and with different hyperparameter values. Below is a breakdown of all the models released, with an indication of the training method and hyperparameters. The models are split into ‘Diachronica’ and ‘ALP’ models, according to the published paper they are associated with. [Diachronica:] Stopponi, Silvia, Nilo Pedrazzini, Saskia Peels-Matthey, Barbara McGillivray & Malvina Nissim. Forthcoming. Natural Language Processing for Ancient Greek: Design, Advantages, and Challenges of Language Models, Diachronica. [ALP:] Stopponi, Silvia, Nilo Pedrazzini, Saskia Peels-Matthey, Barbara McGillivray & Malvina Nissim. 2023. Evaluation of Distributional Semantic Models of Ancient Greek: Preliminary Results and a Road Map for Future Work. Proceedings of the Ancient Language Processing Workshop associated with the 14th International Conference on Recent Advances in Natural Language Processing (RANLP 2023). 49-58. Association for Computational Linguistics (ACL). https://doi.org/10.26615/978-954-452-087-8.2023_006 Diachronica models Training data Diorisis corpus (Vatri & McGillivray 2018). Separate models were trained for: Classical subcorpus Hellenistic subcorpus Whole corpus Models are named according to the (sub)corpus they are trained on (i.e. hel_ or hellenestic is appended to the name of the models trained on the Hellenestic subcorpus, clas_ or classical for the Classical subcorpus, full_ for the whole corpus). Models Count-based Software used: LSCDetection (Kaiser et al. 2021; https://github.com/Garrafao/LSCDetection) a. With Positive Pointwise Mutual Information applied (folder PPMI spaces). For each model, a version trained on each subcorpus after removing stopwords is also included (_stopfilt is appended to the model names). Hyperparameter values: window=5, k=1, alpha=0.75. b. With both Positive Pointwise Mutual Information and dimensionality reduction with Singular Value Decomposition applied (folder PPMI+SVD spaces). For each model, a version trained on each subcorpus after removing stopwords is also included (_stopfilt is appended to the model names). Hyperparameter values: window=5, dimensions=300, gamma=0.0. Word2Vec Software used: CADE (Bianchi et al. 2020; https://github.com/vinid/cade). a. Continuous-bag-of-words (CBOW). Hyperparameter values: size=30, siter=5, diter=5, workers=4, sg=0, ns=20. b. Skipgram with Negative Sampling (SGNS). Hyperparameter values: size=30, siter=5, diter=5, workers=4, sg=1, ns=20. Syntactic word embeddings Syntactic word embeddings were also trained on the Ancient Greek subcorpus of the PROIEL treebank (Haug & Jøhndal 2008), the Gorman treebank (Gorman 2020), the PapyGreek treebank (Vierros & Henriksson 2021), the Pedalion treebank (Keersmaekers et al. 2019), and the Ancient Greek Dependency Treebank (Bamman & Crane 2011) largely following the SuperGraph method described in Al-Ghezi & Kurimo (2020) and the Node2Vec architecture (Grover & Leskovec 2016) (see https://github.com/npedrazzini/ancientgreek-syntactic-embeddings for more details). Hyperparameter values: window=1, min_count=1. ALP models Training data Archaic, Classical, and Hellenistic portions of the Diorisis corpus (Vatri & McGillivray 2018) merged, stopwords removed according to the list made by Alessandro Vatri, available at https://figshare.com/articles/dataset/Ancient_Greek_stop_words/9724613. Models Count-based Software used: LSCDetection (Kaiser et al. 2021; https://github.com/Garrafao/LSCDetection) a. With Positive Pointwise Mutual Information applied (folder ppmi_alp). Hyperparameter values: window=5, k=1, alpha=0.75. Stopwords were removed from the training set. b. With both Positive Pointwise Mutual Information and dimensionality reduction with Singular Value Decomposition applied (folder ppmi_svd_alp). Hyperparameter values: window=5, dimensions=300, gamma=0.0. Stopwords were removed from the training set. Word2Vec Software used: Gensim library (Řehůřek and Sojka, 2010) a. Continuous-bag-of-words (CBOW). Hyperparameter values: size=30, window=5, min_count=5, negative=20, sg=0. Stopwords were removed from the training set. b. Skipgram with Negative Sampling (SGNS). Hyperparameter values: size=30, window=5, min_count=5, negative=20, sg=1. Stopwords were removed from the training set. References Al-Ghezi, Ragheb & Mikko Kurimo. 2020. Graph-based syntactic word embeddings. In Ustalov, Dmitry, Swapna Somasundaran, Alexander Panchenko, Fragkiskos D. Malliaros, Ioana Hulpuș, Peter Jansen & Abhik Jana (eds.), Proceedings of the Graph-based Methods for Natural Language Processing (TextGraphs), 72-78. Bamman, D. & Gregory Crane. 2011. The Ancient Greek and Latin dependency treebanks. In Sporleder, Caroline, Antal van den Bosch & Kalliopi Zervanou (eds.), Language Technology for Cultural Heritage. Selected Papers from the LaTeCH [Language Technology for Cultural Heritage] Workshop Series. Theory and Applications of Natural Language Processing, 79-98. Berlin, Heidelberg: Springer. Gorman, Vanessa B. 2020. Dependency treebanks of Ancient Greek prose. Journal of Open Humanities Data 6(1). Grover, Aditya & Jure Leskovec. 2016. Node2vec: scalable feature learning for networks. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD ‘16), 855-864. Haug, Dag T. T. & Marius L. Jøhndal. 2008. Creating a parallel treebank of the Old Indo-European Bible translations. In Proceedings of the Second Workshop on Language Technology for Cultural Heritage Data (LaTeCH), 27–34. Keersmaekers, Alek, Wouter Mercelis, Colin Swaelens & Toon Van Hal. 2019. Creating, enriching and valorizing treebanks of Ancient Greek. In Candito, Marie, Kilian Evang, Stephan Oepen & Djamé Seddah (eds.), Proceedings of the 18th International Workshop on Treebanks and Linguistic Theories (TLT, SyntaxFest 2019), 109-117. Kaiser, Jens, Sinan Kurtyigit, Serge Kotchourko & Dominik Schlechtweg. 2021. Effects of Pre- and Post-Processing on type-based Embeddings in Lexical Semantic Change Detection. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics. Schlechtweg, Dominik, Anna Hätty, Marco del Tredici & Sabine Schulte im Walde. 2019. A Wind of Change: Detecting and Evaluating Lexical Semantic Change across Times and Domains. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 732-746, Florence, Italy. ACL. Vatri, Alessandro & Barbara McGillivray. 2018. The Diorisis Ancient Greek Corpus: Linguistics and Literature. Research Data Journal for the Humanities and Social Sciences 3, 1, 55-65, Available From: Brill https://doi.org/10.1163/24523666-01000013 Vierros, Marja & Erik Henriksson. 2021. PapyGreek treebanks: a dataset of linguistically annotated Greek documentary papyri. Journal of Open Humanities Data 7.
Abstract Numismatic inscriptional evidence consistently employs the ΕΥΕΡΓ- word group in describing a superior providing some material public benefit to an inferior, typically an entire city, nation or kingdom. This is evidenced in the present study's comprehensive survey of several hundred numismatic types, extant in many thousands of specimens from the second century bce to the first century ce. Within this context, 1 Timothy 6.2 is discussed, wherein it is noted that the apparent identification of a slave's labour as ɛὐɛργɛσία not only heightens the significance and value of that service but is a deliberate inversion of expected social and linguistic norms.
Abstract Background Obesity and insulin resistance (IR) may negatively influence affect regulation in young adulthood. The paucity of studies that have investigated such associations have predominantly used functional magnetic resonance imaging. However, electroencephalography (EEG) can evaluate neural signatures in real‐time that may be complementary. Thus, this study investigated how adiposity and IR moderated brain activity and underlying affect regulation using EEG. Method Real‐time EEG was recorded in 28 lean and obese young adults with and without IR (18‐39 years, 46.4% female). Two event‐related potential (ERP) components of affect regulation, early posterior negativity (EPN) and late positive potential (LPP), were quantified from EEG data. Participants completed the Interactional Picture Affective System task to measure affect regulation, from which mean valence ratings and stimulus‐to‐response‐onset reaction times were calculated. ERP components and affect regulation parameters were then contrasted in three conditions, as follows: 1) Negative – Neutral; 2) Positive – Neutral; and 3) Negative – Positive. Height, weight, body fat percentage (%BF), and serum proteomics were collected. Fasting glucose and insulin readings were obtained to quantify Homeostatic Model Assessment for Insulin Resistance (HOMA‐IR). Hierarchical moderated regression analysis was utilized to test the interrelationships between adiposity, IR, neural activity, and affect regulation. Result Across all contrasted conditions, HOMA‐IR and %BF were found to moderate the relationships between frontal and parietal LPP amplitudes during the late latency window and affect regulation. Specifically, higher late frontal LPP amplitudes were associated with less negative and more positive valence ratings to unpleasant and pleasant stimuli, respectively, among participants with low, but not high, IR and %BF levels. In the Negative – Neutral and Negative – Positive conditions, lower late parietal LPP amplitudes were also associated with less negative overall valence ratings in response to unpleasant stimuli, but only in lean, insulin sensitive participants. These results overall suggest that modulation of negative and positive affectivity only occurs in lean young adults without IR. Conclusion Young adults with greater adiposity and IR showed worse affect regulation than those without obesity and IR. Furthermore, AD‐associated characteristics may start early in life, and EEG signatures may be a useful neuroimaging approach for tracking such events.
Parallelisms and features that deviate from the rules of standard language form a large part of the optional linguistic features used in poetry, but also in other forms of stylized language use, such as proverbs or advertising slogans. This preprint contains a manual with instructions for annotating poetic features that are either parallelistic or based on deviations from linguistic norms in any kind of rhymed and metered text. It serves to capture the degree to which texts make use of parallelism and deviation, which can in turn be used to predict their effect on readers and listeners.
Broad or inclusive language has reached the attention of media and public debate during the last decades. In Italy, the publication by Alma Sabatini “Raccomandazioni per un uso non sessista della lingua italiana” was an important starting point. The paper focuses on the linguistic innovation as a moment of the society/language interaction, within the theoretical framework of the law and language parallel. The opposition between descriptivism and prescriptivism in linguistics will be challenged, by interpreting the linguistic norm and the legal norm as driver for social transformations.