Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
Introduction: Effective communication often involves expressing disagreement while maintaining social harmony, which is influenced by cultural and linguistic norms. Native speakers of English typically employ various politeness strategies in their disagreement speech acts. However, Iraqi EFL learners may navigate these strategies differently due to variations in cultural norms and language proficiency. Therefore, the current study aimed to contrastively analyze the way Iraqi EFL learners and native English speakers perform the speech act of disagreement in light of politeness. Methodology: In this regard, a discourse completion test (DCT) was administered to 66 participants, comprising 33 Iraqi EFL students studying English as a foreign language (TEFL) and 33 native English speakers. The DCT was made up of scenarios that mirrored real-life circumstances in order to provoke responses from people who disagreed with them. Brown and Levinson’s (1987) theory of politeness was employed to analyze participants’ utterances. Results: The findings indicated that while expressing disagreement with people of higher, participants in both groups were more concerned with keeping their interlocutors’ positive faces. Furthermore, the study findings indicated that despite differences in the two groups of participants, Iraqi EFL learners utilized positive indirect politeness strategies more frequently than English native speakers. On the other hand, English native speakers applied direct and negative politeness strategies. Conclusion: Generally, the findings indicated that both groups tended to use the most direct type of disagreement as the social distance and power relation decreased.
Atuy Galon is a unique fanfiction that began as a series of short stories on social media, eventually gaining prominence and being published in a comic strip format. What sets this work apart is its distinctive use of slang language, an ever-evolving linguistic phenomenon that resonates particularly with younger audiences. The central figures in the narrative, Atuy, Anton, and Sahrul, not only use slang but each has developed their own distinct slang lexicon, reflecting varied facets of cultural identity. As language remains paramount in shaping one's cultural self-awareness, this research meticulously examines the linguistic choices of these characters. Employing a descriptive qualitative approach alongside a rigorous discourse analysis methodology, the study aims to decode the linguistic intricacies within Atuy Galon and their broader implications for cultural identity formation. The significance of this exploration extends beyond literary analysis; it offers a window into the dynamic interplay between language, culture, and identity among today's youth. Additionally, it underscores the transformative power of technology in reshaping linguistic norms and practices.
Purpose: The growing interdisciplinary nature of medicine has prompted increased leadership by physicians, both in administration and as informal team leaders. A formal leadership curriculum is essential to best prepare trainees for these roles.1 Physician Executive Leadership (PEL) is an organization at Thomas Jefferson University that provides medical students with leadership education by facilitating sessions led by established medical experts in executive positions. Each session is rooted in 1 or more of PEL’s leadership pillars: Applied Leadership, Quality Improvement, Health Finance, Entrepreneurship/Innovation, Health Policy, and Law/Ethics. In response to COVID-19 public health precautions, PEL sessions were shifted online. Literature is contradictory on the relative benefits and harms of online coursework in a leadership curriculum; although it increases accessibility, it may decrease student engagement.2 We therefore aimed to analyze the effectiveness of in-person versus virtual PEL curriculum sessions and to determine if certain topics were better taught in-person or virtually. Method: Six hundred eighty-four total responses (from 361 first-, 226 second-, 78 third-, and 19 fourth-year medical student submissions) were collected from 24 sessions from January 2018 to May 2022, including 13 in-person (2018–2019) and 11 virtual (2020–2022) sessions. Six Applied Leadership, 4 Quality Improvement, 8 Health Finance, 5 Entrepreneurship/Innovation, 2 Health Policy, and 3 Law/Ethics sessions were assessed. Following each session, students rated changes in their engagement and comprehension of the topic area (Likert scale 1–7: 1 = strongly disagree; 7 = strongly agree). One-way nonparametric ANOVA and unpaired t tests were used to compare responses between the 2 curricula. Results: Health Finance engagement ratings were higher online than in-person (6.49 versus 5.85; P <.001). Similarly, Health Finance (6.31 versus 5.79; P <.01) and Care Quality sessions (6.29 versus 5.73; P <.01) demonstrated improved comprehension with online delivery, while Health Policy demonstrated the opposite trend (5.72 versus 6.64; P <.01). Overall, student engagement and comprehension following the session were nonsignificantly higher virtually compared to in-person (6.33 versus 6.21; P =.08 for engagement, and 6.22 versus 6.09; P =.09 for comprehension). Comprehension and engagement scores were not significantly affected for the remaining pillars. Medical school class of participants (MS1 through MS4) did not affect ratings. Discussion: Overall ratings for both comprehension and engagement of each event remained high (> 6.0) across pillars in-person (2018–2019) and online (2020–2022). Most notably, the Health Finance sessions scored higher in comprehension and engagement in online formats. This may be explained by the ability of students to look up clarifying information during the session. Conversely, the decline in comprehension of Health Policy may be attributed to the inability to quickly find answers to complex health policy questions online. Students may therefore benefit from directly asking questions to the experts during in-person sessions. A controlled study that compares particular events both in-person and virtually would be necessary to confirm these findings. Significance: Comprehension and engagement improvements for Health Finance sessions indicate that an online format may be better suited to teach this core leadership pillar, while the decline in comprehension of Health Policy indicates this pillar may benefit from being taught in-person. Importantly, there was not a drastic decrease in engagement nor comprehension for most pillars once sessions became virtual; therefore, the increased accessibility, availability of multimedia tools, and decreased cost (for travel, lodging, etc.) suggests virtual events may be superior to their in-person counterparts for the majority of leadership curricular events.
In real time, the size of information resources in natural language is growing rapidly. The processing of these resources urgently requires the presence of linguistic databases and knowledge. Processing of information resources in natural language requires the presence of text corpora and thesauri. To create and process them, markup languages and ontological models of subject areas are required. Insufficient use of linguistic and ontological knowledge used in information retrieval and automatic text processing applications leads to various problems: irrelevant search, poor-quality categorization and referencing of documents. The existing markup languages mainly contain concepts of Romano-Germanic and Slavic language groups. These puzzles are considered burning in the field of computational linguistics. For these purposes, it is proposed to create a metalanguage and an ontological model of the grammar of the Kazakh language. Keywords: ontological model, Kazakh language, natural language, linguistic, Kazakh grammar, semantic. Қазіргі уақытта табиғи тілдегі ақпараттық ресурстардың көлемі тез өсуде. Бұл ресурстарды өңдеу жедел түрде лингвистикалық мәліметтер базасы мен білімнің болуын талап етеді. Ақпараттық ресурстарды табиғи тілде өңдеу мәтіндік корпус пен тезауристан құралады. Ақпаратты іздеу және мәтінді автоматты өңдеу қолданбаларында қолданылатын лингвистикалық және онтологиялық білімдерді жеткіліксіз пайдалану әртүрлі мәселелерге әкеледі. Қолданыстағы белгілеу тілдерінде негізінен роман-герман және славян тілдері топтарының ұғымдары бар. Оларды жасау үшін белгілеу тілдері және пәндік облыстардың онтологиялық үлгілері қажет болады. Бұл есептеуіш лингвистика саласында кеңінен таралған деп саналады. Сонымен қатар, мәтінді автоматты өңдеудің заманауи әдістеріне тіл мен әлем туралы қосымша білім көлемін енгізу күрделі мәселе болып табылады. Осы мақсатта қазақ тілі грамматикасының метатілі мен онтологиялық моделін жасау ұсынылады. Түйiн сөздер: онтологиялық модель, қазақ тілі, табиғи тіл, лингвистикалық, қазақ грамматикасы, семантикалық. В настоящее время объем информационных ресурсов на естественном языке стремительно растет. Развитие этих ресурсов требует наличия актуальной лингвистической базы данных и знаний. Обработка информационных ресурсов на естественном языке состоит из корпуса текстов и тезауруса. Недостаточное использование Абай атындағы ҚазҰПУ-нің ХАБАРШЫСЫ, «Физика-математика ғылымдары» сериясы, №3(79), 2022 лингвистических и онтологических знаний, используемых в приложениях для поиска информации и обработки текстов, приводит к различным проблемам. Существующие языки нотации в основном содержат понятия романо- германской и славянской языковых групп. Для их создания потребуются языки разметки и онтологические модели предметных областей. В то же время в современные методы автоматической обработки текстов трудно внедрить дополнительные знания о языке и мире. Для этого предлагается создать метамодель и онтологическую модель грамматики казахского языка. Ключевые слова: онтологическая модель, казахский язык, естественный язык, лингвистика, казахская грамматика, семантика.
Abstract This study aims to empirically test whether identifying as a supporter of either New South Wales (NSW) or Queensland (QLD) rugby league teams influences the extent that their respective team colors blue and maroon are associated with positively and negatively valenced words. We used a valence categorization experiment and affective rating task (valence and preference) to investigate if team affiliation and shared ingroup experience influenced affective associations with team colors. NSW supporters were faster and more accurate when categorizing positive words presented in blue than maroon font and negative words in maroon than blue font. While QLD supporters did not significantly differ when categorizing words in either blue or maroon, they rated blue and maroon equally positively in contrast to the NSW supporters. Results from this study give us greater insights into how color‐valence associations can be formed through subcultural ingroup affiliations.
Previous research on linguistic relativity and economic decisions hypothesized that speakers of languages with obligatory tense marking of future time reference (FTR) should value future rewards less than speakers of languages which permit present tense FTR. This was hypothesized on the basis of obligatory linguistic marking (e.g., will) causing speakers to construe future events as more temporally distal and thereby to exhibit increased "temporal discounting": the subjective devaluation of outcomes as the delay until they will occur increases. However, several aspects of this hypothesis are incomplete. First, it overlooks the role of "modal" FTR structures which encode notions about the likelihood of future outcomes (e.g., might). This may influence "probability discounting": the subjective devaluation of outcomes as the probability of their occurrence decreases. Second, the extent to which linguistic structures are subjectively related to temporal or probability discounting differences is currently unknown. To address these, we elicited FTR language and subjective ratings of temporal distance and probability from speakers of English, which exhibits strongly grammaticized FTR, and Dutch, which does not. Several findings went against the predictions of the previous hypothesis: Framing an FTR statement in the present ("Ellie arrives later on") versus the future tense ("…will arrive…") did not affect ratings of temporal distance; English speakers rated future statements as relatively more temporally proximal than Dutch speakers; and English and Dutch speakers rated future tenses as encoding high certainty, which suggests that obligatory future tense marking might result in less discounting. Additionally, compared with Dutch speakers, English speakers used more low-certainty terms in general (e.g., may) and as a function of various experimental factors. We conclude that the prior cross-linguistic observations of the link between FTR and psychological discounting may be caused by the connection between low-certainty modal structures and probability discounting, rather than future tense and temporality.
Background: Type 2 diabetes mellitus (DM) is the most common metabolic disorder in the world and an important risk factor for peripheral arterial disease (PAD). CT angiography represents the method of choice for the diagnosis, pre-operative planning, and follow-up of vascular disease. Low-energy dual-energy CT (DECT) virtual mono-energetic imaging (VMI) has been shown to improve image contrast, iodine signal, and may also lead to a reduction in contrast medium dose. In recent years, VMI has been improved with the use of a new algorithm called VMI+, able to obtain the best image contrast with the least possible image noise in low-keV reconstructions. Purpose: To evaluate the impact of VMI+ DECT reconstructions on quantitative and qualitative image quality in the evaluation of the lower extremity runoff. Materials and Methods: We evaluated DECT angiography of lower extremities in patients suffering from diabetes who had undergone clinically indicated DECT examinations between January 2018 and January 2023. Images were reconstructed with standard linear blending (F_0.5) and low VMI+ series were generated from 40 to 100 keV, in an interval of 15 keV. Vascular attenuation, image noise, signal-to-noise ratio (SNR), and contrast-to-noise ratio (CNR) were calculated for objective analysis. Subjective analysis was performed using five-point scales to evaluate image quality, image noise, and diagnostic assessability of vessel contrast. Results: Our final study cohort consisted of 77 patients (41 males). Attenuation values, CNR, and SNR were higher in 40-keV VMI+ reconstructions compared to the remaining VMI+ and standard F_0.5 series (HU: 1180.41 ± 45.09; SNR: 29.91 ± 0.99; CNR: 28.60 ± 1.03 vs. HU 251.32 ± 7.13; SNR: 13.22 ± 0.44; CNR: 10.57 ± 0.39 in standard F_0.5 series) (p < 0.0001). Subjective image rating was significantly higher in 55-keV VMI+ images compared to the other VMI+ and standard F_0.5 series in terms of image quality (mean score: 4.77), image noise (mean score: 4.39), and assessability of vessel contrast (mean value: 4.57) (p < 0.001). Conclusions: DECT 40-keV and 55-keV VMI+ showed the highest objective and subjective parameters of image quality, respectively. These specific energy levels for VMI+ reconstructions could be recommended in clinical practice, providing high-quality images with greater diagnostic suitability for the evaluation of lower extremity runoff, and potentially needing a lower amount of contrast medium, which is particularly advantageous for diabetic patients.
Languages are known to describe the world in diverse ways. Across lexicons, diversity is pervasive, appearing through phenomena such as lexical gaps and untranslatability. However, in computational resources, such as multilingual lexical databases, diversity is hardly ever represented. In this paper, we introduce a method to enrich computational lexicons with content relating to linguistic diversity. The method is verified through two large-scale case studies on kinship terminology, a domain known to be diverse across languages and cultures: one case study deals with seven Arabic dialects, while the other one with three Indonesian languages. Our results, made available as browseable and downloadable computational resources, extend prior linguistics research on kinship terminology, and provide insight into the extent of diversity even within linguistically and culturally close communities.
Human affect recognition has been a significant topic in psychophysics and computer vision. However, the currently published datasets have many limitations. For example, most datasets contain frames that contain only information about facial expressions. Due to the limitations of previous datasets, it is very hard to either understand the mechanisms for affect recognition of humans or generalize well on common cases for computer vision models trained on those datasets. In this work, we introduce a brand new large dataset, the Video-based Emotion and Affect Tracking in Context Dataset (VEATIC), that can conquer the limitations of the previous datasets. VEATIC has 124 video clips from Hollywood movies, documentaries, and home videos with continuous valence and arousal ratings of each frame via real-time annotation. Along with the dataset, we propose a new computer vision task to infer the affect of the selected character via both context and character information in each video frame. Additionally, we propose a simple model to benchmark this new computer vision task. We also compare the performance of the pretrained model using our dataset with other similar datasets. Experiments show the competing results of our pretrained model via VEATIC, indicating the generalizability of VEATIC. Our dataset is available at https://veatic.github.io.
Abstract Stress, anxiety, and depressive symptoms can be reduced by listening to music, but the underlying mechanisms remain unclear. To address this gap, we measured brain connectivity while participants listened to songs of different genres: ambient, pop, and metal. Additionally, affective ratings were obtained while participants ( n = 30) listened to the six different songs, and subjective ratings of state anxiety were solicited at the terminus of each song. Electroencephalography (EEG) connectivity combining weighted Phase Lag Index and graph theory was utilised to document brain activity during listening. Repeated-measures ANOVA indicated that listening to more pleasant and less arousing songs was associated with lower self-reported state anxiety levels than songs rated unpleasant and highly arousing. Of interest, EEG alpha connectivity differed across two ambient songs, particularly in the frontal lobes, despite being from the same genre and rated as highly pleasant and low in arousal. We also observed a sex effect on EEG results, where female participants ( n = 18) displayed stronger connectivity than male participants ( n = 12). Combined, these results suggest that ambient songs reduce state anxiety but have divergent brain responses, possibly reflecting the complex nature of music listening, including sensory processing, emotion and cognition.
The paper traces the dynamics of the interpretation of the grammatical nature of the vocative in Ukrainian grammars from the 16th century until the present. The subject of the analysis is the content and presentation of this category in two sections of Ukrainian grammar books: (1) morphological, which clarifies the status of the vocative in the inflectional paradigm of the noun, and (2) syntactic, in which the means of expressing address are characterized. Based on the findings of the research, various trends in the description of the vocative in different historical periods have been identified, in particular: (1) until the beginning of the 20th, it was unequivocally qualified as an equal member of the inflectional paradigm of the noun, equal to other cases; (2) from the beginning of the 20th century to 1933 was a period of competition between two theories (the vocative is a case the same as others or the vocative is not a true case, but a “special” form in the inflectional paradigm of the noun); (3) the canonization of the “fake case” status theory; (4) from 1991 to the present there has been an unanimity of authors in qualifying the vocative as a case. Comparing the stages of fundamental changes in the scientific definition of the vocative in grammars with defining events in the history of Ukraine provides the basis for discussions about the influence of socio–political factors on the representation of linguistic theories and the codification of the linguistic norms.
The Wall Street Journal section of the Penn Treebank has been the de-facto standard for evaluating POS taggers for a long time, and accuracies over 97\% have been reported. However, less is known about out-of-domain tagger performance, especially with fine-grained label sets. Using data from Elder Scrolls Fandom, a wiki about the \textit{Elder Scrolls} video game universe, we create a modest dataset for qualitatively evaluating the cross-domain performance of two POS taggers: the Stanford tagger (Toutanova et al. 2003) and Bilty (Plank et al. 2016), both trained on WSJ. Our analyses show that performance on tokens seen during training is almost as good as in-domain performance, but accuracy on unknown tokens decreases from 90.37% to 78.37% (Stanford) and 87.84\% to 80.41\% (Bilty) across domains. Both taggers struggle with proper nouns and inconsistent capitalization.
Numerous studies have been conducted on the interpretation and translation of English terms into other languages. The purpose of this study was to identify the adequate Indonesian equivalent terminology for hotel amenities, services, and facilities applied in English and the strategies utilized by both domestic and international hotel guests in understanding the equivalent terms in their native language. Qualitative research methodology was used. The subjects included 10 domestic guests from a 5-star hotel, 10 domestic guests from a 4-star hotel, 5 international guests from a 3-star hotel, and 2 hotel staff from a 5-star hotel, 3 staff from a 4-star hotel, and 1 staff from a 3-star hotel. The findings demonstrated that some of the English terms commonly used in hotels had Indonesian equivalents, and some did not. The international guests strategies were: 1) searching in an online dictionary or a Google search; 2) asking people they met nearby immediately; and 3) guessing the meaning. Domestic guests’ strategies included: (a) asking other guests or hotel staff for clarification; and (b) guessing the meaning. Future research should overcome the limitations of this study, considering translations and linguistic norms training strategies.
Abstract We propose a Slovak language model for the spaCy library in Python. These models are easy-to-use for basic natural language processing tasks in a single package. The package contains several components for basic preprocessing tasks, such as tokenization, sentence boundary detection, syntactic parsing, lemmatization, named entity recognition, morphology analysis, and word vectors. It is based on the state-of-the-art monolingual SlovakBERT model. Named entity recognition is trained on a separate, publicly available WikiAnn database. The other statistical classifiers use a Slovak Dependency Treebank corpus. Morphological tags are compatible with the conventions of the Slovak National Corpus. The part of speech tags use conventions of the Universal Dependencies framework. We trained a separate word vector model on a web-based corpus. The training uses fastText with Floret modification. We present a series of experiments that confirm that the model performs similarly to other languages for all tasks. Training scripts and data are publicly available.
В статье анализируется историческая динамика политической корректности, ее положительные и отрицательные стороны, а также прослеживается ее связь с вежливостью, которая заключается в неиспользовании конфронтационных, ликоповреждающих коммуникативных стратегий при указании на гендер, расу, этнос, возраст, физическое состояние и социальное положение адресата. Исследуются факторы, затрудняющие формирование политкорректного русского языкового сознания: 1) не вполне сформировавшееся понятие ПК применительно к отечественному социальному контексту, осложненное его концептуализацией через призму западного восприятия; 2) отсутствие правовых механизмов ПК, несмотря на недопустимость дискриминации, закрепленную в Конституции РФ; 3) неразработанность языковых основ применения ПК в российском публичном дискурсе; 4) англоцентричность правил ПК для международного общения. Сделан вывод о необходимости выработки российских норм политической корректности с участием широкого лингвистического сообщества. The paper explores the historical dynamics of political correctness (PC), its positive and negative aspects, as well as its connection with politeness as an avoidance of confrontational, face-threatening strategies in reference to the interlocutor’s gender, race, ethnic background, age, physical condition and social status. The study also deals with the factors hindering the formation of the Russian PC awareness, which include: 1) the incompletely formed notion of political correctness in the Russian social context complicated by its conceptualization through the prism of Western comprehension; 2) absence of legal PC mechanisms, in spite of non-discrimination enshrined in the RF Constitution; 3) underdeveloped linguistic norms of political correctness in Russian public discourse; 4) anglocentrism of PC rules in intercultural communication. The article concludes by proposing a wide discussion of Russian political correctness norms involving a wider linguistic community.
Introduction A study was conducted to investigate if an individual’s trust in law enforcement affects their perception of the emotional facial expressions displayed by police officers. Methods The study invited 77 participants to rate the valence of 360 face images. Images featured individuals without headgear (condition 1), or with a baseball cap (condition 2) or police hat (condition 3) digitally added to the original photograph. The images were balanced across sex, race/ethnicity (Asian, African American, Latine, and Caucasian), and facial expression (Happy, Neutral, and Angry). After rating the facial expressions, respondents completed a survey about their attitudes toward the police. Results The results showed that, on average, valence ratings for “Angry” faces were similar across all experimental conditions. However, a closer examination revealed that faces with police hats were perceived as angrier compared to the control conditions (those with no hat and those with a baseball cap) by individuals who held negative views of the police. Conversely, participants with positive attitudes toward the police perceived faces with police hats as less angry compared to the control condition. This correlation was highly significant for angry faces ( p &lt; 0.01), and stronger in response to male faces compared to female faces but was not significant for neutral or happy faces. Discussion The study emphasizes the substantial role of attitudes in shaping social perception, particularly within the context of law enforcement.
- output-{ciep,treebanks}-full.csv: frequency and entropy for all the categories, using four types of combinations of layers;<br> - plots.R: R script to draw plots from the output files;<br> - readReport-{CIEP+,treebanks}.R: R script to extract frequency and compute entropy from the report files (not included);<br> - ud-wordorder.py: Python script to extract word order pairs from conllu files and write them in report files. Unfortunately, I cannot include the report files, as CIEP+ is protected by copyright; the analysis can be however replicated with respect to the UD Treebanks.
Fear overgeneralization and perceived uncertainty about future outcomes have been suggested as risk factors for clinical anxiety. However, little is known regarding how they influence each other. In this study, we investigated whether different levels of threat uncertainty influence fear generalization. Three groups of healthy participants underwent a differential fear conditioning protocol followed by a generalization test. All groups learned to associate one female face (conditioned stimulus, CS+) with a female scream (unconditioned stimulus, US) while the other face (CS-) was not associated with the scream. In order to manipulate threat uncertainty, one group (low uncertainty, n = 26) received 80%, the second group (moderate uncertainty, n = 32) received 60%, and the third group (high uncertainty, n = 30) 40% CS-US contingency. In the generalization test, all groups saw CS+ and CS- again as well as four morphs that varied in similarity with the CS+ in steps of 20%. Subjective (expectancy, valence, and arousal ratings), psychophysiological (skin conductance response, SCR), and visuocortical (steady-state visual evoked potentials, ssVEPs) indices of fear were registered. Participants expected the US in accordance with their reinforcement schedules but displayed stronger skin conductance with more uncertainty. However, acquisition of conditioned fear was not evident in ssVEPs. During the generalization test, we found no effect of threat uncertainty in any of the measured variables, but the strength of generalization for threat expectancy ratings was positively correlated with dispositional intolerance of uncertainty. This study suggests that mere threat uncertainty does not modulate fear generalization.
User ratings are widely used in web systems and applications to provide personalized interaction and to help other users make better choices. Previous research has shown that rating scale features and user personality can both influence users' rating behaviour, but relatively little work has been devoted to understanding if the effects of rating scale features may vary depending on users' personality. In this paper, we study the impact of scale granularity and colour on the ratings of individuals with different personalities, represented according to the Big Five model. To this aim, we carried out a user study with 203 participants, in the context of a web-based survey where users were assigned an image rating task. Our results confirm that both colour and granularity can affect user ratings, but their specific effects also depend on user scores for certain personality traits, in particular agreeableness, openness to experience and conscientiousness.
Reviewed by: La norme du français et sa diffusion dans l’histoire ed. by Dorothée Aquino-Weber, Sara Cotelli Kureth and Carine Skupien Dekens Bryan Donaldson Aquino-Weber, Dorothée, Sara Cotelli Kureth, and Carine Skupien Dekens, eds. La norme du français et sa diffusion dans l’histoire. Honoré Champion, 2021. ISBN 978-2-7453-5626-0. Pp. 250. Linguistic norms, a perennial topic in French, receive a historical perspective in this edited volume. Aquino-Weber and Cotelli Kureth’s introduction reminds readers of Haugen’s seminal model and the distinction between linguistic norms (actual usage) and prescriptive norms (desired usage). Skupien Dekens traces the diffusion of prescriptive norms in the teaching of French as a foreign language. Kristol discusses François Bonivard, a 16th-century Swiss polyglot whose remarques prioritize objective norms over prescriptive norms, more reminiscent of a linguist than a grammarian. Amatuzzi examines how three 17th-century grammarians (Du Val, Chiflet, De Courtin) view variation and changing norms. All three have a globally negative view of variation, but Chiflet appears tolerant of stylistic variation that characterizes familiar speech. Grosse discusses the 18th-century educator (translator, historian, literary critic...) Eléazar de Mauvillon; this émigré from Provence gave French lessons in Dresden, where his pupils’ interlanguage (“vitement”, “il est vingt ans”) still resonates today. Cotelli Kureth and Nissille examine the teaching of (standard) French in Switzerland in the 18th and 19th centuries, when the home language was often a patois or regional French. Interestingly, some prescriptive manuals defended regional variants (gringe “grumpy”), either on account of their meritorious etymology or the génie of the local variety. Surcouf addresses the challenges of describing spoken language through writing and asks if (educated, professional) linguists can accurately and objectively describe the speech of less educated speakers. Surcouf notes the implicit comparison to the (conservative) written norm when we describe, for example, ne “deletion” in spoken French. He proposes, instead, that we speak of adding ne for negation in writing. Glikman and Bouard research the conjunctive phrase pour que, which, despite widespread normative reprobation (e.g., by Vaugelas), replaced afin que as the dominant expression of purpose. Capin examines historical metalinguistic commentary on subject pronoun expression. Caron focuses on the choice between passé simple ~ passé composé and social factors like the “crushing” of the linguistically conservative Parlement de Paris that may have favored the rise of the latter. Laferrière examines tensions between prescriptive norms and usage of incises de citation; after early condemnation of examples like reconnus-je, decried by grammarians as “horreurs” or “infamies,” Laferrière documents evolutions to the norm that encourage variety beyond the [End Page 257] “lassant” and formulaic dit-il. Sthioul examines treatments of the elusive passé surcomposé, identified as early as Meigret (1550) but often marginalized, stigmatized, or simply ignored; Sthioul’s hope is that the current openness to variation leads to more acceptance and understanding of this form. This volume contributes valuable scholarship on norms in French, despite its heterogeneity and occasional unevenness. Strengths include the inclusion of figures who observed and commented on French, beyond the usual suspects like Vaugelas or Bouhours and the emphasis on regional Frenches (especially in Switzerland), including local norms. The contributors consistently highlight the subjectivity of prescriptive norms and the shifting rationales for them (usage, raison, Dieu...) and, chemin faisant, unearth historical sociolinguistic morsels. Surcouf’s chapter, in particular, should be required reading for any linguist, en herbe or otherwise, working on spoken French. [End Page 258] Bryan Donaldson University of California, Santa Cruz Copyright © 2023 American Association of Teachers of French
The goal of this contribution is to present The Digital Rosetta Stone, which is a project developed at Leipzig University by the Chair of Digital Humanities and the Egyptological Institute/Egyptian Museum Georg Steindorff in collaboration with the British Museum and the Digital Epigraphy and Archaeology Project at the University of Florida. The aims of the project are to produce a collaborative digital edition of the Rosetta Stone, address standardization and customization issues for the scholarly community, create data that can be used by students to understand the language and content of the document, and produce a high-resolution 3D model of the stone. First, the three versions of the text were transcribed and encoded in XML according to the EpiDoc guidelines. Next, the versions were aligned with the Ugarit iAligner tool that supports the alignment of ancient texts with modern languages, such as English and German. All three texts were then parsed syntactically and morphologically through Treebank annotation. Finally, the project explored new 3D-digitization techniques of the Rosetta Stone in the British Museum in order to enhance traditional archaeological methods and facilitate the study of the artifact. The results of this work were used in different courses in Digital Humanities, Digital Philology, and Egyptology.
This paper attempts to examine 'World Englishes' (WE) with connectivity to English as an International Language (EIL), Applied Linguistics and socio-linguistics. In the light of Kachru's model of English Language in the late 20th century. This model has three circles, inner circle, where English is used as native language, Outer Circle, mostly former colonies of British Empire, such as Singapore, India, Kenya, Ghana, Malaysia, Pakistan and others, and 3rd is Expanding Circle, include countries in which English is known as Foreign Language in schools and universities, mostly for communication and business or economic purposes as well with Inner and Outer circles. The term "English language" refers to various interesting and notable features, patterns, or aspects of the English language. These phenomena can encompass a wide range of linguistic phenomena, including grammar, vocabulary, pronunciation, syntax, idioms, and more. English holds significant importance around the world because English is the most widely spoken language globally. It serves as a common language of communication among people from different linguistic backgrounds. Proficiency in English enables individuals to connect with a broader range of people, both in personal and professional contexts. English is the language of international business and economics as well. It facilitates global trade, negotiations, and collaboration between companies and individuals from different countries. Proficiency in English enhances employability and career opportunities, particularly in multinational corporations and industries with international reach. It recognizes the importance of both native and non-native varieties of English and acknowledges that each circle has its own linguistic norms, purposes, and language development. The study informs us that Kachru was an original thinker not in the field of English Language including applied linguistics, multilingualism, bilingualism, language policy, language creativity, code mixing, code switching, cross-cultural communication, sociolinguistics but also in the domain of politics of language and so many other issues including cross-cultural awareness.
PURPOSE: Slow speech rate and abnormal temporal prosody are primary diagnostic criteria for differentiating between people with aphasia who do and do not have apraxia of speech. We sought to identify appropriate cutoff values for abnormal word syllable duration (WSD) in a word repetition task, interpret them relative to a data set of people with chronic aphasia, and evaluate the extent to which manually derived measures could be approximated through an automated process that relied on commercial speech recognition technology. METHOD: Fifty neurotypical participants produced 49 multisyllabic words during a repetition task. Audio recordings were submitted to an automated speech recognition (ASR) service (IBM Watson) to measure word duration and generate an orthographic transcription. The transcribed words were compared to a lexical database, and the number of syllables was identified. Automatic and manual measures were compared for 50% of the sample. Results were interpreted relative to WSD scores from an existing data set of 195 people with mostly chronic aphasia. RESULTS: ASR correctly identified 83% of target words and 98% of target syllable counts. Automated word duration calculations were longer than manual measures due to imprecise cursor placement. Upon applying regression coefficients to the automated measures and examining the frequency distributions for both manual and estimated measures, a WSD of 303-316 ms was found to indicate longer-than-normal performance (corresponding to the 95th percentile). With this cutoff, 40%-45% of participants with aphasia in our comparison sample had an abnormally long WSD. CONCLUSIONS: We recommend using a rounded WSD cutoff score between 303 and 316 ms for manual measures. Future research will focus on customizing automated WSD methods to speech samples from people with aphasia, identifying target words that maximize production and measurement reliability, and developing WSD standard scores based on a large participant sample with and without aphasia.
Embodied cognition research identifies mechanisms by which our cognitive activity is connected to body experiences. This approach encompasses not only experimental manipulations but also the quantification of variables related to group and individual differences, i.e., participant-related variables. Moreover, stimuli-related characteristics, such as sensorimotor word ratings, can either be used for the selection of experimental materials or can be the main output of a study themselves. This quantitative information about individuals or stimuli can be collected through non-experimental methods, such as questionnaires and cognitive tests. This chapter gives an overview of questionnaires and cognitive tests often used in embodied cognition research. A questionnaire is a list of questions asking participants to provide information on certain aspects, such as their sociodemographic or medical status. A test is a series of tasks which participants perform for further evaluation by researchers, such as tests of mathematical ability, reading speed, or counting direction. Rating studies collect subjective evaluations of various parameters, typically for large sets of items. The present chapter is divided into two main sections: Participant-related variables and stimuli-related characteristics. We present examples from cognitive linguistics, psycholinguistics, psychophysics, as well as from research on numerical cognition, peripersonal space, and attitudes towards social robots.
In this study, we collected affective ratings of emotional valence and arousal for 882 Serbian words and compared their values at three points in time: before the onset of the COVID-19 pandemic (2018), during the COVID-19 lockdown (2020) and after the government measures were abandoned (2022). Although valence ratings were more stable than arousal ratings, we did not observe a significant change in either valence or arousal ratings across the time points. A more detailed look into the data revealed the change in arousal that was different across the valence values. Our analyses demonstrated that, upon the onset of the COVID-19 pandemic, emotionally negative words elicited higher arousal ratings, whereas emotionally positive words elicited lower arousal ratings. It revealed that our participants became more sensitive to the negative content and less sensitive to the positive content. We hypothesized that this pattern could be linked to reduced resilience and consequently could represent a mental health risk. (This article is published in Applied Psycholinguistics. Popović Stijačić, M., Mišić, K., &amp; Filipović Đurđević, D. (2023). Flattening the curve: COVID-19 induced a decrease in arousal for positive and an increase in arousal for negative words. Applied Psycholinguistics, 44(6), 1069–1089. doi:10.1017/S0142716423000425)
The circumplex model posits a circular representation of affect and some personality traits. There is an increasing need to examine the viability of the circumplex model with multivariate time series data collected on the same individuals due to the development of new data collection methods such as smartphone applications and wearable sensors. Estimating the circumplex model with time series data is more complex than with cross-sectional data because scores at nearby time points tend to be correlated. We adapt Browne’s circumplex model to accommodate time series data. We illustrate the proposed method with an empirical data set of daily affect ratings of an individual over 70 days. We conducted a simulation study to explore the statistical properties of the proposed method. The results show that the method provides more satisfactory confidence intervals and test statistics than a method that treats time series data as if they were cross-sectional data.
This corpus-based study investigates the distributions of Korean multiple anaphors, with respect to their morphological types and discourse-pragmatic properties in Long-Distance (LD)-binding. The study is based on the theory of Long-Distance Anaphors (LDAs), such as ‘form-function correlation’ argument (Cole, Hermon, & Sung, 1990; Reinhart & Reuland, 1993; Reuland, 2011, 2017), as well as exempt binding and logophoricity (Sells, 1987; Pollard & Sag, 1992). Based on Sejong Treebank (Parsed Corpus), six hundred sentences containing various Korean anaphors went through manual coding of 5 linguistic factors related to LD-binding: locality, discourse, exempt, logophoric, and logophoric roles. The encoded sentences in distinct conditions were analyzed in terms of frequency by using Chi-square tests. The overall results demonstrated the following: 1) Korean anaphors did not show form-function correlations in terms of binding type and morphological form. 2) Korean anaphors can occur even in syntactically non-exempt positions. 3) Logophoricity conditions were found with the LD-antecedents of the anaphors. The results seem to support discourse-pragmatic analysis with Korean anaphors.
Russian constructicon is an open-access linguistic database containing detailed descriptions of over 3,800 Russian grammatical constructions. In this paper we present a new, enlarged and updated version of Russian Constructicon (RusCxn) as well as new trajectories of development which were opened for the resource after the update. Since its first release, RusCxn, has undergone many significant changes. Our team has expanded the number of constructions present in the database 1,5 times, introduced new meta-information features such as glosses, significantly reworked the architecture and the design of Russian Constructicon’s website, and improved the search facilities. The above-mentioned changes not only make RusCxn more attractive and convenient-to-use, but they can also greatly facilitate typological research in the field of Construction Grammar and improve the mapping between constructicography-orinented resources for different languages.
Previous studies have made great advances in RST discourse parsing through specific neural frameworks or features, but they usually split the parsing process into two subtasks and heavily depended on gold discourse segmentation. In this article, we introduce an end-to-end method for sentence-level RST discourse parsing via transforming it into a text-to-text generation task, which can also be simply applied to document-level parsing. Our method unifies the traditional two-stage parsing and generates the parsing tree directly from the input text through our constrained decoding and postprocessing algorithms, without requiring a complicated model. Moreover, the discourse segmentation can be simultaneously generated and extracted from the parsing tree. Experimental results on the RST Discourse Treebank demonstrate that our proposed method outperforms existing methods in both the tasks of discourse parsing and segmentation. We further carry out ablation studies and more targeted comparisons with traditional patterns to analyze our method in more detail. Considering the lack of annotated data in RST parsing, we also create high-quality augmented data and implement self-training, which further improves the performance of our method.
Decoding Part of Speech(POS) tagging directly from electroencephalography (EEG) signals whilst user overtly spoke (voiced speech) sentences could improve direct speech brain-computer interfaces (BCIs) using imagined or inner speech. To the best of our knowledge, earlier work uses machine learning approach using 74,953 sentences/tokens recorded in 75 EEG sessions. The tokens can be found in 4,479 phrases consisting of terms from the English Online treebank which contains the record of weblogs, newsgroups, reviews, and Yahoo Answers. The results demonstrated the feasibility of POS decoding from EEG based on word class, word frequency, and word length with accuracy of 71%, 86%, 89%, respectively. We believe that there is significant room for improvement with more advanced artificial intelligence. In this paper, we further extend the existing work with end-to-end transformers. Our results presents transformer model outperforms benchmark traditional ML results with +20% in length, +13% for the open vs closed class and +12% in frequency. In our empirical analysis, we find the decoding performance was better when using multi-electrode recordings as compared to single-electrode recordings.
We present an approach for assessing how multilingual large language models (LLMs) learn syntax in terms of multi-formalism syntactic structures. We aim to recover constituent and dependency structures by casting parsing as sequence labeling. To do so, we select a few LLMs and study them on 13 diverse UD treebanks for dependency parsing and 10 treebanks for constituent parsing. Our results show that: (i) the framework is consistent across encodings, (ii) pre-trained word vectors do not favor constituency representations of syntax over dependencies, (iii) sub-word tokenization is needed to represent syntax, in contrast to character-based models, and (iv) occurrence of a language in the pretraining data is more important than the amount of task data when recovering syntax from the word vectors.
The purpose of the article is to consider the morphological peculiarities of the system of the nouns in New Bulgarian translation of the “Catechismos” written by Theodore the Studite, which is a part of the manuscripts no. 1/154 kept in Odessa National Scientific Library. The subject of the research is the morphological specifics of nouns in Odessa copy of the “Catechismos” dating from the 18th. The morphological peculiarities of nouns is considered in the context of the formation of a linguistic norm, which allowed the combination with different intensity of linguistic means of several language systems functioning at the time (traditional Middle Bulgarian written language, Church Slavonic Eastern recensions, and vernacular language form). The analysis proposed in this paper presents the extensive system of cases, which does not reflect the real vernacular Bulgarian speech in the 18th; the specifics of the functioning of the gramemes of the case paradigm of masculine, feminine and neuter nouns in the singular and plural forms is analyzed. Usage case endings mistakes, which indicates their artificial nature, are considered. The lack of article of nominal parts of speech is noted; the predominance of compound declension form of the adjectives and participles over short forms is revealed; the relatively high frequency of use of active present participles is registered. The results of the study make it possible to outline some probable factors that determine the writer’s preference for using the linguistic tools of the so-called “bookish”, “traditional”, “archaic” writing systems. An another reason which to some extent explains he usage of case inflections in the text of this relatively late stage of the historical development of the Bulgarian language might be the use of East Slavic copies of the Studite’s sermons by the scriber. Еhe comparison of “Catechismos” copies of South and East Slavic origin is necessary for verification of this assumption, in which we see prospects for further research.
The onset of the COVID-19 pandemic accentuated the need for access to biomedical literature to answer timely and disease-specific questions. During the early days of the pandemic, one of the biggest challenges we faced was the lack of peer-reviewed biomedical articles on COVID-19 that could be used to train machine learning models for question answering (QA). In this paper, we explore the roles weak supervision and data augmentation play in training deep neural network QA models. First, we investigate whether labels generated automatically from the structured abstracts of scholarly papers using an information retrieval algorithm, BM25, provide a weak supervision signal to train an extractive QA model. We also curate new QA pairs using information retrieval techniques, guided by the clinicaltrials.gov schema and the structured abstracts of articles, in the absence of annotated data from biomedical domain experts. Furthermore, we explore augmenting the training data of a deep neural network model with linguistic features from external sources such as lexical databases to account for variations in word morphology and meaning. To better utilize our training data, we apply curriculum learning to domain adaptation, fine-tuning our QA model in stages based on characteristics of the QA pairs. We evaluate our methods in the context of QA models at the core of a system to answer questions about COVID-19.
This chapter synthesizes the most prominent Natural Language Processing (NLP) studies conducted on Persian, focusing on text processing. The first section contains selected tasks from the NLP pipeline, such as text preprocessing, tokenization, POS tagging, syntactic parsing, treebank annotation or semantic analysis along with examples of how researchers approached the problem for Persian and, where applicable, examples of tools developed to perform given tasks. The following section discusses the application of Persian NLP like spell-checking, information retrieval, machine translation or sentiment analysis. Finally, the last section summarizes the Persian NLP corpora and other resources.
This paper provides a comparative analysis of the patterns of formation and qualities of a modern linear text and an Internet text. The article is the result of the study of Internet text stylistics, based mainly on Russian-language texts from Russia and Ukraine. The paper considers the process of formation of a new type of text - the Internet text. Being essentially different from the classical linear text, the Internet text does not lend itself well to the description based on the classical text theory. Thus, the Internet text is not complete, vectorial, not united by the completeness of thought expression. An important characteristic of the Internet text is its interactivity, which means that the roles of addressee and addressee are constantly changing. In addition, it is difficult, or even impossible, to define the boundaries of the Internet text due to its hypertextuality, which has become habitual intertextuality. All the above-mentioned aspects make up the pragmatics of the Internet text as a subspecies of the media text. Another crucial problem addressed in the paper is the study of the regularities of Internet communication in general and the stylistics of the Internet text. In the course of the research it became obvious that speech aggression and violation of norms of speech culture are stylistic dominants of online communication. This influenced the formation of other stylistic dominants such as, hate speech, fake, hype, clickbait, etc. Internet style is clearly characterized by being provocative, aggressive, hostile. The problems of bullying, humiliation of human dignity, invective and obscenity are actively studied from the standpoint of linguoecology, because, according to most researchers, the constant neglect of communicative and linguistic norms leads to the degradation of the national language style and literary norms. Despite the fact that there is still a division between public and interpersonal online communication, i.e. formal and informal, the problems of speech behavior of Internet users are becoming increasingly relevant.
In the investigation of musical features that influence musical affect, timbre has received relatively little attention. We studied the acoustic properties describing the timbral qualities of sound and analyzed how they predict perceived and induced affect. First, we considered the timbre of single tones played by different instruments by re-analyzing two previously published studies on perceived affect by Eerola et al. (2012, Mus. Percept.) and McAdams et al. (2017, Front. Psychol.) and comparing them to Experiment 1 from Korsmit et al. (2023, submitted), which investigated both perceived and induced affect on valence, tension arousal, and energy arousal ratings. For all datasets, positive valence and decreased tension were predicted by an increasingly prominent fundamental frequency. In experiments with pitch variation, energy arousal was predicted by increased pitch and decreased inharmonicity. In experiments with variations in playing technique, energy arousal was predicted by a faster attack or less sustain. Second, we compared Experiment 1 results on dimensional affect (valence, tension, energy) to results on discrete affect (anger, fear, sadness, happiness, tenderness). Like valence and tension, angrier and more fearful sounds and less happy and tender sounds showed a less prominent fundamental frequency. Happiness and tenderness had a shorter perceived duration, and sad sounds were more sustained. Third, Experiment 2 from Korsmit et al. (2023) tested the affective response to chromatic scales. As with single tones, energy was predicted by an increase in pitch and decrease in inharmonicity. However, dimensional and discrete affect were most frequently predicted by median inharmonicity and spectral spread range. The synthesis of multiple datasets revealed consistent findings, but also discrepancies that may be explained by differences in stimulus selection. Furthermore, although findings on perceived and induced affect were largely similar, some findings on discrete affect and chromatic scales were not revealed with dimensional affect and single tones.
An analysis of robots (simulators) in education is provided. Promising directions for their development are highlighted, such as realism, interactivity, adaptation and personalization. The features of using simulators in dentistry are considered. The main disadvantages of existing simulators in dentistry have been identified, namely the lack of a communicative component and imitation of patient behavior. The anthropomorphic dental simulator is based on the Robo-C robot, which is a unique combination of advanced technologies and human facial expressions, which allows it to communicate with people, reproduce movements of different parts of the body and express emotions. As dental components, the following components were created and implemented into the Robo-C control system: a Smart jaw, including cameras and a temperature sensor, and a Smart tooth, including a pressure sensor. The Robo-C control system has been upgraded taking into account the Smart jaw and Smart tooth, which made it possible to connect the dental treatment process with the robot’s servos through its linguistic base. The process of analyzing data obtained from Smart jaw cameras using a neural network is described. A two-stage classification scheme for dental defects has been proposed and its effectiveness has been proven. The linguistic base contains a set of rules with the help of which devices (microphone, speakers, servos, Smart jaw, Smart tooth) interact with each other. An example of compiling a linguistic database rule is given. The linguistic base, Smart-jaw and Smart-tooth are configured for one of four cases: caries treatment, tooth preparation for a crown, tooth extraction, endodontic treatment. Treatment quality control is carried out using a comprehensive assessment of communication interaction with the robot and analysis of Smart-jaw and Smart-tooth data. An example of work in one of the cases is given. The anthropomorphic dental simulator presented in the article allows the use of new technologies in the training of dentists, as well as the simulation of various dental procedures, which will significantly improve the practical preparation of students for working with patients.
In this paper, we propose a method for removing linguistic information from speech for the purpose of isolating paralinguistic indicators of affect. The immediate utility of this method lies in clinical tests of sensitivity to vocal affect that are not confounded by language, which is impaired in a variety of clinical populations. The method is based on simultaneous recordings of speech audio and electroglotto-graphic (EGG) signals. The speech audio signal is used to estimate the average vocal tract filter response and amplitude envelop. The EGG signal supplies a direct correlate of voice source activity that is mostly independent of phonetic articulation. These signals are used to create a third signal designed to capture as much paralinguistic information from the vocal production system as possible-maximizing the retention of bioacoustic cues to affect-while eliminating phonetic cues to verbal meaning. To evaluate the success of this method, we studied the perception of corresponding speech audio and transformed EGG signals in an affect rating experiment with online listeners. The results show a high degree of similarity in the perceived affect of matched signals, indicating that our method is effective.
The strongest formulations of grounded cognition assume that perceptual intuitions about concepts involve the re-activation of sensorimotor experience we have made with their referents in the world. Within this framework, concreteness and imageability ratings are indeed of crucial importance by operationalising the amount of perceptual interaction we have made with objects. Here we tested such an assumption by asking whether visual intuitions about concepts are provided accurately even when direct visual experience is absent. To this aim, we considered concreteness and imageability intuitions in blind people and tested whether these judgments are predicted by Image-based Frequency (IF, i.e. a data-driven estimate approximating the availability of the word referent in the visual environment). Results indicated that IF predicts perceptual intuitions with a larger extent in sighted compared to blind individuals, thus suggesting a role of direct experience in shaping our judgements. However, the effect of IF was significant not only in sighted but also in blind individuals. This indicates that having direct visual experience with objects does not play a critical role in making them concrete and imageable in a person’s intuitions: people do not need visual experience to develop intuition about the availability of things in the external visual environment and use this intuition to inform concreteness/imageability judgments. Our findings fit closely the idea that perceptual judgments are the outcome of introspection/abstraction tasks invoking high-level conceptual knowledge that is not necessarily acquired via direct perceptual experience.