Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This paper investigates how different textile materials on the wrist bands may mediate the affective experience of two forms of tactile icons, namely stroking and squeezing. To this end, we designed Emoband, a wrist-worn textile-based tactile display that can stroke or squeeze on wearers’ wrists through the moving fabrics. We conducted two studies on the participants’ valence and arousal ratings towards the stroking and squeezing feedback from Emoband, respectively. Results showed that the valence ratings were significantly affected by the stroking or squeezing parameters and the wristband’s material, interactively. The movement-related parameters dominated the arousal rating. Meanwhile, the valence ratings demonstrated different tendencies between the two tactile feedback, with the valence increasing and decreasing along with the strength of stroking and squeezing, respectively. Besides, the valence ratings of the stroking stimuli demonstrated varied distributions across different materials, indicating the materials’ varying affective expressiveness in the wrist-worn haptic devices.
O artigo analisa variações no PetroGold, um treebank padrão ouro desenvolvido para o processamento de linguagem natural (PLN). Os resultados mostram que considerar a classe gramatical das expressões multipalavras na anotação de todas as palavras que as compõem, assim como simplificar o tagset sintático do treebank, produz modelos com melhor desempenho em algumas métricas, destacando a importância da modelagem linguística durante a anotação para resultados adequados no PLN. Os datasets utilizados no estudo estão disponíveis em um repositório dedicado, podendo ser ainda mais modificados para treinar melhores modelos de linguagem.
Emotional vocabulary development represents a growing field of interest.Studies show that children use emotional words starting from the age of two (Izard & Harris, 1995; Michalson & Lewis, 1985).Between 3 and 5 years, children develop their ability to name basic emotions (e.g., joy, sadness, surprise, disgust, anger, fear).Li and Yu (2015) observed that 2-13-years-old Chinese children comprehend positive emotion words earlier than negative and neutral ones.This result is probably link to the fact that valence is an early key dimension in emotion conceptual representation (Nook et al., 2017).If many studies investigated emotion words comprehension, only few studies investigated spontaneous vocabulary young children use to refer to emotions in reaction to emotional and non-emotional stimuli.The present experiment measured the use of emotional vocabulary during an emotional valence rating task of words, pictures, words-pictures combined.More precisely, 178 young French children aged 4-, 5-and 6-years-old were observed while rating stimuli.These ratings were made using a three points emotional valence rating scale (negative, neutral, and positive) based on AEJE scale (Largy, 2018).The 90 words, 90 pictures, 90 words-pictures combined were divided in sets of 15 stimuli.Each child rated all sets of stimuli in separate sessions in random order.Children's utterances containing emotion words were recorded.The content form of these emotional words produced was analyzed thanks to EMOVAL software (Leveau, Jhean-Larose & Denhire, 2011).EMOVAL is an automatic evaluation of emotional valence and arousal of texts, sentences using a 5656 root-words meta norm in French and in English.It also extracts emotional semantic topics from texts and sentences.First, analyses highlighted that young children use more positive emotion vocabulary compared to negative one.This result is congruent with the positive bias observed while young children rated emotional valence of stimuli (Syssau & Monnier, 2009).Like comprehension of emotional words, the use of positive emotion vocabulary occurs earlier in the development than negative one.Second, it was found that with increasing of age, the use of emotion vocabulary enhanced.If children aged 4 used topics that refer to primary emotions, children aged 5 and 6 used larger and more complex emotional topics.Discussion focused on the understanding of children's daily emotional language environments, and the implications of these for early educators and teachers.
Psycholinguistic properties such as concept familiarity and concreteness have been investigated in relation to technological innovations in teaching and learning. Due to ongoing advances in semantic representation and machine learning technologies, the automatic extrapolation of lexical psycholinguistic properties has received increased attention across a number of disciplines in recent years. However, little attention has been paid to the reliable and interpretable assessment of familiarity ratings for domain concepts. To address this gap, we present a regression model grounded in advanced natural language processing and interpretable machine learning techniques that can predict domain concepts’ familiarity ratings based on their lexical features. Each domain concept is represented at both the orthographic–phonological level and semantic level by means of pretrained word embedding models. Then, we compare the performance of six tree-based regression models (adaptive boosting, gradient boosting, extreme gradient boosting, a light gradient boosting machine, categorical boosting, and a random forest) on domain concepts’ familiarity rating prediction. Experimental results show that categorical boosting with the lowest MAPE (0.09) and the highest R2 value (0.02) is best suited to predicting domain concepts’ familiarity. Experimental results also revealed the prospect of integrating tree-based regression models and interpretable machine learning techniques to expand psycholinguistic resources. Specifically, findings showed that the semantic information of raw words and parts of speech in domain concepts are reliable indicators when predicting familiarity ratings. Our study underlines the importance of leveraging domain concepts’ familiarity ratings; future research should aim to improve familiarity extrapolation methods. Scholars should also investigate the correlation between students’ engagement in online discussions and their familiarity with domain concepts.
The Word Lexical Database is a vast collection of English words that categorizes nouns, verbs, adjectives, and adverbs into sets of cognitive synonyms. This database may seem similar to a thesaurus, as it also groups words based on their meanings, but there are key differences. Unlike a thesaurus that only links word forms, the Word Lexical Database links specific senses of words, thereby disambiguating words found in close proximity. Additionally, the Word Lexical Database annotates the semantic relationships between words, whereas a thesaurus does not follow any explicit pattern besides meaning similarity in the grouping of words.
According to the Facial Feedback Hypothesis (FFH), feedback from facial muscles can initiate and modulate a person’s emotional state. However, this assumption is debated, and existing research has arguably suffered from a lack of control over which facial muscles are activated, when, to what degree, and for how long. To overcome these limitations, we carried out a pre-registered experiment recruiting 58 participants in 2023 in which we applied facial neuromuscular electrical stimulation (fNMES) to the Zygomaticus Major (ZM) and Depressor Anguli Oris (DAO) muscles for 5 seconds at 100% and 50% of the participants individual motor threshold (MT). After each trial, participants reported their emotions’ valence and arousal. Heart rate and electrodermal activity were recorded throughout. Results showed that muscle activation through fNMES, even when controlling for fNMES-induced discomfort, modulated participants’ emotional state as expected, with more positive emotions reported after stronger stimulation of the ZM than the DAO muscle. The addition of expression-congruent emotional images increased the effect. Moreover, fNMES intensity predicted arousal ratings and skin conductance responses. The finding that changes in felt emotion can be induced through brief, controlled activation of specific facial muscles is in line with the FFH and offers exciting opportunities for translational intervention.
Abstract We propose a novel graph-based approach for semantic parsing that resolves two problems observed in the literature: (1) seq2seq models fail on compositional generalization tasks; (2) previous work using phrase structure parsers cannot cover all the semantic parses observed in treebanks. We prove that both MAP inference and latent tag anchoring (required for weakly-supervised learning) are NP-hard problems. We propose two optimization algorithms based on constraint smoothing and conditional gradient to approximately solve these inference problems. Experimentally, our approach delivers state-of-the-art results on GeoQuery, Scan, and Clevr, both for i.i.d. splits and for splits that test for compositional generalization.
Extinction training has proved effective to diminish the expectancy of the aversive unconditioned stimulus (US). However, the negative valence of the conditioned stimulus (CS) may still stay intact. In fact, several studies have suggested that the CS negative valence may be a factor that promotes the return of fear. Our study focuses on the role of changes in the CS valence as a potential mechanism to reduce the spontaneous recovery of threat expectancies. To do that, we evaluated counterconditioning (CC), a technique aimed to reduce the CS negative valence by paring it with a positive stimulus and compared its efficacy to that of a novelty-facilitated extinction (NFE) and a standard extinction interventions. Using a 2-day protocol, participants first learned the relationship between a figure and an aversive sound, using a differential conditioning paradigm, and were then randomly assigned to one of three different groups. For the CC group, CS+ or cue A was paired with a positive US. The standard extinction group was exposed to cue A alone. For a third NFE group, cue A was followed by a neutral US. Finally, on the second day, spontaneous recovery was tested. Our findings did not provide evidence to suggest that CC could be more effective to prevent or reduce the return of threat expectancies or influence valence ratings when compared with NFE and standard extinction.
Two distinct literatures have evolved to study within-person changes in affect over time. One literature has examined affect dynamics with millisecond-level resolution under controlled laboratory conditions, and the second literature has captured affective dynamics across much longer timescales (e.g., hours or days) within the relatively uncontrolled but more ecologically valid conditions of daily life. Despite the importance of linking these literatures, very little research has been done so far. In the laboratory, peak affect intensities and reaction durations were quantified using a paradigm that captures second-to-second changes in subjective affect elicited by provocative images. In two studies, analyses attempted to link these micro-dynamic indexes to fluctuations in daily affect ratings collected via daily protocols up to 4 weeks later. Although peak intensity and reaction duration scores from the laboratory did not consistently relate to daily scores pertaining to affect variability or instability, the total magnitude of changes in affect following images did display relationships of this type. In addition, higher peaks in the laboratory predicted larger intensity reactions to salient daily events. Together, the studies provide insights into the mechanisms through which correspondences and noncorrespondences between laboratory reactivity indices and daily affect dynamic measures can be expected.
The growing interaction between humans and machines raises the necessity to more sophisticated tools for natural language understanding. Dependency parsing is crucial for capturing the semantics of a sentence. Although graph-based dependency parsing approaches outperform transition-based methods because they are not exposed to error propagation as their compeer, their feature space is comparatively limited. Thus, the main issue with graph-based parsing is how to expand the set of features to improve performance. In this research, we propose to expand the feature space of graph-based parsers. To benefit from the global meaning of the entire sentence content, we employee the sentence representation as an additional token feature. Also, to highlight local word collaborations that build sub-tree structures, we use convolutional neural network layers over token embeddings. We achieve the state-of-art results for Turkish, English, Hungarian, and Korean by getting the unlabeled and labeled attachment scores respectively on the test sets; 82.64% and 76.35% on Turkish IMST, 93.36% and 91.34% on English EWT, 90.85% and 87.39% on Hungarian Szeged, 92.44% and 89.58% on Korean GSD treebanks. Our experimental findings show that augmented global and local features empower the performance of graph-based dependency parsers.
OBJECTIVES: We used three-dimensional (3D) virtual images to undertake a subjective evaluation of how different factors affect the perception of facial asymmetry among orthodontists and laypersons with the aim of providing a quantitative reference for clinics. MATERIALS AND METHODS: A 3D virtual symmetrical facial image was acquired using FaceGen Modeller software. The left chin, mandible, lip and cheek of the virtual face were simulated in the horizontal (interior/exterior), vertical (up/down), or sagittal (forward or backward) direction in 3, 5, and 7 mm respectively with Maya software to increase asymmetry for the further subjective evaluation. A pilot study was performed among ten volunteers and 30 subjects of each group were expected to be included based on 80% sensitivity in this study. The sample size was increased by 60% to exclude incomplete and unqualified questionnaires. Eventually, a total of 48 orthodontists and 40 laypersons evaluated these images with a 10-point visual analog scale (VAS). The images were presented in random order. Each image would stop for 30 s for observers with a two-second interval between images. Asymmetry ratings and recognition accuracy for asymmetric virtual faces were analyzed to explore how different factors affect the subjective evaluation of facial asymmetry. Multivariate linear regression and multivariate logistic regression models were used for statistical data analysis. RESULTS: Orthodontists were found to be more critical of asymmetry than laypersons. Our results showed that observers progressively decreased ratings by 1.219 on the VAS scale and increased recognition rates by 2.301-fold as the degree of asymmetry increased by 2 mm; asymmetry in the sagittal direction was the least noticeable compared with the horizontal and vertical directions; and chin asymmetry turned out to be the most sensitive part among the four parts we simulated. Mandible asymmetry was easily confused with cheek asymmetry in the horizontal direction. CONCLUSIONS: The degree, types and parts of asymmetry can affect ratings for facial deformity as well as the accuracy rate of identifying the asymmetrical part. Although orthodontists have higher accuracy in diagnosing asymmetrical faces than laypersons, they fail to correctly distinguish some specific asymmetrical areas.
Abstract We describe the first steps in preparation of a treebank of 14 th -century Czech in the framework of Universal Dependencies. The Dresden and Olomouc versions of the Gospel of Matthew have been selected for this pilot study, which also involves modification of the annotation guidelines for phenomena that occur in Old Czech but not in Modern Czech. We describe some of these modifications in the paper. In addition, we provide some interesting observations about applicability of a Modern Czech parser to the Old Czech data.
Charles Dickens (1812-1870) achieved a recognizableplace among English writers through the use of the stylisticfeatures in his fictional language. This study is concernedwith Dickens' unique fictional language, used in one of hisnovels entitled "Hard Times", in relation to phonologicaldeviation from settled norms in English. It endeavors toshow Dickens' manipulating language and the effectsachieved through this manipulation.This research investigates Dickens' use of languagewhich deviates from the linguistic norm phonologically. Assuch, it is hypothesized that Dickens used phonologicaldeviation to show the character's social class.The study aims to analyze the types of phonologicaldeviations in Dickens' "Hard Times". It determines thereasons behind these deviations, and how that reflects
This research paper aims to present the feminist side of translation. It aims to reconsider and reshape translation from the feminist point of view. It does so by introducing the feminist theory to translation and applying it to some excerpts. These excerpts are taken from the great English novelist Virginia Woolf’s (1977) A Room of one’s own. The theory in this study is the feminist theory in translation by the famous theorist Louise von Flotow (1991). This theory contains three strategies. The first is supplementing. It means adding elements to the translation to compensate for what the language lacks, such as gender agreement in English. This strategy fills the gap between conventional linguistic norms and feminist purposes. The second strategy is Footnotes and prefaces which present explanations for translational choices and linguistic references at the beginning of the text or throughout it to make women’s voices visible. The third strategy is Hijacking which is an appropriation of the text, including the changes applied to it to suit feminist translators’ aims.
Purpose: The growing interdisciplinary nature of medicine has prompted increased leadership by physicians, both in administration and as informal team leaders. A formal leadership curriculum is essential to best prepare trainees for these roles.1 Physician Executive Leadership (PEL) is an organization at Thomas Jefferson University that provides medical students with leadership education by facilitating sessions led by established medical experts in executive positions. Each session is rooted in 1 or more of PEL’s leadership pillars: Applied Leadership, Quality Improvement, Health Finance, Entrepreneurship/Innovation, Health Policy, and Law/Ethics. In response to COVID-19 public health precautions, PEL sessions were shifted online. Literature is contradictory on the relative benefits and harms of online coursework in a leadership curriculum; although it increases accessibility, it may decrease student engagement.2 We therefore aimed to analyze the effectiveness of in-person versus virtual PEL curriculum sessions and to determine if certain topics were better taught in-person or virtually. Method: Six hundred eighty-four total responses (from 361 first-, 226 second-, 78 third-, and 19 fourth-year medical student submissions) were collected from 24 sessions from January 2018 to May 2022, including 13 in-person (2018–2019) and 11 virtual (2020–2022) sessions. Six Applied Leadership, 4 Quality Improvement, 8 Health Finance, 5 Entrepreneurship/Innovation, 2 Health Policy, and 3 Law/Ethics sessions were assessed. Following each session, students rated changes in their engagement and comprehension of the topic area (Likert scale 1–7: 1 = strongly disagree; 7 = strongly agree). One-way nonparametric ANOVA and unpaired t tests were used to compare responses between the 2 curricula. Results: Health Finance engagement ratings were higher online than in-person (6.49 versus 5.85; P <.001). Similarly, Health Finance (6.31 versus 5.79; P <.01) and Care Quality sessions (6.29 versus 5.73; P <.01) demonstrated improved comprehension with online delivery, while Health Policy demonstrated the opposite trend (5.72 versus 6.64; P <.01). Overall, student engagement and comprehension following the session were nonsignificantly higher virtually compared to in-person (6.33 versus 6.21; P =.08 for engagement, and 6.22 versus 6.09; P =.09 for comprehension). Comprehension and engagement scores were not significantly affected for the remaining pillars. Medical school class of participants (MS1 through MS4) did not affect ratings. Discussion: Overall ratings for both comprehension and engagement of each event remained high (> 6.0) across pillars in-person (2018–2019) and online (2020–2022). Most notably, the Health Finance sessions scored higher in comprehension and engagement in online formats. This may be explained by the ability of students to look up clarifying information during the session. Conversely, the decline in comprehension of Health Policy may be attributed to the inability to quickly find answers to complex health policy questions online. Students may therefore benefit from directly asking questions to the experts during in-person sessions. A controlled study that compares particular events both in-person and virtually would be necessary to confirm these findings. Significance: Comprehension and engagement improvements for Health Finance sessions indicate that an online format may be better suited to teach this core leadership pillar, while the decline in comprehension of Health Policy indicates this pillar may benefit from being taught in-person. Importantly, there was not a drastic decrease in engagement nor comprehension for most pillars once sessions became virtual; therefore, the increased accessibility, availability of multimedia tools, and decreased cost (for travel, lodging, etc.) suggests virtual events may be superior to their in-person counterparts for the majority of leadership curricular events.
Previous research on linguistic relativity and economic decisions hypothesized that speakers of languages with obligatory tense marking of future time reference (FTR) should value future rewards less than speakers of languages which permit present tense FTR. This was hypothesized on the basis of obligatory linguistic marking (e.g., will) causing speakers to construe future events as more temporally distal and thereby to exhibit increased "temporal discounting": the subjective devaluation of outcomes as the delay until they will occur increases. However, several aspects of this hypothesis are incomplete. First, it overlooks the role of "modal" FTR structures which encode notions about the likelihood of future outcomes (e.g., might). This may influence "probability discounting": the subjective devaluation of outcomes as the probability of their occurrence decreases. Second, the extent to which linguistic structures are subjectively related to temporal or probability discounting differences is currently unknown. To address these, we elicited FTR language and subjective ratings of temporal distance and probability from speakers of English, which exhibits strongly grammaticized FTR, and Dutch, which does not. Several findings went against the predictions of the previous hypothesis: Framing an FTR statement in the present ("Ellie arrives later on") versus the future tense ("…will arrive…") did not affect ratings of temporal distance; English speakers rated future statements as relatively more temporally proximal than Dutch speakers; and English and Dutch speakers rated future tenses as encoding high certainty, which suggests that obligatory future tense marking might result in less discounting. Additionally, compared with Dutch speakers, English speakers used more low-certainty terms in general (e.g., may) and as a function of various experimental factors. We conclude that the prior cross-linguistic observations of the link between FTR and psychological discounting may be caused by the connection between low-certainty modal structures and probability discounting, rather than future tense and temporality.
Background: Type 2 diabetes mellitus (DM) is the most common metabolic disorder in the world and an important risk factor for peripheral arterial disease (PAD). CT angiography represents the method of choice for the diagnosis, pre-operative planning, and follow-up of vascular disease. Low-energy dual-energy CT (DECT) virtual mono-energetic imaging (VMI) has been shown to improve image contrast, iodine signal, and may also lead to a reduction in contrast medium dose. In recent years, VMI has been improved with the use of a new algorithm called VMI+, able to obtain the best image contrast with the least possible image noise in low-keV reconstructions. Purpose: To evaluate the impact of VMI+ DECT reconstructions on quantitative and qualitative image quality in the evaluation of the lower extremity runoff. Materials and Methods: We evaluated DECT angiography of lower extremities in patients suffering from diabetes who had undergone clinically indicated DECT examinations between January 2018 and January 2023. Images were reconstructed with standard linear blending (F_0.5) and low VMI+ series were generated from 40 to 100 keV, in an interval of 15 keV. Vascular attenuation, image noise, signal-to-noise ratio (SNR), and contrast-to-noise ratio (CNR) were calculated for objective analysis. Subjective analysis was performed using five-point scales to evaluate image quality, image noise, and diagnostic assessability of vessel contrast. Results: Our final study cohort consisted of 77 patients (41 males). Attenuation values, CNR, and SNR were higher in 40-keV VMI+ reconstructions compared to the remaining VMI+ and standard F_0.5 series (HU: 1180.41 ± 45.09; SNR: 29.91 ± 0.99; CNR: 28.60 ± 1.03 vs. HU 251.32 ± 7.13; SNR: 13.22 ± 0.44; CNR: 10.57 ± 0.39 in standard F_0.5 series) (p < 0.0001). Subjective image rating was significantly higher in 55-keV VMI+ images compared to the other VMI+ and standard F_0.5 series in terms of image quality (mean score: 4.77), image noise (mean score: 4.39), and assessability of vessel contrast (mean value: 4.57) (p < 0.001). Conclusions: DECT 40-keV and 55-keV VMI+ showed the highest objective and subjective parameters of image quality, respectively. These specific energy levels for VMI+ reconstructions could be recommended in clinical practice, providing high-quality images with greater diagnostic suitability for the evaluation of lower extremity runoff, and potentially needing a lower amount of contrast medium, which is particularly advantageous for diabetic patients.
Abstract We propose a Slovak language model for the spaCy library in Python. These models are easy-to-use for basic natural language processing tasks in a single package. The package contains several components for basic preprocessing tasks, such as tokenization, sentence boundary detection, syntactic parsing, lemmatization, named entity recognition, morphology analysis, and word vectors. It is based on the state-of-the-art monolingual SlovakBERT model. Named entity recognition is trained on a separate, publicly available WikiAnn database. The other statistical classifiers use a Slovak Dependency Treebank corpus. Morphological tags are compatible with the conventions of the Slovak National Corpus. The part of speech tags use conventions of the Universal Dependencies framework. We trained a separate word vector model on a web-based corpus. The training uses fastText with Floret modification. We present a series of experiments that confirm that the model performs similarly to other languages for all tasks. Training scripts and data are publicly available.
В статье анализируется историческая динамика политической корректности, ее положительные и отрицательные стороны, а также прослеживается ее связь с вежливостью, которая заключается в неиспользовании конфронтационных, ликоповреждающих коммуникативных стратегий при указании на гендер, расу, этнос, возраст, физическое состояние и социальное положение адресата. Исследуются факторы, затрудняющие формирование политкорректного русского языкового сознания: 1) не вполне сформировавшееся понятие ПК применительно к отечественному социальному контексту, осложненное его концептуализацией через призму западного восприятия; 2) отсутствие правовых механизмов ПК, несмотря на недопустимость дискриминации, закрепленную в Конституции РФ; 3) неразработанность языковых основ применения ПК в российском публичном дискурсе; 4) англоцентричность правил ПК для международного общения. Сделан вывод о необходимости выработки российских норм политической корректности с участием широкого лингвистического сообщества. The paper explores the historical dynamics of political correctness (PC), its positive and negative aspects, as well as its connection with politeness as an avoidance of confrontational, face-threatening strategies in reference to the interlocutor’s gender, race, ethnic background, age, physical condition and social status. The study also deals with the factors hindering the formation of the Russian PC awareness, which include: 1) the incompletely formed notion of political correctness in the Russian social context complicated by its conceptualization through the prism of Western comprehension; 2) absence of legal PC mechanisms, in spite of non-discrimination enshrined in the RF Constitution; 3) underdeveloped linguistic norms of political correctness in Russian public discourse; 4) anglocentrism of PC rules in intercultural communication. The article concludes by proposing a wide discussion of Russian political correctness norms involving a wider linguistic community.
PURPOSE: Slow speech rate and abnormal temporal prosody are primary diagnostic criteria for differentiating between people with aphasia who do and do not have apraxia of speech. We sought to identify appropriate cutoff values for abnormal word syllable duration (WSD) in a word repetition task, interpret them relative to a data set of people with chronic aphasia, and evaluate the extent to which manually derived measures could be approximated through an automated process that relied on commercial speech recognition technology. METHOD: Fifty neurotypical participants produced 49 multisyllabic words during a repetition task. Audio recordings were submitted to an automated speech recognition (ASR) service (IBM Watson) to measure word duration and generate an orthographic transcription. The transcribed words were compared to a lexical database, and the number of syllables was identified. Automatic and manual measures were compared for 50% of the sample. Results were interpreted relative to WSD scores from an existing data set of 195 people with mostly chronic aphasia. RESULTS: ASR correctly identified 83% of target words and 98% of target syllable counts. Automated word duration calculations were longer than manual measures due to imprecise cursor placement. Upon applying regression coefficients to the automated measures and examining the frequency distributions for both manual and estimated measures, a WSD of 303-316 ms was found to indicate longer-than-normal performance (corresponding to the 95th percentile). With this cutoff, 40%-45% of participants with aphasia in our comparison sample had an abnormally long WSD. CONCLUSIONS: We recommend using a rounded WSD cutoff score between 303 and 316 ms for manual measures. Future research will focus on customizing automated WSD methods to speech samples from people with aphasia, identifying target words that maximize production and measurement reliability, and developing WSD standard scores based on a large participant sample with and without aphasia.
Russian constructicon is an open-access linguistic database containing detailed descriptions of over 3,800 Russian grammatical constructions. In this paper we present a new, enlarged and updated version of Russian Constructicon (RusCxn) as well as new trajectories of development which were opened for the resource after the update. Since its first release, RusCxn, has undergone many significant changes. Our team has expanded the number of constructions present in the database 1,5 times, introduced new meta-information features such as glosses, significantly reworked the architecture and the design of Russian Constructicon’s website, and improved the search facilities. The above-mentioned changes not only make RusCxn more attractive and convenient-to-use, but they can also greatly facilitate typological research in the field of Construction Grammar and improve the mapping between constructicography-orinented resources for different languages.
Previous studies have made great advances in RST discourse parsing through specific neural frameworks or features, but they usually split the parsing process into two subtasks and heavily depended on gold discourse segmentation. In this article, we introduce an end-to-end method for sentence-level RST discourse parsing via transforming it into a text-to-text generation task, which can also be simply applied to document-level parsing. Our method unifies the traditional two-stage parsing and generates the parsing tree directly from the input text through our constrained decoding and postprocessing algorithms, without requiring a complicated model. Moreover, the discourse segmentation can be simultaneously generated and extracted from the parsing tree. Experimental results on the RST Discourse Treebank demonstrate that our proposed method outperforms existing methods in both the tasks of discourse parsing and segmentation. We further carry out ablation studies and more targeted comparisons with traditional patterns to analyze our method in more detail. Considering the lack of annotated data in RST parsing, we also create high-quality augmented data and implement self-training, which further improves the performance of our method.
* Introduction This is the Khmer ALT of the Asian Language Treebank (ALT) Corpus. Please refer to<br> http://www2.nict.go.jp/astrec-att/member/mutiyama/ALT/index.html<br> for an introduction of the ALT project. The process of building the Khmer ALT began with sampling about 20,000 sentences from English Wikinews, and then these sentences were translated into Khmer language.<br> <br> The English Wikinews<br> https://en.wikinews.org/wiki/Main_Page<br> is available under the terms of the Creative Commons Attribution 2.5 License.<br> https://creativecommons.org/licenses/by/2.5/ * License Khmer ALT has been developed by NICT and CADT (a.k.a. NIPTICT). The license of Khmer ALT is Academic Research Non-Commercial Limited CC-BY-NC-SA Reference-Type License Agreement Terms and Conditions The use of this material (copyrighted material, data, etc.) is permitted under the conditions similar to the Creative Commons "Attribution-NonCommercial-ShareAlike" License. However, the purpose of use shall not only be "NonCommercial" but must be “for Academic Research and Non-Commercial”.<br> The detailed legal provisions are listed at the end of the document, but the outline is as follows. You are free to:<br> Share — copy and redistribute the material in any medium or format<br> Adapt — remix, transform, and build upon the material<br> The licensor cannot revoke these freedoms as long as you follow the license terms. Under the following terms:<br> Attribution — You must give appropriate credit, provide a link to the license, and indicate if changes were made. You may do so in any reasonable manner, but not in any way that suggests the licensor endorses you or your use.<br> Academic Research&Non-Commercial — You may use the material only for academic research purposes and may not use it for commercial purposes.<br> ShareAlike — If you remix, transform, or build upon the material, you must distribute your contributions under the same license as the original. In this Terms and Conditions, academic research purposes are not associated with the interpretation of the Creative Commons License, and shall refer to "the purpose of intellectual creative activities conducted by persons belonging to institutions of higher education and research institutes, etc. in order to develop a better and affluent society through exploration of the truth about nature, humans, society, etc., discovery of new principles and laws, and application of these in a wide range of fields from natural sciences to human and social sciences." Non-commercial means: not for the purpose of any contribution to a for-profit or commercial business. The outcome of research and development activities initially undertaken without the purpose of contributing in any way to a for-profit or commercial business shall not be permitted to be used later to contribute to one. The legal provisions of this license are: the terms of Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode) plus the following rider clauses. On that note, the Terms and Conditions of this license are not equal to that of the Creative Commons license—they were inspired by them and share most of the terms and conditions with them, but shall be considered as a different license. Rider Clauses<br> 1. "Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License" in the title and preamble shall read "Academic Research Non-Commercial Limited CC-BY-NC-SA Compliant License Agreeement Terms and Conditions".<br> 2. Article 1 c and g shall be deleted.<br> 3. Article 1 k shall be amended as follows:<br> k. "Academic Research Purposes" shall refer to the purpose of intellectual creative activities conducted by persons belonging to institutions of higher education and research institutes, etc. in order to develop a better and affluent society through exploration of the truth about nature, humans, society, etc., discovery of new principles and laws, and application of these in a wide range of fields from natural sciences to human and social sciences, and does not consider contribution to for-profit or commercial businesses as its main or secondary purpose. Activities initially undertaken without the purpose of contributing to a for-profit or commercial business, shall no longer be considered as "academic research purposes" as soon as they decide to contribute to one.<br> 4. Article 3 b.1. shall be replaced by the following:<br> 1. The Adapter's License that You apply must be this License.<br> 5. In all provisions, "NonCommercial" shall read "Academic Research Purposes" and "this Public License" shall read "this License". Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International Public License<br> <https://creativecommons.org/licenses/by-nc-sa/4.0/legalcode> * Usage 1. Compile the source codes.<br> 2. Run with no option to see the instruction.<br> 3. Follow the instruction to generate the treebank.
Decoding Part of Speech(POS) tagging directly from electroencephalography (EEG) signals whilst user overtly spoke (voiced speech) sentences could improve direct speech brain-computer interfaces (BCIs) using imagined or inner speech. To the best of our knowledge, earlier work uses machine learning approach using 74,953 sentences/tokens recorded in 75 EEG sessions. The tokens can be found in 4,479 phrases consisting of terms from the English Online treebank which contains the record of weblogs, newsgroups, reviews, and Yahoo Answers. The results demonstrated the feasibility of POS decoding from EEG based on word class, word frequency, and word length with accuracy of 71%, 86%, 89%, respectively. We believe that there is significant room for improvement with more advanced artificial intelligence. In this paper, we further extend the existing work with end-to-end transformers. Our results presents transformer model outperforms benchmark traditional ML results with +20% in length, +13% for the open vs closed class and +12% in frequency. In our empirical analysis, we find the decoding performance was better when using multi-electrode recordings as compared to single-electrode recordings.
This chapter synthesizes the most prominent Natural Language Processing (NLP) studies conducted on Persian, focusing on text processing. The first section contains selected tasks from the NLP pipeline, such as text preprocessing, tokenization, POS tagging, syntactic parsing, treebank annotation or semantic analysis along with examples of how researchers approached the problem for Persian and, where applicable, examples of tools developed to perform given tasks. The following section discusses the application of Persian NLP like spell-checking, information retrieval, machine translation or sentiment analysis. Finally, the last section summarizes the Persian NLP corpora and other resources.
This paper provides a comparative analysis of the patterns of formation and qualities of a modern linear text and an Internet text. The article is the result of the study of Internet text stylistics, based mainly on Russian-language texts from Russia and Ukraine. The paper considers the process of formation of a new type of text - the Internet text. Being essentially different from the classical linear text, the Internet text does not lend itself well to the description based on the classical text theory. Thus, the Internet text is not complete, vectorial, not united by the completeness of thought expression. An important characteristic of the Internet text is its interactivity, which means that the roles of addressee and addressee are constantly changing. In addition, it is difficult, or even impossible, to define the boundaries of the Internet text due to its hypertextuality, which has become habitual intertextuality. All the above-mentioned aspects make up the pragmatics of the Internet text as a subspecies of the media text. Another crucial problem addressed in the paper is the study of the regularities of Internet communication in general and the stylistics of the Internet text. In the course of the research it became obvious that speech aggression and violation of norms of speech culture are stylistic dominants of online communication. This influenced the formation of other stylistic dominants such as, hate speech, fake, hype, clickbait, etc. Internet style is clearly characterized by being provocative, aggressive, hostile. The problems of bullying, humiliation of human dignity, invective and obscenity are actively studied from the standpoint of linguoecology, because, according to most researchers, the constant neglect of communicative and linguistic norms leads to the degradation of the national language style and literary norms. Despite the fact that there is still a division between public and interpersonal online communication, i.e. formal and informal, the problems of speech behavior of Internet users are becoming increasingly relevant.
OBJECTIVE: Hoarding behaviour is a common but poorly characterised problem in real-world clinical practice. Although hoarding behaviour is the key component of Hoarding Disorder (HD), there are people who exhibit hoarding behaviour but do not suffer from HD. The aim of the present study was to characterise a clinical sample of patients with clinically relevant hoarding behaviour and evaluate the differential characteristics between patients with and without HD. METHODS: This study included patients who received treatment at the home visitation program in Barcelona (Spain) from January 2013 through December 2020, and scored ≥ 4 on the Clutter Image Rating scale. Sociodemographic, DSM-5 diagnosis, clinical data and differences between patients with and without an HD diagnosis were assessed. RESULTS: A total of 243 subjects were included. Hoarding behaviour had been unnoticed in its early stages and the median length in the sample was 10 years (IQR 15). 100% of the cases had hoarding-related complications. HD was the most common diagnosis in 117 patients (48.1%). CONCLUSIONS: The study found several differential characteristics between patients with and without HD diagnosis. Alcohol use disorder could play an important role among those without HD diagnosis. Home visitation programs could improve earlier detection, preventing hoarding-related complications.
This dataset offers a multilingual corpus for studying Persian translations of Plato’s Crito in comparison with the Ancient Greek source text. It is designed for scholars and students working in fields such as translation studies, computational linguistics, and philology. This repository includes: The Ancient Greek text of Crito (Burnet edition) Five Persian student translations, word-level and sentence-level alignments Three finalized Persian translations, word-level andsentence-level alignments, including UD treebank annotations Two English translations (Jowett and Fowler), sentence-level aligned One German translation (Schleiermacher), sentence-level aligned A Greek–Persian lexical wordlist for Crito Source Texts & References Greek text: Plato. Platonis Opera, ed. John Burnet. Oxford University Press. 1903.Source: Perseus Digital Library – https://www.perseus.tufts.edu/hopper/text?doc=Perseus%3Atext%3A1999.01.0169%3Atext%3DCrito%3Asection%3D43a Student Persian translations: Produced during a Greek translation course by elementary students after a 30-hour introductory course to Ancient Greek. Each translation was made using Ancient Greek commentaries, lexicons, and parallel English and German translations available in this dataset. Alignments were created by the translators using Ugarit.These translations are available as word-level and sentence-level alignments. The word alignments are extracted from Ugarit, but they are still visualized on Ugarit. Translators’ Ugarit profiles: - Shouresh Assimi: https://ugarit.ialigner.com/userProfile.php?userid=50956- Aylar Mahmoudzadeh Sarabi: https://ugarit.ialigner.com/userProfile.php?userid=63464&tgid=9576- Nima Mohammadi: https://ugarit.ialigner.com/userProfile.php?userid=52434&tgid=9362- Kimia Nikpour: https://ugarit.ialigner.com/userProfile.php?userid=52378- Farshid Rahimi: https://ugarit.ialigner.com/userProfile.php?userid=50932&tgid=9727 Finalized Persian translations: Three finalized Persian translations, revised under the supervision of Farnoosh Shamsian.These translations are available as word-level and sentence-level alignments, accompanied by UD treebank annotation English translations: Fowler’s translation:Plato, H. N. Fowler, W. Lamb, Plato in Twelve Volumes, Vol. 1, translated by Harold North Fowler; Introduction by W.R.M. Lamb, volume 1, Harvard University Press and Wiliam Heinemann Ltd., Cambridge, MA and London, 1966.Fowler's translation on Perseus Digital Library:https://www.perseus.tufts.edu/hopper/text?doc=plat.+crito+43aAlignment: Sentence-level Jowett's translation:Plato, B. Jowett, Crito, The Internet Classics Archive, Massachusetts Institute of Technology, http://classics.mit.edu/Plato/crito.html.Alignment: Sentence-level German translation: Plato, F. Schleiermacher, Platons Werke, In der Realschulbuchhandlung, 1809Link to the text on Project Gutenberg:https://www.projekt-gutenberg.org/platon/platowr1/kriton.html Alignment: Sentence-level Dataset Contents To ensure proper rendering of Persian RTL text, files are available in both CSV and XML formats when necessary. Student translations aligned at sentence-level:Sentence-level alignments of the Ancient Greek text of Plato’s Crito with all five student Persian translations, two English translations (Jowett and Fowler), and one German translation (Schleiermacher).The sentence number is the treebank sentence ID followed by the Stephanus paginations.Files and Formats available:crito_sentence_alignments_student_translations.csvcrito_sentence_alignments_student_translations.xlsxNotes: All eight translations aligned to the same sentence boundaries as the Greek text. Student alignment data extracted from Ugarit:Word-level alignments of the student translations, which were done throughout the course by each student on Ugarit. The zip file includes five folders, each containing the Ugarit user ID of the student who did the translation and the alignment, followed by their last name in parentheses. Each folder includes the alignments of Crito done by the student, exported from Ugarit as a JSON file. The name of each JSON file is the identifier of the alignment on Ugarit.Formats available:student_alignments_ugarit.zip, containing JSON files Three finalized Persian translations aligned at sentence-level:Sentence-level alignments of the Ancient Greek text of Crito with the three finalized Persian translations under the supervision of Farnoosh Shamsian. The sentence number is the treebank sentence ID followed by the Stephanus paginations. Formats available:crito_finalized_translations_sentence_alignment.csvcrito_finalized_translations_sentence_alignment.xlsx Three finalized Persian translations aligned at word-level to the UD treebanks:Word-level alignments of the three finalized Persian translations to the Ancient Greek text, accompanied by UD Treebank annotations. These alignments are done by a single annotator according to alignment guidelines previously published on Zenodo here:Shamsian, F. (2023). Alignment Guidelines for Greek-Persian. Zenodo. https://doi.org/10.5281/ZENODO.8039931Formats available:crito_finalized_translations_treebank_word_alignment.csvcrito_finalized_translations_treebank_word_alignment.xlsxNotes: The translations are aligned to the treebanks provided by the Pedalion project, which were automatically generated and manually corrected. See more here: https://perseids-publications.github.io/pedalion-trees/ A Greek-Persian wordlist:A bilingual wordlist containing Greek–Persian lexical correspondences for every Greek word attested in Plato’s Crito.Formats available:crito_greek_persian_wordlist.csvcrito_greek_persian_wordlist.xlsx Notes This is the actively maintained and most complete release. An earlier version (only the three finalized translations) is now outdated, but still archived on Zenodo for reference here: Shamsian, F., Assimi, S., Mahmoudzadeh Sarabi, A., Mohammadi, N., Nikpour, K., & Rahimi, F. (2023). Persian Translations of Crito, Three Versions with Word-level Alignment. Zenodo. https://doi.org/10.5281/zenodo.8333681 Refrences to the consulted print sources and other Persian translations: Adam, J. (1888). Platonis Crito. Cambridge University Press. Emlyn-Jones, C. (1999). Crito. London: Bloomsbury Academic. Fowler, H. N., Lamb, W. R. M., & Shorey, P. (1966). Plato in Twelve Volumes: With an English Translation. Harvard University Press. Jowett, B. (1909). The Apology, Phaedo, and Crito of Plato, with introduction, notes and illustrations. New York P.F. Collier & son. http://archive.org/details/apologyphaedocri0000plat Plato (1938) Hekmate Soghrāt va Aflātun (M. A. Foroughi, Trans.) Tehran: Majles Publishing house. Plato (1979) Dore-ye Āsāre Aflātun (H. Lotfi, Trans.) Tehran: Kharazmi. Plato (1989) Vāpasin Ruzhāye Soghrāt (J. Jahanshahi, Trans.) Tehran: Porsesh. Plato (2013) Mohākeme-ye Soghrāt (L. Golestan, Trans.) Tehran: Markaz. Plato (2023) Kriton (I. Shafi’-Beyk, Trans.) Tehran: Ney. Schleiermacher, F. (1910). Platons Apologie und Kriton. Verlag Philipp Reclam.
Aesthetic processing has profound implications for everyday life. Although liking and beauty judgements are outcomes of aesthetic processing and derive from a common hedonic value, there may be some differences in how they engage working memory. This study used maintenance and aesthetic judgement tasks to examine whether liking and beauty judgements make different demands on domain-specific working memory resources. Sixty participants (30 males) were instructed to rate picture for liking or beauty while maintaining the subjective affect or brightness of the presented pictures. Results indicated that liking judgements selectively impaired participants' performance in the affect maintenance task, and beauty judgements selectively impaired their performance in the brightness maintenance task. In addition, maintaining affect and brightness feelings in the mind increased image ratings on beauty but not on liking. Our findings provide evidence that liking judgements draw more on affective working memory resources than beauty judgements, and beauty judgements draw more on visual working memory resources than liking judgements.
This article is determined to analyze the way Russian political and ideological system is using linguistics in the means of its informational and psychological aggression against Ukraine. The war launched by Russian Federation has become the culmination of its prolonged aggressive actions on many fronts: ideological, economical, cultural, linguistic and so on. Current cremlin hosts are trying to regain full control over Ukraine in every imaginable way, drawing on the previous political regimes' experience. This article investigates one of the definitions of chauvinism as harassment of so-called "small nations" on the domestic and international levels that is shaping into political oppression and assimilation of languages. Have it that cremlin ideologues have always exploited any opportunity to diminish Ukrainians' self-consciousness, in particular by means of different quasi-scientific theories. Here we explored one example of the forementioned such as the creation of fake exception to the rule of spelling of prepositions "in" and "on" with administrative geographical names by Russian philologers that only applied to place-name "Ukraine". In contradiction to linguistic norms of Russian language Russians use "on Ukraine" in accusative and locative cases. We respectively analyzed arguments of Soviet linguists D. Rozental and K. Bilinskyy, and also modern Russian geographers and their theory of geo concept that basically comes to one statement: such version has been made up through history and is backed up by the expression "on the outskirts". The Ukrainian linguist I. Ohiyenko's work, in which he explained that the expression "in Ukraine" positioned it as a separate state, is mentioned in the article. After the collapse of the Soviet Union the logical norm "in Ukraine" was being used in the official documents of Russian Federation for a while, but after the cremlin tacked its course to re-establish its dominance over the former Soviet republics, they returned to the previous version. The sources that have been studied in the article point out that the political and ideological system is of a paramount importance for Russian linguistic science and is using it as a non-lethal weapon against Ukraine.
Nouns are one of the most frequently used word classes in English. As one variety of “World Englishes”, China English has received much concern in recent years. However, most research on nouns is still at the cognitive aspect or in the traditional qualitative way. As a new quantitative method, dependency grammar reveals the internal relations among the words in the sentence. Therefore, the syntactic features of nouns in China News English can provide a new perspective. The raw material of China English in this study is randomly crawled from China Daily, and the News part in the Corpus of Contemporary American English (COCA) from 2013 to 2016is selected as the reference corpus. By using the Jupyter Notebook based on Python, the dependency treebanks are generated. After comparing the proportion of the word classes, dependency direction, and dependency distance of different dependents of nouns in China Daily and COCA, this paper finds that 1) On the whole, the categories of dependents in China English and American English are similar. Determiners, adjectives, nouns, prepositions, verbs, numbers, proper nouns, conjunctions, pronouns and particles are the top ten frequently-used modifiers of nouns. Compared with American English, China English tends to use more adjectives and nouns, while the determiners especially the articles and pronouns are relatively less used. 2) Both in China English and American English, the percentage of HI constructions and HF constructions is close to 50%. China English tends to be more right-embedded than American English. Compared with AmE, in ChiE, prepositions, conjunctions and participles are more frequently used as post-modifiers. 3) The mean dependency distance value of China English (2.68) is larger than American English (2.61). Both in China English and American English, the mean dependency distance of pre-modifiers is shorter than in post-modifiers. The longer the mean dependency distance value is, the more difficult is in processing the information. It indicates that in the News domain, China English is more difficult to be processed in expression. This research provides a new perspective to studying English variants and enriches the current research on nouns.
A dictionary’s basic function is to provide the meaning for all of the words contained within it. However, in addition to this key role, dictionaries also serve as a model for linguistic correctness, especially with reference to spelling norms and, consequently, they afford users an important guide to current orthographic standards. This chapter offers a series of reflections concerning a dictionary’s dual normative and semantic functions. First, we examine the extent to which linguistic norms should constitute part of lexicographic entries and then focus on the dictionary’s scope as a prescriptive reference tool that affirms and promotes a particular set of linguistic norms. Second, we debate the question of whether or not a dictionary should resolve any type of doubt that exists with respect to linguistic correctness, given the fact that other prescriptive tools, such as reference grammars, carry out the same prescriptive language function. We end this section with a reflection on the interdependence between language use and variation and how this interplay should be reflected in the formation of a set of ideal pan-Hispanic spelling norms. The second part of the chapter deals with the normative role played by the Royal Spanish Academy in conjunction with the Association of Spanish Academies in determining the nature and spread of orthographic standards for the Spanish language.
<em>This article touches upon the problems of improving the speech culture of students. It is known that one of the indicators of the level of culture, thinking, intellect of a person is his speech, which must comply with linguistic norms. It is in elementary school that children begin to master the norms of oral and written literary language, learn to use language tools in different communication conditions in accordance with the goals and objectives of speech.</em>
We present new supertaggers trained on English grammar-based treebanks and test the effects of the best tagger on parsing speed and accuracy. The treebanks are produced automatically by large manually built grammars and feature high-quality annotation based on a well-developed linguistic theory (HPSG). The English Resource Grammar treebanks include diverse and challenging test datasets, beyond the usual WSJ section 23 and Wikipedia data. HPSG supertagging has previously relied on MaxEnt-based models. We use SVM and neural CRF- and BERT-based methods and show that both SVM and neural supertaggers achieve considerably higher accuracy compared to the baseline and lead to an increase not only in the parsing speed but also the parser accuracy with respect to gold dependency structures. Our fine-tuned BERT-based tagger achieves 97.26\% accuracy on 950 sentences from WSJ23 and 93.88% on the out-of-domain technical essay The Cathedral and the Bazaar (cb). We present experiments with integrating the best supertagger into an HPSG parser and observe a speedup of a factor of 3 with respect to the system which uses no tagging at all, as well as large recall gains and an overall precision gain. We also compare our system to an existing integrated tagger and show that although the well-integrated tagger remains the fastest, our experimental system can be more accurate. Finally, we hope that the diverse and difficult datasets we used for evaluation will gain more popularity in the field: we show that results can differ depending on the dataset, even if it is an in-domain one. We contribute the complete datasets reformatted for Huggingface token classification.
Word embeddings trained on the lemmatised TOROT Treebank, using Word2Vec and the following parameters: sg = True min_count = <1,3,5> window = <3,5> vector_size = <100,200,300> epochs = 5 One model was trained for each combination of the parameters enclosed in angled brackets (< >). The release contains both the full models (.model) and the plain vector files (_vectors.txt). The models are named according to the parameters they were trained with. Note that these are the result of very preliminary experiments and no systematic evaluation of their quality was carried out, so use with caution.
Foregrounding denotes a linguistic phenomenon which allows a text to stand out against the backdrop of the text that conforms with the linguistic norms. The article sheds the light on its origin, the tendency of investigations of the theory of foregrounding as well as a wide range of scientific approaches deployed in an attempt to investigate the notion “foregrounding” and enlarges on what part it plays in generating meaning in a text. In the course of the research the following conclusions were drawn: a text that comprise foregrounded language means stimulates mental work and is perceived to be intricate and thought-provoking by a reader. Foregrounding gives a rise to the creation of a special perception of the object, the creation of a ‘vision’ of it, and not ‘recognition’”. Foregrounding in a language is normally realized through the usage of stylistic devices as well as consistent and systematic nature of actualization. It enables a writer to fulfill writer’s specific aims and motivations through the text.
As the most widely studied and spoken language worldwide, English is a medium of communication between speakers of varied language backgrounds. Global English users need not converge on any one variety; rather, they need adaptive, flexible skills that support communication with speakers of diverse World English (WE) varieties (Kirkpartick 2007). Nevertheless, many English learners and teachers around the world continue to adhere to ‘native speaker’1 models that prioritize language varieties from English-dominant nations, such as the United States or the United Kingdom (Tseng 2019). We have encountered this perspective first-hand in our work in Indonesia. Hanung has taught English and trained teachers at an Islamic University in Central Java for over twenty years. Tabitha has nearly twenty years of experience in ELT, including three years as a visiting instructor at Hanung’s institution. Among the students and teachers we have worked with in Central Java, the dominant language learning model is the ‘native English speaker’. For instance, on a survey given to our first-year English-major students, 81 percent reported wanting to ‘sound like a native speaker’. To challenge this tendency, Simanjuntak and Lien (2021) encourage the use of materials from global and local language communities. Using diverse listening materials can help students understand a wide variety of language varieties and accents and adapt to dialects they have never encountered before. We found podcasts, free audio recordings which are automatically delivered to users’ devices, to be a particularly rich and convenient source of listening materials featuring WE speakers. In the sections below, we describe our teaching context and use of podcasts. We close by discussing how our approach could transfer to other contexts where students and teachers continue to aspire to ‘native speaker’ models, and offer recommendations for teachers interested in shifting to a WE approach and exposing students to varied linguistic norms through podcasts. We collaborated to design lessons for ‘Listening for General Communication’, a course for first-year English majors. The students met weekly from September to December, 2021. Because of COVID-19, the course was taught asynchronously online, with the exception of in-person sessions in weeks 13 and 14. Podcasts were well suited to asynchronous online teaching because their easy accessibility facilitated students’ independent listening. Teachers can create their own podcasts (for a discussion of teacher-created podcasts, see Ingham 2022), but we found several high-quality podcasts online which featured WE speakers. Class activities revolved around seven episodes of the 22.33 podcast, which features participants in exchange programmes sponsored by the US Department of State. We used this podcast because it features speakers with diverse accents, language use, and perspectives, and focuses on intercultural encounters. With the wide variety of podcasts available online, teachers can find podcasts featuring diverse speakers which match their course objectives, contexts, and aims. In Indonesia, English is compulsory at the secondary and tertiary levels, so our students had studied English for approximately six years. Many students at Islamic universities such as ours come from secondary schools whose curriculum focuses primarily on the teaching of Islam. English instruction at these schools focuses on receptive skills, to allow students to access texts about Islam and Muslims written for international audiences. Students typically also study Arabic, but have stronger English skills due to English’s status as the first foreign language in Indonesia, and its use in mass media. The language model used in that media is most commonly standard American English. Students in our English Department take courses in language, linguistics, and teaching methods, and most enter with language skills at the B1 or B2 level. After graduation, most students gain employment as English teachers. At the beginning of the semester, we asked students to reflect upon their future use of English. Although most hoped their own language use would come to match ‘native speaker’ norms, students also acknowledged that they would probably use English to communicate with people using a variety of language norms. They expressed interest in listening to WE speakers and learning about people from around the world. We introduced the podcasts and explained that they featured WE speakers, thereby exposing students to diverse language varieties. In the first asynchronous session for each podcast, students completed pre-listening activities. Podcasts featuring WE speakers typically discussed contexts unfamiliar to our students, so these activities were essential to support comprehension. When students had some relevant prior knowledge, we activated that knowledge through reflection, opinion sharing, and hypothetical situations. When significant content was new to students, short readings, visuals, and online research helped build background knowledge. Taking the Observing Ramadan podcast as an example, students shared their own memories of observing Ramadan, considered which aspects of Ramadan might be surprising to a non-Muslim, and looked up new vocabulary, such as ‘muscle through’ and ‘famished’. Students were expected to complete while-listening activities between the two sessions. A benefit of podcasts is that students can adjust the audio speed and listen to the podcast repeatedly, if needed. To further support students’ comprehension of the WE speakers, we offered instructional scaffolds. For some lessons, students received a graphic organizer to track the various voices and topics. In other lessons, questions were divided into sections corresponding to timestamps in the podcast, helping students know when to listen intensively. For Observing Ramadan, students completed a chart with the names and nationalities of the six speakers by entering each speakers’ perspective on challenges related to fasting and non-Muslims’ awareness of Ramadan. In the second asynchronous session, students completed post-listening activities to build on and apply what they had learnt from listening. These discussions aimed to help students make personal connections and identify universal human values across the WE speakers’ varied contexts. For example, the application questions for Observing Ramadan asked students to reflect on what non-Muslims could learn from experiencing Ramadan. Students responded favourably to the use of podcasts featuring WE speakers, and felt their language skills had improved. One student explained, ‘My understanding of vocabulary and pronunciation increased, without me realizing it.’ On a survey the end of the semester, 89 percent of students said they felt their vocabulary had increased, and 76 percent reported that they appreciated the opportunity to hear good pronunciation. Students were also positively inclined toward the WE language models, which included speakers from Bangladesh, Canada, Ghana, Jordan, Lithuania, Saudi Arabia, the United States, and Yemen. Students said the speakers’ accents and vocabulary were not as difficult to understand as they had anticipated. In fact, some preferred listening to WE speakers, as the following comment reveals: ‘For me personally, listening to speakers who use English as their additional language is more easier than native speaker because we are both learning, so we have some similarities in pronunciation, or use general words.’ Students appreciated the slower speech and simpler vocabulary used by WE speakers. Students especially appreciated podcasts focused on familiar topics. Favourite episodes included Observing Ramadan, an episode about a Muslim prison chaplain in Canada, and one about a journalist who studied in Indonesia. A student explained that she preferred these episodes because ‘the experiences of the interviewees relate to our lives’. We found that the best-received podcasts were those with topics connected to students’ experiences. We see great potential in the use of podcasts featuring WE speakers in other contexts where the ‘native speaker’ model continues to hold power. We have several recommendations based on our experiences. First, we recommend orienting students to the purpose of listening to diverse WE speakers. As mentioned above, most of our students initially wished to sound like a ‘native speaker’, so it was important to offer a rationale for our seemingly contradictory WE paradigm. We believe that helping students consider the importance of engaging with a wide variety of WE speakers resulted in higher motivation and interest. Second, we suggest finding materials that include content related to students’ lives. Although our goal was to expose students to new language varieties and perspectives, we found that students were more interested when at least some of the content was familiar. They needed to be able to make personal connections with the material. When those connections do not come naturally, teachers should use teaching strategies such as guided reflection, discussion, and hypothetical scenarios to help students see how the speakers’ experiences are similar to or different from their own. Lastly, we encourage the use of materials which not only include diverse language varieties but also offer exposure to new cultures. Podcasts from diverse cultural contexts support the development of intercultural competence in addition to language competence. Given its focus on intercultural exchange, the 22.33 podcast was a rich source of such content. Other podcasts with diverse global speakers and intercultural themes include: Global Voices, Voices of Exchange, Rough Translation, The Europeans, and Sound Africa. As other practitioners use these and other podcasts to expose students to WE speakers, we hope they will also share their experiences. Last version received November 2022 Hanung Triyoko is the Head of the Language Development Unit at UIN Salatiga, Indonesia. He is an enthusiast for mutual relationships among educators across the globe and for encouraging his students and colleagues to have everyone’s unique contribution to use language as means of peace and better understanding of life and humankind. He is known as a master trainer and developer of a massive open online course (MOOC) for various levels of English teachers in Indonesia sponsored by RELO of the US Embassy. Email: [email protected] Tabitha Kidwell is a Professorial Lecturer in the TESOL programme at American University, and was previously a visiting lecturer at UIN Salatiga, Indonesia. She teaches academic writing, applied linguistics, and TESOL methods courses, and has conducted professional development for language teachers around the world. Her research focuses on innovative pedagogies, intercultural teaching approaches, and language teacher education. Email: [email protected]
The paper examines team building in multicultural and multiethnic work environments – particularly in NGOs – by linking classic accounts of effective teams with sociolinguistic and socio-emotional perspectives. It conceptualises workplace teams as speech communities in which members share, negotiate, and sometimes contest linguistic norms, expectations, and emotional display rules. Drawing on management and organisational studies, the text outlines key structural conditions for effective teams, including clear goals, appropriate leadership, resource allocation, and mechanisms for accountability and cooperation. These insights are then integrated with research on emotions, culture, and communication, which highlights the role of emotional attachment, trust, and creative problem-solving in sustaining team cohesion. The paper argues that effective team building in such contexts requires not only formal structures but also deliberate cultivation of shared communicative practices and affective climates that support participation, innovation, and mutual understanding across cultural and linguistic differences.
Focusing on recognition of multi-word expressions (MWEs), we address the problem of recording MWEs in WordNet.In fact, not all MWEs recorded in that lexical database could with no doubt be considered as lexicalised (e.g.elements of wordnet taxonomy, quantifier phrases, certain collocations).In this paper, we use a cross-encoder approach to improve our earlier method of distinguishing between lexicalised and non-lexicalised MWEs found in WordNet using custom-designed rulebased and statistical approaches.We achieve F1-measure for the class of lexicalised word combinations close to 80%, easily beating two baselines (random and a majority class one).Language model also proves to be better than a feature-based logistic regression model.
Sociolects, as social varieties of language, should be classified and described on the basis of both linguistic and sociological criteria. In practice, however, most often only either the first one (e.g. according to Antoni Furdal and Danuta Buttler) or the second one is used (e.g. according to Aleksander Wilkoń and Tomasz Piekot). Both these criteria are less frequently combined (Stanisław Grabias). The article proposes those aspects of the sociological and linguistic functioning of language varieties that should be used by sociolinguistics especially in the characterisation of sociolects, but also in their typology. The sociological criteria include: the type of contacts in the group (including contacts made via the Internet), the degree of group formalisation, the durability of the group and the type of bond that binds the group together. Among the linguistic criteria the following were distinguished: nomination and expressiveness, the method of acquiring a sociolect and the degree of its codification and the rank of the linguistic norm. The article also clarifies the concept of the distinguishing features of the sociolectic vocabulary.