Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Replication package for "The Home Bias in Sovereign Ratings" (Journal of the European Economic Association, 2017), by Andreas Fuchs and Kai Gehring. NOTE: uploaded as a single ZIP archive because the package expands to ~3.6 GB. Abstract: Credit rating agencies are frequently criticized for producing biased sovereign ratings. This article discusses how the home country of rating agencies could affect rating decisions as a result of political economy influences and cultural distance. Using data from nine agencies based in six countries, we test whether agencies assign better ratings to their home countries, as well as to countries economically, geopolitically and culturally aligned with them.
The current study examines how the presentation of separable components of rapport building in forensic interviews with children affect lay perceptions of both the child victim and defendant guilt. Mock jurors will read a forensic interview in which a child either alleges sexual abuse by an adult male perpetrator or does not. Interviews will either contain a ground rules phase or only a brief introduction. Thus, the study adheres to a 2 (ground rules: present or absent) x 2 (disclosure: present or absent) between-subjects design). Effects of ground rules on perceptions of defendant guilt are not expected. However, it is expected that disclosure will affect ratings of defendant guilt. It is also expected that both ground rules and disclosure will affect perceptions of the child. It is expected that the child will be perceived more positively than when both are present.
The current study examines the Pakistani English and Indian English newspapers’ discursive construction of the 2025 flood crisis, grounding the analysis within the framework of World Englishes and Critical Discourse Analysis (CDA). Drawing on Fairclough’s three-dimensional model (1989), the research investigates textual features, discursive practices, and socio-cultural contexts to reveal how language mediates disaster narratives in two neighbouring South Asian countries. A sample of thirty news reports from six leading English-language newspapers in Pakistan and India were taken, employing a qualitative, comparative analysis. The findings demonstrate clear divergences in disaster representation. Pakistani English newspapers predominantly frame floods as humanitarian emergencies, employing emotive lexicalization, passive constructions, and crisis-oriented narratives that foreground vulnerability, climate risk, and governance limitations. Indian English newspapers, by contrast, adopt a more procedural and bureaucratic discourse, emphasizing administrative control, technical expertise, and institutional accountability through active agency and policy-focused framing. Despite these differences, both varieties rely heavily on elite institutional sources, marginalizing the voices of affected communities. From a World Englishes perspective, the study shows how Pakistani and Indian English function as localized outer-circle varieties that balance global journalistic norms with national socio-political ideologies. The article contributes to disaster discourse scholarship by highlighting how English, as a shared transnational medium, simultaneously enables cross-border circulation of information and reproduces distinct national identities, power relations, and models of governance in climate crisis reporting.
The task of information retrieval is to find information that best satisfies the user's information needs. In the context of social media, information retrieval is complicated by the high dynamism of content, thematic heterogeneity, and the diversity of users' mental models. This paper proposes an approach to solving the problem of information retrieval under such conditions by constructing a multi-domain dynamic knowledge system. Its novelty lies in the combination of three levels of semantics: problem-oriented, represented by the ontology of the metatask (describing the search objectives); domain-specific, implemented through a dynamic multi-layer knowledge graph built on the basis of user content of social media; and domain-independent, based on a lexical database and a large language model. The knowledge graph allows us to reflect various contexts of concept usage corresponding to thematic clusters in the document collection. Such integration allows us to take into account the evolution of concepts, discourse features, and mental stereotypes of communication participants. To evaluate the effectiveness of the proposed system, an experiment was conducted using a dataset of publications from the VKontakte social network for problem-oriented monitoring of publications, where the selection of relevant publications from non-thematic sources is required. To solve this problem, a technology based on the use of the distance metric between query terms and publication terms in a multi-layer knowledge graph was proposed. The results of the experiment using this technology confirm the effectiveness of the proposed model for information retrieval tasks compared to standard keyword search and embedding models. In continuation of this study, it is planned to create a lexical database and also to consider the possibility of expanding the model by using a measure of pointwise mutual information and graph embedding methods.
Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811 EmoSet images at r=0.636, reaches AUC 0.906 on ESC-50 audio (p<2.2e-15), and AUC 0.720+/-0.055 on EEG from 123 subjects (p<3.65e-8). The direction is mechanistically active: ablating it collapses sentiment accuracy by 5.5-37.2 pp across three LLMs vs. at most 0.88 pp for matched random directions (z>12). A 2-parameter classifier trained on text labels transfers to images (AUC 0.961), audio (0.764), and brain recordings (0.828) without target-modality labels; a generic 16-D subspace stays at chance (0.525). The recipe is bounded to continuous attributes -- seven tests on categorical concepts return near-chance -- and steering is family-specific (Llama/Mistral yes, Qwen/Gemma no).
To better understand the relationship between social anxiety and excessive reassurance seeking, the current experimental study will focus on the interpersonal consequences of reassurance seeking from the reassurance provider's perspective. In the current study, we aim to answer the following research questions: 1. How does excessive reassurance seeking affect the quality of reassurance provided over multiple attempts for reassurance? 2. Does the closeness of the relationship (close other versus acquaintance) moderate the quality of the reassurance provided? 3. How does the reassurance provider’s momentary emotional experience (i.e., Self-Assessment Maniken emotional valence and arousal ratings after each trial) change across successive reassurance attempts, and is this trajectory moderated by relationship condition (close other vs. acquaintance)? 4. What are the emotional and interpersonal consequences of excessive reassurance seeking based on the perspective of the reassurance provider? Specifically, how will providing reassurance multiple times impact the reassurance provider, in terms of their negative and positive affect, their perception of the interaction task, and feelings of relationship closeness with the interaction partner after the task? In general, will the reassurance provider experience emotional and interpersonal strain after providing excessive reassurance? 5. Will emotional and interpersonal strain for the provider be greater when imagining providing reassurance to an acquaintance or a close other? Using an online design, participants are randomized to either an imagined close other or acquaintance condition. In each condition, participants are prompted to provide repeated reassurance to a central character whose reassurance-seeking attempts are scripted by the researchers to align with social anxiety themes. All participants also complete various ratings and survey questionnaires.
This article examines the role of advertisements and signboards in shaping and reflecting public attitudes toward language. In modern society, linguistic culture is not only preserved in literature and education, but also manifested in everyday public texts such as commercial advertisements, street signs, shop names, and information boards. The study analyzes the linguistic quality of advertising texts, the influence of globalization on language use, and the social consequences of neglecting linguistic norms. Special attention is given to the relationship between language accuracy and cultural identity. The article also discusses the responsibility of businesses, media representatives, and educational institutions in maintaining linguistic standards in public communication.
Emotional information often modulates recognition memory, but the relative roles of arousal and valence remain debated, partly because many studies rely on normative emotion categories rather than individualized affective experience. This study examined whether post-test individualized arousal and valence ratings were associated with recognition-confidence responses to emotional Chinese words and whether the learning task influenced later recognition. Forty participants studied neutral, positive, and negative words under semantic-judgment and recognition-judgment learning conditions. After a 24 h delay, they completed a six-point old/new confidence test and then rated each final-test word for valence and arousal. Linear mixed-effects models showed that individualized arousal was more consistently associated with stronger old-response confidence than individualized valence, with a nonlinear increase at higher arousal levels. Semantic-judgment learning was followed by higher old-item confidence and hit probability than recognition-judgment learning. Descriptive signal-detection summaries indicated that valence-category effects were more evident in false-alarm rates, response criterion, and d′ (sensitivity index) estimates than in the primary old-response confidence model. Because affective ratings were collected after recognition, these findings should be interpreted as associative rather than causal, and they highlight the need to distinguish recognition confidence, response bias, and memory sensitivity in emotional memory research.
This paper documents the construction of the Nigerian Salafi Discourse Corpus (NSDC), a purpose-built corpus of 5,791 Facebook posts (2,485,963 words; 2,518,453 tagged tokens) published by Nigerian pages and profiles over a ten-year frame (2015-2025). The NSDC was assembled from eight keyword-filtered exports of the Meta Content Library and processed through a fully documented, machine-verified pipeline executed by Hermes Agent, an agentic artificial-intelligence research assistant. The pipeline performs schema validation, three-pass deduplication, Unicode normalisation and noise removal, SGML-annotated corpus assembly, cross-keyword merging, and calendar-year segmentation. The corpus is annotated in three layers: Penn Treebank part-of-speech and lemma tagging; assignment of tokens to fifteen purpose-built semantic fields; and sentence-level coding of fifteen discourse strategies under the discourse-historical approach. This paper reports aggregate statistics, data-quality challenges and solutions, verification outcomes, a manual reliability audit of the annotation layers, and research applications.
This study employs Critical Discourse Analysis (CDA) to examine how toxic masculinity is discursively constructed, reinforced, and occasionally contested within Men’s Rights subreddits. While existing research has documented the presence of anti-feminist and stoic-masculine ideologies in online male communities, less attention has been paid to the specific linguistic strategies through which members negotiate emotional vulnerability, gender hierarchy, and victimhood. Using a purposive sample of 250 posts and comment threads from r/Mens Rights and r/Left Wing Male Advocates (collected between January and March 2025), the analysis focuses on three discursive strategies: (1) lexical choices framing emotional expression as weakness (e.g., “crying is beta behavior”), (2) presupposition and implication in narratives of male victimization (e.g., “women’s tears are weaponized, men’s tears are invisible”), and (3) intertextual references to pop psychology and evolutionary biology to legitimize emotional suppression. Findings reveal a paradoxical discourse: members explicitly reject “traditional masculinity” as imposed by feminism, yet they reproduce core tenets of hegemonic masculinity, particularly emotional stoicism and the denigration of femininity. However, a minority of threads display resistance, redefining “real men” as emotionally intelligent. The study concludes that Men’s Rights subreddits are not monolithic but operate within a tension between victimized self positioning and the re inscription of patriarchal norms. These findings have implications for online gender politics and mental health interventions.
This corpus-based multifactorial study aims to explore potential predictors of repetition versus lexical variety in translation of repeated reporting verbs from English into Slovak in literary novels. First, we provide a theoretical overview of research on repetition and reporting verbs in Slovak and Czech translation studies. Next, using a sample of 14 literary novels extracted from InterCorp corpus (v.15), we fit multiple negative binomial regression models with mixed effects to assess the effect that several predictor variables (frequency, semantic category, verb length, number of senses, translators) have on the response variable, i.e. the number of Slovak target-text reporting verbs an English source-text (ST) reporting verb is translated into. The findings revealed that factors such as frequency of use of ST reporting verbs, the semantic category of neutral ST reporting verbs, as well as the translators as a random effect, influence Slovak translators’ decisions of using a wide variety of Slovak reporting verbs instead of preserving the originals’ patterns of repetition. More precisely, the model allowed us to explain 70% of the variation (per conditional r-squared) in the response variable. Against the backdrop of prevailing stylistic norms in Slovak, the findings shed light on the translator’s choices in rendering recurrent reporting verbs introducing direct speech, a stylistically salient feature of literary texts.
Шевчук С. Українська мова у сфері міжнародного бізнесу: між локалізацією та англізацією strategies for adapting Anglicisms into Ukrainian linguistic norms, including transcription, calque translation, semantic adaptation, and the creation of new lexical units.The study concludes that effective use of the Ukrainian language in international business requires a balanced approach between localization of global terminology and adherence to linguistic norms.Ukrainian possesses sufficient word-formation potential to develop its own modern business terminology, which contributes to strengthening the national economic space and preserving linguistic identity.
This research is situated at the intersection of digital humanities, the history of emotions, and computational linguistics. The article presents the results of the sentiment analysis of the epistolary heritage and diary entries of three key figures of the Russian monarchy: Catherine II, Alexander I, and Nicholas I. The total corpus of analyzed texts amounted to over 2 million word usages. Using models of deep learning BERT (XLMRoBERTaLarge and Conversational RuBERT) adapted for historical texts, the authors reconstruct the emotional dynamics of communication throughout a turbulent century—from Enlightened absolutism to the crisis of the Nicholas system. The study confirms the hypothesis of a stable correlation between the genre of the document (official letter vs. private diary) and the degree of emotional expressiveness, as well as identifies specific lexical markers of anxiety during periods of political instability (on the eve of the Decembrist revolt and during the Crimean War). Methodologically, the research is based on three conceptual foundations. Firstly, it is the theory of emotional communities by B. Rosenwein, according to which emotions are constructed within social groups with shared values and norms of expression. Secondly, it adopts the semiotic approach of Yu. M. Lotman in studying the everyday behavior of the Russian nobility. Thirdly, it employs methodologies of computational text analysis. The hypothesis of this study is as follows: the emotional tone of the personal correspondence and diaries of Russian monarchs is not so much a spontaneous expression of an individual psychological state, but rather a ritualized social action, subject to the cultural codes of the era and genre canon. The aim of the study is to conduct a comprehensive historical-linguistic analysis of the emotional tone of the epistolary and diary heritage of Russian statesmen of the 18th-19th centuries using digital text processing methods, to identify stable emotional patterns and their connection with historical-biographical context. It has been established that the sentimentalist tradition of the late 18th century paradoxically combined with hypertrophied emotional restraint in official communication, creating an effect of "emotional dissonance," which was resolved in the literature and journalism of the 19th century. This work contributes to the methodology of analyzing historical texts, demonstrating the possibilities and limitations of NLP tools when working with archaic vocabulary and bilingual corpora (Russian-French linguistic dualism).
Abstract Background Something in discourse with a person experiencing psychosis often “feels off” before formal assessment is completed, yet this disturbance has not been quantified at the level of ongoing dyadic conversation. Prior work has largely treated patient speech in isolation, limiting our capacity to measure how communicative disruption emerges within clinical exchange. Methods We applied a three-level decomposition of conversational alignment in 109 patients with psychotic disorders (26 female) and 60 healthy controls (22 female) at baseline and 12 months ( n = 115). Register divergence (dAUC norm ) captured lexical distance between interviewer and patient; embedding-based synchrony (r embed ) measured semantic trajectory coupling; within-speaker coherence was computed separately for each speaker. We used linear mixed-effects models adjusted for timepoint and participant clustering. Results Patients showed significantly greater lexical-semantic divergence from the interviewer ( d = 0.48, p <.001) and reduced embedding-based synchrony ( d = −0.59, p <.001), both effects replicating at each timepoint. Critically, the interviewer’s within-speaker coherence was reduced during conversations with patients ( d = −0.33, p =.016), indicating that the disruption extends beyond the patient to the interaction itself. Register divergence tracked impoverished thinking and synchrony tracked disorganized thinking (both FDR-corrected q =.038). Group differences were persistent at 12 months, indicating a partially stable profile. Conclusions Conversational alignment in psychosis reveals a dyadic failure of semantic coordination that destabilizes the interviewing clinician’s coherence even when patient narrative continuity is preserved. These transcript-derived alignment metrics offer a scalable approach to quantifying interpersonal communicative function from routine clinical encounters.
This article examines the representation of gender roles in Uzbek family vocabulary. Using a qualitative lexical-semantic andsociolinguistic approach, the study analyzes kinship terms and forms of address to identify how language reflects cultural norms and socialexpectations. The findings reveal a clear asymmetry in the representation of male and female roles, where men are associated with authorityand leadership, while women are linked to caregiving and domestic responsibilities. The study highlights the role of language inmaintaining traditional gender structures and suggests that these patterns may evolve under the influence of social change
The intersection of race and language has emerged as a central concern in contemporary sociolinguistics, challenging longstanding assumptions about the autonomy of linguistic analysis from social structures of power and inequality. This article presents a critical theoretical review of raciolinguistics as a framework for understanding the co-constitution of race and language in US sociolinguistic contexts. Drawing on foundational works by Alim, Smitherman, Flores, Rosa, and others, I examine how raciolinguistic perspectives disrupt traditional sociolinguistic approaches that treat race as a demographic variable rather than a dynamic social process enacted through language. The review identifies three key theoretical contributions: (1) the reconceptualization of language practices as racializing projects that construct and maintain racial hierarchies; (2) the critique of "appropriateness" ideologies that position white linguistic norms as neutral standards while rendering racialized speech as deficient; and (3) the insistence on historical and colonial frameworks for understanding contemporary linguistic inequality. I argue that raciolinguistics represents not merely a subfield of sociolinguistics but a fundamental reorientation that demands attention to the material and ideological mechanisms through which race and language become mutually constitutive. The article concludes by considering future directions for research that might extend raciolinguistic inquiry beyond educational contexts to examine labor, healthcare, and other institutional domains where racialized language ideologies operate. This theoretical inquiry contributes to ongoing debates about the scope and significance of raciolinguistic perspectives for sociolinguistic theory and practice.
Abstract We present a study in model comparison and model selection to explore cartographic theoretical proposals and test their generalization ability. Specifically, we test whether an asymmetry arises between stages of French with respect to the merge nature of topicalized elements, as proposed in Wolfe (2022). We explore quantitative and simple computational methods to test the model’s predictions compared to a series of control group models. In particular, we evaluate predictions with respect to intervention locality effects. We explore seven treebanks, representing three stages of French (Old French, Middle French, and contemporary French). Our results partially confirm Wolfe’s (2022) proposal. Additionally, two key elements emerge: one concerning the type of locality involved in movement, and the other relating to a cartographic approach to the syntax of genres/registers.
Introduction. In the present-day scientific discourse, there is a great number of theoretical research focusing on the problem of objectifying the semiotic nature of law and analysing the functional construct of legal semantics. Whereas, many practical legal issues, such as: interpretative ambiguity in the meaning-formation and meaning-application of normative acts, lexical vagueness and contextual dependence of legal notions and the incoherence of legal terminology across different legal systems, remain neglected, which leads to contradictions and inaccuracies in legal practice. The aim of the study is to define the methodological principles fostering establishment of the acceptable scope of semantic interpretation of legal notions in the context of building a legal thinking culture. Materials and Methods. The research methodology was based on the principle of jurisprudential definition of legal norm meaning-formation in socio-legal discourse. Analytical, systematizing and pragmatic methods were used to reveal a complex nature of the semantics of law in the context of legal thinking development. The semiotic analysis of the objectivity and normativity of legal notions taking into account the contextual differences of legal definitions, was used as a specialised research method. Results. It was established that normative notions are the complex semantic constructs encompassing a conceptual sphere (normativity) and social reality. For building sustainable models of legal behaviour and legal culture, it is necessary to overcome external and internal conflicts in interpretation of law. In this regard, a number of advisory measures were proposed aimed at establishing acceptable scope of semantic interpretation: differentiation between the informational nature of prescriptive and descriptive notions, semantic monitoring of legal phenomena, and implementation of the principle of discourse contextualism, which makes it possible to formulate the normativity of law requirements based on the specific contextual interpretations. Discussion and Conclusion. A justified conclusion about possibility of a properly selected semantic toolkit to determine the objectivity of perception of the legal norms and, consequently, to improve the process of building a legal culture was drawn. The main advantage of the principle of discourse contextualism such as conjunction of the semantics and pragmatics of legal notions was identified, which provides a fruitful foundation for further theorizing on the nature and metaphysics of law.
Sentiment analysis can be considered a critical task in the field of Natural Language Processing (NLP). It is applied not only in product reviews but also in areas like healthcare, education, services assessment, and so on. While there have been remarkable advancements in building text-based sentiment models, the problem of the interaction between structured and unstructured data for sentiment analysis has not been addressed adequately. Specifically, current systems use scale ratings and descriptive texts separately without considering how one can complement the other. This chapter provides an in-depth discussion of the dual-channel sentiment model based on the Average Cumulative Rating Matrix Factorization (ACRMF) model presented in Kumar et al. [7]. In addition to utilizing internal review comments, namely structured numerical scale ratings, and external review comments, open-ended descriptive text, the approach creates a common sentiment score based on the normalized averaging method. The mathematical concepts underpinning the framework include Bayesian probability theory, the Naïve Bayes classifier, Part-of-Speech tagging, the Stanford Sentiment Treebank, and Recursive Neural Networks. Three research hypotheses are formulated, which are tested on two real datasets: Tourpedia and Kaggle Travel Review Rating. Experiments are conducted for two different train-test splits, three different cut-off values, and seven baseline methods. Under all experimental conditions, the dual-channel sentiment model outperforms its single-channel counterpart in a statistically significant way. This section is concluded with a limitations analysis and an outline for further research directions.
Sparse Mixture-of-Experts (MoE) routers commonly use the same scores both to select experts and to weight their already-computed outputs. We study whether these two roles, dispatch and aggregation, should be coupled. On pretrained OLMoE-1B-7B, we keep selected Top-8 expert IDs, expert computation, and total selected router mass fixed and change only within-set aggregation. A structured oracle improves full-horizon cross-entropy by 0.0160 +/- 0.0039 across three seeds; the router's top-scored expert is the counterfactual-best vertex only 17.2% of the time, with router-utility Spearman 0.030. We therefore train Fixed-Dispatch Adaptive Aggregation (FDAA), a 301K-parameter post-compute head optimized directly with the language-modeling objective while freezing the backbone, router, and experts. On OLMoE, FDAA improves fresh WikiText-103 test by Delta CE = -0.1523 +/- 0.0031 across three seeds, and mixed-domain training gives robust gains on WikiText-103, C4, and held-out Penn Treebank under frozen confirmatory evaluation. We also replicate the fixed-dispatch audit on DeepSeek-V2-Lite, which uses Top-6 routed experts plus shared experts. Best-vertex headroom remains significant on WikiText and C4, while router Top1 identifies the best selected expert in only 12.5% and 16.7% of audited examples. In a one-seed mixed-domain replication, FDAA improves locked WikiText and PTB, while C4 is statistically neutral. These results support a cross-architecture distinction between expert selection and expert commitment.
Abstract Depression is marked by blunted affective responses to context, which interoceptive accounts trace to altered neural representations of bodily states. Yet this evidence mainly concerns response magnitude, not how quickly affect is updated when contexts change. Here we tested whether depressive symptom severity is related to delayed affective updating, and whether cortical dynamics tracking cardiac states account for this delay. To this end, we applied a movie-watching paradigm with independently defined contextual shifts, continuous affect ratings, electroencephalography, and electrocardiography in individuals spanning a continuum of depressive symptoms. Combining deep representation learning and a dynamical systems framework, we quantified how quickly (speed) and how sharply (angle) cardiac-coupled cortical representations reorganized at each shift. Greater symptom severity predicted longer latency to enter the context-congruent affective state across contextual shifts, regardless of valence. In a cross-sectional mediation analysis, slower speed, but not angle, accounted for this association. This mediation was specific to depressive symptoms, contextual shifts, and cardiac-coupled neural dynamics. These findings extend the embodied account of depression from blunted affective intensity toward its inflexible updating at moments of contextual shift, and offer a broadly applicable framework for quantifying brain-body dynamics across affective dysfunctions.
Social media platforms, particularly TikTok, have become primary arenas for linguistic experimentation among adolescents, yet systematic analyses of how platform-specific affordances shape lexical and semantic innovation remains limited. This study investigated lexical and semantic variations in adolescent digital communication on TikTok, addressing three research questions concerning the types of lexical innovations, processes of semantic change, and the role of platform affordances in shaping language evolution. Methods: A mixed-methods design integrated quantitative corpus linguistics with qualitative discourse analysis. A corpus of 2,848 TikTok comments was compiled across four major trends (September–December 2024). Lexical analysis identified neologisms, graphical variations, and acronyms; semantic analysis documented broadening, narrowing, metaphoric extension, and pejoration/amelioration; platform affordances analysis examined meme-driven language and intertextual policing. Analysis revealed 15 lexical innovations with 63 occurrences across semantic categories. Neologisms (fr, bestie, delulu) and graphical variations (tryna, cuz, ion) served dual functions of efficiency and identity performance. Semantic shifts included ameliorative broadening (slay, fire), pejoration (basic, cringe), metaphoric extension (era, main character), and reclamatory usage (ghetto). Platform analysis identified 11 meme-driven phrases generating 2,848 occurrences with near-neutral sentiment, and 347 policing instances (12.2%) concentrated during rising and peak trend phases, demonstrating active semantic negotiation through definition, debate, and correction. TikTok functions as an accelerated laboratory for language change where adolescents deploy multiple mechanisms of linguistic innovation simultaneously. Platform affordances fundamentally reshape traditional sociolinguistic processes, with intertextual policing serving as the mechanism by which communities enforce emerging semantic norms. The findings extend communities of practice frameworks to algorithmically-mediated digital environments. Educators should recognize digital language as systematic innovation; lexicographers should develop protocols for documenting ephemeral platform-specific terms; platform designers should account for in-group reclamation practices; and researchers should prioritize cross-platform longitudinal studies to track whether observed innovations represent enduring change or age-graded phenomena.
The article provides a comprehensive analysis of inconsistencies and variations observed in the orthographic norms of compound words in the modern Kazakh language. The research material consists of 54 lexical items, including 18 names of animals, 12 names of plants, 14 medical terms, and 10 words representing diverse semantic and morphological models. The study employs comparative analysis, phonetic-pattern analysis, structural-morphological and semantic modeling, as well as a comparative examination of orthographic dictionaries and normative reference sources. The findings reveal that approximately 30% of the compound words under consideration appear in two or more parallel written forms across different orthographic dictionaries and reference publications. Major problematic areas of Kazakh orthography identified in the study include the inconsistent application of vowel harmony rules, the lack of reflection of phonetic assimilation in writing, discrepancies between pronunciation and orthographic representation, the violation of morphological integrity, and the presence of unsystematic spelling patterns in the names of animals, plants, and medical terms. The results underscore the necessity of revising the spelling conventions of compound words in accordance with the internal linguistic laws of Kazakh, its natural phonetic structure, and its agglutinative nature. The conclusions presented in the article hold practical significance for the development of orthographic rules based on the new alphabet, the updating of orthographic dictionaries, and the scientific justification of orthographic directions within state language policy
This article analyzes the regional vocabulary of the Provence region found in regional print media focusing on sports. The subject of the study is the regional lexical units characteristic of the Provence region. The object of the research is the functioning of regional vocabulary in the texts of print media on sports. The author examines in detail how the most significant sporting events in the region are described and which lexical units are used. The author's attention is directed solely to lexical units, as an analysis of language units at other levels (phonetic and grammatical) based on written texts is not possible for a number of reasons: a detailed study of the phonetic features of the regional language requires a corpus of audio and video texts. At the grammatical level, no differences between the literary and regional languages are identified, as the texts of print media are composed in accordance with literary norms, while the researcher's focus is not on colloquial forms but on the spoken language of educated speakers of this region's language. For comparison, texts from the most popular national and regional publications covering the same sporting events were selected and the lexical units used in their descriptions were analyzed. A method of complete sampling and contextual analysis was employed for this purpose. The novelty of this research is due to the fact that texts from print media are a very important source for analyzing the national and cultural features of the regional variant of the French language. Existing lexicographic sources at this stage do not provide reliable information about the functioning of linguistic units in everyday speech. Therefore, studying regional language features based on media material has become relevant. It is not by chance that sports themes were chosen for the analysis of lexical units, as they are the most akin to colloquial speech. The conducted analysis showed that regional print media texts on sports are characterized by a wide integration of regional lexical units. This leads to the conclusion that these lexical units are indeed used in the written language of educated speakers and serve as a cultural marker of the French language variant in the Provence region.
Abstract Cities are experienced in motion, yet urban soundscape research has largely assumed stationary viewpoints, overlooking the perceptual role of the walking perspective. Here, we show that this mismatch fundamentally blinds us to how walking views reshape urban soundscape perception. Using a within-subject, repeated-measures audiovisual design, 34 participants evaluated 18 urban streets under standing-view (SV) and walking-view (WV) conditions paired with identical binaural audio, providing both continuous real-time affective ratings and retrospective soundscape evaluations. We found that, first, rapid perceptual stabilisation occurred: taking the standing view as the baseline, real-time affective responses under the walking view initially diverged but consistently converged within the first 10 s across all participants and streets. Second, despite this early stabilisation, the walking perspective reshaped overall perceptual outcomes by selectively reweighting the salience of urban sounds. Among streets showing significant effects, transient sounds, including alarms and sirens, became more salient, whereas continuous background sounds, including traffic and human activity, became less salient. The walking perspective also polarised overall sound environmental evaluations, making positively evaluated streets more positive and negatively evaluated streets more negative, while reducing inter-individual variability. These findings demonstrate that the walking perspective is an active component of urban soundscape perception, shaping both the temporal dynamics of perceptual adaptation and the overall perceptual weighting of urban sound environments, with broader implications for understanding how environmental perception unfolds in motion.
Background: The accuracy and safety of generating medication orders by large language models (LLMs) must be demonstrated. Without standardization, performance evaluation is limited to time and resource-intensive clinician grading. This evaluation aimed to develop a standardized medication format that supports automated performance evaluation (MedMatch). Methods: First, a survey of 40 medication prompts was given to clinicians to assess agreement in medication order communication. Second, a clinician panel developed a standardized medication format (MedMatch) for oral and intravenous medications. Third, a clinician-annotated dataset of medication prompts and standardized answers in the MedMatch format was developed for LLM testing. Finally, LLMs were retested with the same dataset, adjusted to exclude route information, to evaluate the appropriate categorization of medication route. Results: The formal medication orders consistently showed low omission rates and high overlap for all entities, compared to the verbal and brief written communication types. Lexical overlap results demonstrated pattern norms amongst clinicians with entities appearing most commonly in positions 1-5 in the order of drug name, dose, unit, route, and frequency. In the second survey, the formal written group performed the highest with 78.3% of prompts considered appropriate as a computer-generated response. LLM accuracy on MedMatch order standardization was highest in oral solid (64.2-72.5%), intravenous intermittent (72.5-84.3%), and intravenous push (62.7-74.5%) categories. LLMs performed the worst at categorizing medication orders accurately into intravenous push (18-61%) and intravenous intermittent (51-100%) routes. Conclusions: Standardized format for computer-based outputs may support automated performance analysis and enhance the clarity of medication communication.
The article is devoted to the theoretical substantiation of the essence of the grammatical aspect of foreign language speech as a key component of foreign language communicative competence. The relevance of the study is determined by the need to understand the structural content of the grammatical aspect of speech in the context of the requirements of modern educational standards. The paper presents a comparative analysis of the approaches of foreign and Russian researchers to understanding the grammatical aspect of speech. D. Larsen-Freeman's three-dimensional model, which includes form, meaning, and use of grammatical phenomena, is examined and illustrated with examples. The position of S. Thornbury, who defines grammar through morphology and syntax and emphasizes its meaning-making potential realized in representational and interpersonal functions, is analyzed. Attention is paid to M. Lewis's lexical approach, which assigns a secondary role to grammar. The article presents the views of Russian methodologists (N.D. Galskova, N.I. Gez, E.N. Solovova), who consider grammar as a fundamental component of speech activity that ensures practical language proficiency for solving communicative tasks. Based on the analysis conducted, the author formulates a definition of the grammatical aspect of foreign language speech as a complex of automated actions for selecting, combining, and using grammatical structures in accordance with communicative intention and language norms. It is concluded that an insufficient level of mastery of the grammatical aspect of speech leads to difficulties in the formation of foreign language communicative competence as a whole.
Transitivity is a lexico-grammatical system for construing experiential meaning through both grammatical structure and lexical choice. This study investigates how the experiential meaning is reproduced in translating cultural imagery in allusions in Chinese political texts, with particular attention given to its structural realization. Drawing on a parallel corpus of 449 instances from Xi Jinping: The Governance of China (Volumes I to IV), it examines how imagery, in translation, is grammatically realized as participant or circumstance through different processes, together with lexical choices. A strong preference is observed for structural equivalence, especially when the imagery conveys universal or ideological significance. Shifts in structural role or process type are employed to clarify abstract reasoning, foreground general principles, or facilitate intercultural understanding. Material and relational processes are predominant, and shifts often signal rhetorical movement between action and evaluation. As to lexical choices, a strong preference is also observed for retaining the original image rather than rendering its sense or omitting it. These results indicate that the selection of imagery reflects the diplomatic principle of proximity, and its retention in translation aligns with harmony-oriented institutional norms. This study offers a replicable model for analysing cultural imagery in political translation from a transitivity perspective.
The recitation of the Qurʾān occupies a central position in the ritual and devotional life of Muslims. In Arabic, this practice is commonly referred to as tilāwat al-Qurʾān and its origins go back to the time of the Prophet Muḥammad. Qurʾānic recitation is a required component of the five daily prayers (ṣalāt), with both Sunni and Shiʿi Muslims regularly reciting al-Fātiḥah, the opening chapter of the Qurʾān, along with additional verses or chapters. Today, most Muslim communities recite the Qurʾān according to the reading attributed to ʿĀṣim ibn Abī al-Najūd (d. 744), as transmitted by Ḥafṣ ibn Sulaymān al-Asadī (d. 796). However, within the Sunni qirāʾāt tradition, there are ten canonical readings, including that of ʿĀṣim. Each of these readings is known as qirāʾah (plural: qirāʾāt). In this entry, "canonical" refers to the ten qirāʾāt recognized within the Sunni qirāʾāt tradition. This canonical status, however, was not established all at once. The seven readings first achieved it through Ibn Mujāhid's selection, while the additional three attained canonical status through later scholarship, especially the works of Ibn al-Jazarī. The term "shawādhdh" (singular: "shādhdh"), here used broadly for non-canonical readings, denotes readings that fall outside these ten, whether because they failed to meet one or more of the established criteria for acceptance, namely a sound chain of transmission (isnād), conformity with the orthography of the ʿUthmānic codices, and adherence to Arabic linguistic norms, or because they did not attain wide recognition among qirāʾāt scholars. The seven qirāʾāt achieved canonical status and became central to Qurʾānic recitation after Ibn Mujāhid (d. 936) selected them, in al-Sabʿa fī al-qirāʾāt, from among the numerous variant readings circulating in his time. The compilations of 25 readings by Abū ʿUbayd al-Qāsim ibn Sallām (d. 838) and 22 readings by Ibn Jarīr al-Ṭabarī (d. 923) attest to the significant plurality of variant readings circulating in the ninth and tenth centuries. Ibn Mujāhid selected seven qirāʾāt from Mecca, Medina, Basra, Damascus, and Kufa on the basis of transmission, consistency with the ʿUthmānic codices, conformity to Arabic linguistic norms, and broad recognition among qirāʾāt scholars. His decision to exclude less well-known readings from his compilation significantly shaped how later scholars approached the study and recitation of other variant readings. For instance, some scholars, such as Ibn ʿAṭiyya (d. 1147) and al-Nawawī (d. 1277), explicitly stated that readings outside the seven qirāʾāt constitute non-canonical readings and are impermissible for recitation during the daily prayers, although other scholars continued to recognize the validity of additional readings that later came to be counted among the canonical readings. Ibn Jinnī (d. 1002) similarly categorized readings beyond the seven as non-canonical, noting that he refrained from reciting them to prevent their dissemination among the broader community. In his al-Fihrist, Ibn al-Nadīm (d. 995) also classifies the readings into two main groups: the seven readings and non-canonical readings. Even though Ibn Mujāhid pioneered the canonization of the seven readings, it was the Andalusian scholars who played a key role in consolidating and disseminating the seven qirāʾāt across broad regions of the Islamic world through their works on the seven readings. For instance, most of the works on the seven qirāʾāt referenced by Ibn al-Jazarī in the introduction to his famous work al-Nashr were authored by Andalusian scholars. In particular, al-Taysīr by al-Dānī (d. 1053) and al-Ḥirz al-Amānī (also known as al-Shāṭibiyya) by al-Shāṭibī (d. 1194) have remained foundational texts in qirāʾāt education on the seven readings to the present day. Al-Sakhāwī (d. 1245), a student of al-Shāṭibī, and his commentary on al-Ḥirz al-Amānī were instrumental in establishing al-Ḥirz al-Amānī's reputation among qirāʾāt scholars. Regarding the prominence these works achieved among qirāʾāt scholars, al-Mashīnī, in his Madrasat al-Tafsīr fī al-Andalus, a work devoted to Andalusian contributions to Qurʾānic scholarship, observes that upon examining biographical entries in Ibn al-Jazarī's Ṭabaqāt al-Qurrāʾ (Biographies of qirāʾāt Scholars), one frequently encounters phrases such as "He read al-Taysīr with so-and-so" or "He memorized the Shāṭibiyya," from which al-Mashīnī concludes that al-Andalus occupied a particularly prominent position in the field of qirāʾāt. In light of the widespread influence of al-Shāṭibiyya and al-Taysīr, Ibn al-Jazarī explained that one reason for composing his work on the ten qirāʾāt was that people had almost come to believe that the only accepted readings were the seven contained in al-Shāṭibiyya and al-Taysīr. Ibn al-Jazarī was not the first scholar to write on the ten qirāʾāt. More than ten scholars from the tenth to the fifteenth century, including Ibn Mihrān al-Iṣbahānī (d. 992) and Ibn Siwār al-Baghdādī (d. 1103), composed works on the ten readings. However, these works did not succeed in establishing the study of the ten qirāʾāt as a widely adopted standard in qirāʾāt curricula in the way that Ibn al-Jazarī's work did from the fifteenth century onward. Ibn al-Jazarī recounts that when ʿAbdallāh b. ʿAbd al-Muʾmin al-Wāsiṭī (d. 1339) traveled to Damascus intending to teach the ten readings, local qirāʾāt scholars sought to prevent him through judicial intervention on the grounds that he sought to teach readings beyond those found in al-Shāṭibiyya and al-Taysīr. The selection of the seven readings by Ibn Mujāhid in the tenth century, combined with the Andalusian scholars' works on the seven qirāʾāt, significantly strengthened the preference for the seven qirāʾāt in both Qurʾānic recitation and qirāʾāt education. This preference extended across Muslim laypeople as well as scholars from the tenth to the fifteenth centuries. Nevertheless, throughout this period, a number of scholars continued to transmit and defend the additional three readings alongside the seven.
Some recent scientific studies in the field of linguistics and philology increasingly suggest that English cannot be considered a completely homogeneous system.Moreover, its main national varieties -British, American, Canadian and Australian -are characterised by stable and systematic differences at the grammatical, morphological and syntactic levels.This article examines the regional patterns of English usage.The article synthesizes the main structural differences and trends in grammatical development, demonstrating that these variations constitute distinct norms of usage that are crucial for accurate interpretation and translation.Numerous scholars have conducted scientific research on the patterns of English language integration from various analytical perspectives, focusing on features that are particularly important for contemporary descriptive and comparative grammatical studies in specific regional contexts.In particular, scholars from Great Britain have often set standards for English grammar, and contemporary works are usually descriptive, drawing heavily on corpus linguistics.This article focuses on a comparative analysis of grammatical differences between the four main national varieties of English: British English (BrE), American English (AmE), Canadian English (CanE) and Australian English (AusE) through the prism of practical application.It focuses on key aspects of grammatical variation, including tense and case preferences, collective noun agreement, irregular verb morphology, modal and auxiliary verb usage, and fixed prepositional constructions.Drawing on descriptive grammar and corpus research, the article argues that these differences are neither accidental nor stylistic anomalies, but reflect deeper historical, functional, and sociolinguistic processes shaping modern English.The article also discusses the practical implications of grammatical variation for translators and interpreters, emphasising the need for grammatical localization alongside lexical choice.
Social interactions are dynamic and complex, relying on tracking variation in your own and your partner’s affective and mental states. Being “in-sync” neurally and building a shared consensus can mark successful social interactions. In addition, social anxiety can impact social experiences. We used a novel, naturalistic paradigm to investigate relations between inter-brain neural similarity within mentalizing regions, affective similarity and social anxiety symptoms. Undergraduate student friend pairs (N = 34, 85% White, 65% Women) engaged in a social interaction while videos previously captured from their individual perspectives were recorded. Participants watched clips of the social interaction from both their own perspective (their personal view of the social interaction) and their friend’s perspective (their friend’s personal view of the social interaction) while fMRI data were collected. They rated their affect after each clip. Inter-brain neural similarity was computed across three conditions: 1) Same-Stimuli: both participants viewed identical visual stimuli as in a traditional neural similarity paradigm, 2) Self-Perspective: both participants viewed the clip from their own perspective like the originally experienced social interaction and 3) Friend-Perspective: both participants viewed the clip from their friend’s perspective, a novel perspective. Participants self-reported their social anxiety symptoms and affective similarity captured affect rating concordance within dyads. In contrast to the Same-Stimuli condition, when participants viewed the clips like they originally experienced social interaction (Self-Perspective), greater affective similarity was associated with greater inter-brain neural similarity. When participants viewed the clips from a novel perspective (Friend-Perspective), participants lower in social anxiety symptoms exhibited greater inter-brain neural similarity with greater affective similarity; whereas, participants higher in social anxiety symptoms exhibited greater inter-brain neural similarity with less affective similarity. The results suggest stimuli from socially relevant perspectives, rather than identical stimuli, may reveal more nuanced brain-behavior dynamics, allowing for a better understanding of individual differences in socioemotional experience.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic.
this article examines the sociolinguistic and normative aspects of loanwords in contemporary Korean. In the context of globalization and digital communication, borrowed vocabulary has become an integral part of everyday language use, particularly in media, technology, and youth discourse. The study analyzes the processes of phonological and orthographic adaptation of foreign lexical items, as well as the challenges of their standardization. Special attention is paid to the regulatory role of the National Institute of the Korean Language in establishing transcription norms and maintaining linguistic consistency. The paper also explores issues of semantic shift, hybrid word formation, and variation in spelling across digital platforms. The findings highlight the need for a balanced language policy that preserves linguistic identity while accommodating global lexical influence.
This repository contains the datasets used to reproduce the experiments presented in the work Early Language Learning via Spreading Activation and Lexical Category Exploration in Complex Networks (to appear on arXiv). The repository includes preprocessed datasets for four languages: German (de), English (en), Dutch (nl), and Rioplatense Spanish (sp). Network Structure The files free_de, free_en, free_nl, and free_sp are derived from the association norms provided by the Small World of Words (SWOW) project. Specifically, we consider the releases SWOW-DE25, SWOW-EN18, SWOW-NL13, and SWOW-RP22, respectively. The files syn_ant_hier_de, syn_ant_hier_en, syn_ant_hier_nl, and syn_ant_hier_sp contain semantic relations corresponding to synonymy, antonymy, and hypo/hypernymy relations, derived from WordNet resources: NLTK WordNet for English, Dutch, and Spanish, and Open German WordNet. The files phon_de, phon_en, phon_nl, and phon_sp contain phonological relations. Phonetic transcriptions are obtained using the CMUdict library for English and the epitran Python package for German, Dutch, and Spanish. Edges are introduced between pairs of words whose phoneme strings have an edit distance d<=2. The files graph_de, graph_en, graph_nl, and graph_sp are the effective network representations obtained by aggregating the files described above, and directly used in all experiments presented. Ground-Truth Validation Orderings, Age-of-Acquisition values, and lexical categories (in the form of CDIs) are derived from the Wordbank database, which is based on the MacArthur-Bates Communicative Development Inventories (CDIs).
Amid ongoing academic debates about the status of China English (CE) within global Englishes scholarship, this study examines the attitudes of 302 Chinese pre-service English teachers through a sequential explanatory mixed-methods design. Data were collected via a survey, semi-structured interviews with 22 participants, and analysis of open-ended responses. Quantitative findings showed that CE was widely perceived as intelligible (88.8%) but only moderately acceptable (around 60%), with syntactic features rated more positively than lexical and discourse-pragmatic ones. Qualitative analyses further revealed a “legitimacy paradox”: many participants recognized CE’s cultural value and role in reflecting local identity, yet most were reluctant to incorporate CE features into classroom practice due to institutional constraints such as standardized testing, limited codification, and perceived professional risks. These findings demonstrate how pre-service teachers’ evaluations of CE are shaped by structural ideologies in Expanding Circle contexts, underscoring the need for teacher education programs to foster critical language awareness and to develop pedagogical approaches that balance global intelligibility with local identity representation.
This repository contains Anomaly Soul Kit, an open simulation framework for observing the emergence, persistence, and evolutionary inheritance of anomalous behavior in populations of LLM-driven agents. Each agent encodes a numeric internal state — vitality (H) and anomaly intensity (Z) — and expresses that state through LLM-generated text each generation. A detection layer scores each expression against the population across three axes: lexical divergence, structural divergence, and novel vocabulary. Agents whose expressions deviate from the population accumulate anomaly intensity, which feeds back into their fitness and is heritable across generations. The project does not claim these anomalies constitute mind or soul. It provides a reproducible kit for observing whether something — a persistent, evolving deviation — reliably emerges from this process, and what it looks like when it does. --- Update — February 2026 v2 of the anomaly detection layer has been released. Two structural issues identified in early testing have been addressed. First, anomaly score inflation: as the population evolved, an increasing proportion of agents were flagged as anomalous, eventually making the designation meaningless. This has been resolved by replacing absolute scoring with a dynamic baseline — scores are now normalized relative to the population median each generation, making it structurally impossible for the entire population to simultaneously score as anomalous. Second, convergence speed: the original selection pressure caused Z-awakening to saturate too quickly (~90% by generation 50). Scaling has been adjusted to allow slower, more observable divergence dynamics. Two new observational metrics have been added: new_normal_threshold tracks whether what was previously anomalous is becoming the new collective norm, and population_drift measures how much the group as a whole is shifting toward anomalous expression across generations.
This article provides a thorough examination of the critical role that intercultural pragmatic competence plays in contemporary English language instruction. This sophisticated construct extends beyond traditional linguistic knowledge to encompass the nuanced understanding of how language functions within diverse cultural frameworks to convey meaning, intent, and social relationships. Contemporary English Language Teaching (ELT) methodologies have undergone a significant paradigmatic transformation, characterized by growing acknowledgment of the complex interdependence between linguistic structures, communicative intentions, and the sociocultural contexts that shape their interpretation. This comprehensive perspective deliberately moves beyond conventional pedagogical approaches that prioritized grammatical accuracy and lexical acquisition in relative isolation. Rather, it actively promotes a more profound comprehension of target cultures, recognizing that successful communication depends substantially on understanding culturally conditioned expectations regarding appropriateness, politeness, and discourse organization. Central to this evolving pedagogical framework is the systematic integration of communicative language teaching principles. This approach provides substantial theoretical foundations for investigating how cultural norms, social conventions, and contextual factors fundamentally influence language learners’ interpretation and production of meaning in authentic communicative situations. Ultimately, the findings presented herein compellingly demonstrate the imperative of equipping language learners not merely with structural accuracy and lexical diversity, but fundamentally with the pragmatic awareness essential for genuinely effective, contextually appropriate, and mutually comprehensible cross-cultural communication, thereby enabling them to navigate the complexities of international discourse with competence and cultural sensitivity.
Multimodal resources for emotion expression analysis in pediatric clinical populations remain limited, particularly for children with Tourette syndrome (TS). MindTS-MMD was developed as a Chinese multimodal dataset to support computational research on emotional expression and tic-related behavior in this population. The dataset contains 10,034 instance-level samples from 60 children with TS aged 6–12 years, collected through semi-structured emotion-elicitation tasks. Each sample includes available symbolic representations from three modalities—visual, acoustic, and semantic—and is linked to a corresponding annotation record. The annotations cover seven categories: anxious, calm, focused, irritable, relaxed, shy, and tense, together with child-reported and experimenter-observed valence–arousal ratings and tic occurrence, anatomical location, and frequency when observable. Annotation reliability, signal quality, audiovisual synchronization, facial tracking, modality completeness, and tic–emotion co-occurrence were evaluated. Unimodal and multimodal baselines are provided to illustrate the computational usability of the released representations.
Studies of emotion often rely on standardized stimulus sets to elicit affective responses. Although established databases provide images with normative valence and arousal ratings, selecting suitable stimuli can be difficult when experiments require specific thematic or content constraints. This challenge is especially pronounced for negative stimuli, which are central to research on maladaptive emotions and behaviors in clinical contexts but are often scarce in necessary quantity or specificity. The present study evaluated the feasibility of using generative AI, specifically text-to-image generators, to create tailored negative and neutral affective stimuli. To assess whether these images can serve as alternatives to traditional stimuli, we compared their affective properties to those reported in standardized image databases. Across two studies, participants rated the valence and arousal of 160 and 200 AI-generated images. Our findings revealed that AI-generated negative and neutral images reproduced the characteristic inverse association between valence and arousal observed in standardized databases, with moderate to strong correlations between these dimensions. These results highlight the potential of generative AI as a practical methodological tool for creating customized affective stimuli aligned with specific research objectives and experimental designs.
Follow-up of an earlier study investigating the link between cognitive impairment and long-term speech recognition in noise. Sentence recognition in noise (MBAA2 test, SNR50); cognition (Montreal Cognitive Assessment MoCA, Victoria Stroop, and ECLA 16 + phonological awareness subtests 3 and 4); subjective benefit (Glasgow Benefit Inventory, GBI); and approaches to speech therapy were assessed three years post implantation in 26 adult subjects. Average MoCA scores were below population norms, with seven subjects scoring below 22: the normal threshold set for the study. Average Stroop and ECLA scores were within the normal range. Poor MoCA performance and poor verbal fluency were linked to poorer sentence-in-noise scores. Subjects with MoCA scores < 22 or verbal fluency scores < 6 were fifteen times more likely to have long-term SNR50 > 5 dB. GBI scores were significantly correlated with SNR50s. Cognitive scores remained stable over time, except for completion times for ECLA 3 and 4, which improved significantly. Speech therapy with cognitive goals showed better outcomes in a small subgroup. The speed of completing the reading tests was improved after prolonged implant use which may indicate improved processing speed, better lexical mapping, and improved working memory. An improved decision tree is proposed to guide rehabilitation in underperformers. The results confirmed that after three years of implant use, top-down cognitive deficits still had an impact on sentence recognition in noise, with poorer scores related to lower patient reported benefit. There may be some speech recognition benefits to targeting speech therapy towards objectives in both perceptual and cognitive domains, but more data is required.
Ο γλωσσικός πόρος san-Corpus περιλαμβάνει σώμα κειμένων γραπτού λόγου της Νέας Ελληνικής, έκτασης περίπου 9 εκατομμυρίων λέξεων. Ο πόρος συγκροτήθηκε στο πλαίσιο διδακτορικής διατριβής, με στόχο τη μελέτη των συγκρίσεων ομοιότητας στη Νέα Ελληνική. Μέγεθος & Πηγές Το corpus περιλαμβάνει τρία ισομεγέθη υποσώματα, ώστε να επιτρέπονται οι συγκρίσεις μεταξύ τους: (α) Δημοσιογραφικός Λόγος: 7.774 άρθρα από τέσσερις διαδικτυακές εφημερίδες (Η ΑΥΓΗ, Η ΚΑΘΗΜΕΡΙΝΗ, ΕΘΝΟΣ, ΤΟ ΒΗΜΑ - έτος 2015). Συνολική Έκταση: 2,9 εκατ. λέξεις. (β) Εκπαιδευτικός Λόγος: 96 σχολικά εγχειρίδια δημοτικού και γυμνασίου. Συνολική Έκταση: 3,4 εκατ. λέξεις (μελετώνται 2,8 εκατ.). (γ) Λογοτεχνικός Λόγος: 28 μυθιστορήματα (βραβεία αναγνωσιμότητας περιόδου 2010-2015). Συνολική Έκταση: 2,5 εκατ. λέξεις. Κατανομή Σχολικών Εγχειριδίων ανά Γνωστικό Αντικείμενο Πλήθος Εγχειριδίων Ελληνική Λογοτεχνία 13 Ελληνική Γλώσσα 12 Ιστορία 9 Φυσική – Χημεία – Βιολογία 9 Μαθηματικά 9 Γεωγραφία – Γεωλογία – Περιβάλλον 8 Θρησκευτικά 7 Αγωγή Αισθητική (Εικαστικά – Μουσική – Θέατρο) 15 Αγωγή Υγείας (Φυσική Αγωγή – Οικιακή Οικονομία) 5 Πληροφορική – Τεχνολογία 5 Αγωγή Κοινωνική – Πολιτική 3 Αγωγή Σταδιοδρομίας (ΣΕΠ) 1 Κατάλογος μυθιστορημάτων: Συγγραφέας, Τίτλος Έτος 1ης έκδοσης Δούκα, Μάρω - Το δίκιο είναι ζόρικο πολύ 2010 Θέμελης, Νίκος - Η συμφωνία των ονείρων 2010 Καρυστιάνη, Ιωάννα - Τα σακιά 2010 Μιχαλοπούλου, Αμάντα - Πώς να κρυφτείς 2010 Ελευθερίου, Μάνος - Πριν απ' το ηλιοβασίλεμα 2011 Ζουργός, Ισίδωρος - Ανεμώλια 2011 Μακριδάκης, Γιάννης - Η άλωση της Κωνσταντίας 2011 Μπουραζοπούλου, Ιωάννα - Η ενοχή της αθωότητας 2011 Πανσέληνος, Αλέξης - Σκοτεινές επιγραφές 2011 Παπαδημητρίου, Χίλντα - Για μια χούφτα βινύλια 2011 Παπαθεοδώρου, Θοδωρής - Οι καιροί της μνήμης 2011 Τριανταφύλλου, Σώτη - Για την αγάπη της γεωμετρίας 2011 Φακίνος, Μιχάλης - Η έρημος έρχεται 2011 Βαμβουνάκη, Μάρω - Κυριακή απόγευμα στη Βιέννη 2012 Διβάνη, Λένα - Εγώ, ο Ζάχος Ζάχαρης 2012 Στεφανάκης, Δημήτρης - Φιλμ νουάρ 2012 Ακρίβος, Κώστας - Αλλάζει πουκάμισο το φίδι 2013 Ζέη, Άλκη - Με μολύβι φάμπερ νούμερο δύο 2013 Κορτώ, Αύγουστος - Το βιβλίο της Κατερίνας 2013 Κωνσταντούρου, Μαρία - Αγεφύρωτες σιωπές 2013 Μαντά, Λένα - Με λένε Ντάτα 2013 Ξανθούλης, Γιάννης - Κωνσταντινούπολη των ασεβών μου φόβων 2013 Ρώσση–Ζαΐρη, Ρένα - Άρωμα βανίλιας 2013 Ανδρουλάκης, Μίμης - Αλλέγκρα 2014 Δημουλίδου, Χρυσηίδα - Το κελάρι της ντροπής 2014 Παπαδοπούλου, Ελισάβετ - Μέρες και νύχτες που δεν ήταν δικές μας 2014 Χατζή, Αθηνά - Η θάλασσα έφυγε 2014 Χωμενίδης, Χρήστος - Νίκη 2014 Τεχνικές προδιαγραφές & Μορφότυπος Για την αναπαράσταση των δεδομένων και των μεταδεδομένων υιοθετήθηκε η πολυεπίπεδη οπτική των XML σχημάτων και τροποποιήθηκε το διεθνές πρότυπο TEI P5, 4.0.0 (Text Encoding Initiative). Δημιουργήθηκε ειδικός χώρος ονομάτων sanCorpus (sanC) με σχήμα τύπου RELAX-NG. Το σώμα κειμένων διατίθεται σε TXT και σε XML σε τρεις εκδοχές: Βάθος 0: απλό κείμενο (TXT). Περιλαμβάνει το main core (κείμενο βάσει του οποίου εξετάζονται οι συγκρίσεις ομοιότητας) και το out of core (κείμενο εκτός εμβέλειας της διατριβής, στο οποίο περιλαμβάνονται κείμενα που πλαισιώνουν το κυρίως κείμενο, π.χ. κείμενα διδασκαλίας, πίνακες περιεχομένων, εξώφυλλα) Βάθος 1 = κείμενα στην απλούστερη δυνατή XML κωδικοποίηση Βάθος 2 = κείμενα με πιο λεπτομερείς XML κωδικοποιήσεις. Αυτή η έκδοση (san-Corpus v1.0, Depth 0: Plain Text) περιλαμβάνει το σώμα κειμένων σε μορφή απλού κειμένου (Βάθος 0) στην αρχική του διάταξη (βλ. Επεξεργασία). Στόχος είναι ο σταδιακός εμπλουτισμός με επισημειωμένα δεδομένα, καθώς και με τις εκδοχές Βάθους 1 και 2. Επεξεργασία (Processing) Η μεθοδολογία συλλογής των δεδομένων, η θεωρητική τεκμηρίωση και το σχήμα επισημείωσης επεξηγούνται στις μελέτες Αφεντουλίδου (2022, 2021, 2013, 2012) και Afentoulidou (2009). Η πρώτη εκδοχή (Βάθος 0) χρησιμοποιήθηκε αποκλειστικά για τη λημματοποίηση που απαιτούσε η Collostruction Analysis (Gries, 2024). Στάδια επεξεργασίας (για τη λημματοποίηση): Τμηματοποίηση σε προτάσεις (sentence segmentation) με τη χρήση της βιβλιοθήκης Stanza (Stanford NLP Group, Qi et al. 2020), η οποία βασίζεται στο μοντέλο Greek Dependency Treebank (GDT) του Ινστιτούτου Επεξεργασίας του Λόγου / ΕΚ «Αθηνά». Τυχαία αναδιάταξη των προτάσεων για την προστασία της ακεραιότητας των πρωτότυπων έργων. Λημματοποίηση με τον ILSP Lemmatizer μέσω της Υποδομής Clarin-EL. Για την ανάλυση συμφράσεων του δείκτη σαν απομονώθηκαν συγκεκριμένοι λεκτικοί τύποι. Αδειοδότηση & Δικαιώματα Το san-Corpus συγκροτήθηκε για τις ανάγκες της διδακτορικής διατριβής και προστατεύεται από το δικαίωμα ειδικής φύσης σύμφωνα με την Οδηγία 96/9/ΕΟΚ και το άρθρο 45Α του Ν. 2121/1993. Η χρήση του περιεχομένου γίνεται αποκλειστικά για ερευνητικούς σκοπούς βάσει των εξαιρέσεων της Οδηγίας 2001/29 και της Οδηγίας (ΕΕ) 2019/790 (Text and Data Mining exceptions / Fair Use). Η πρόσβαση είναι περιορισμένη (Restricted Access) και παρέχεται αποκλειστικά σε μέλη της ακαδημαϊκής κοινότητας για σκοπούς επαλήθευσης των αποτελεσμάτων της διατριβής και περαιτέρω μη εμπορική έρευνα. Η πηγή προέλευσης δικαιούται να ζητήσει οποιαδήποτε τροποποιητική ενέργεια (π.χ. αφαίρεση) επί του πρωτότυπου περιεχομένου. Τέλος, η άδεια CC BY-NC-ND 4.0 ισχύει για την επιμέλεια (curation), τα μεταδεδομένα και τη γλωσσολογική επισημείωση του σώματος κειμένων. Βιβλιογραφικές αναφορές Αφεντουλίδου, Β. (2022). Σώμα ελληνικών κειμένων για τη μελέτη δομών ομοιότητας της Νέας Ελληνικής: σχεδιασμός και υλοποίηση. Στο Πρακτικά του 10ου Συνεδρίου Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας (σσ. 67-94). ΕΚΠΑ. Αφεντουλίδου, B. (2021). Δομές ομοιότητας στη Νέα Ελληνική. Σωματοκειμενικές παρατηρήσεις για τον πολυλειτουργικό δείκτη σαν. Προφορική ανακοίνωση στην 41η Ετήσια Συνάντηση του Τομέα Γλωσσολογίας, 13–15 Μαΐου 2021. ΑΠΘ. Αφεντουλίδου, Β. (2013). Και σου απάντησα κάτι σαν ‘τέλεια, εντάξει’. Δείκτης σαν + ευθύς λόγος;. Προφορική ανακοίνωση στο 7ο Συνέδριο Μεταπτυχιακών Φοιτητών και Υποψηφίων Διδακτόρων του Τμήματος Φιλολογίας, 16–18 Μαΐου. ΕΚΠΑ. Αφεντουλίδου, Β. (2012). Συγκρίσεις ομοιότητας στα Νέα Ελληνικά: ο δείκτης σαν. Στο Z. Gavriilidou, A. Efthymiou, E. Thomadaki & P. Kambakis-Vougiouklis (Επιμ.), Selected papers of the 10th International Conference on Greek Linguistics (σσ. 696-707). DUTH. Afentoulidou, V. (2009). Sketching the σαν conditional construction in Modern Greek. Submitted essay, 2009 Linguistic Institute, Linguistic Structure and Language Ecologies, Linguistic Society of America and UC Berkeley. Gries, Stefan Th. 2024. Coll.analysis 4.1. A script for R to compute perform collostructional analyses. https://www.stgries.info/teaching/groningen/index.html Institute for Language and Speech Processing - Athena Research Center (2015). ILSP Lemmatizer. Version 1. [Software (Tool/Service)]. CLARIN:EL. http://hdl.handle.net/11500/ATHENA-0000-0000-23EE-D Qi, P., Zhang, Y., Zhang, Y., Bolton, J., & Manning, C. D. (2020). Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations (pp. 101–108). Online: Association for Computational Linguistics.
Abstract This study examines clitic placement and the use of contracted forms in 13th-century medieval Spanish, focusing on the works of Gonzalo de Berceo. The results indicate that, although proclisis was the predominant norm, unexpected enclitic usages appear, especially in initial position and after metrical caesuras. These occurrences suggest that Spanish at the time was in a transitional stage regarding the fixation of clitic position. Metrical patterns and parallel syntactic structures also influenced clitic distribution, revealing an interaction between linguistic norms and stylistic constraints. Furthermore, the analysis identifies contracted forms affecting not only third-person pronouns but also first- and second-person forms, although their usage was not fully standardized. These findings support the view that the evolution of clitics in Spanish was a gradual process, shaped by phonological, syntactic, and discourse-related factors, with poetic composition offering a space in which grammatical rules and artistic expression coexisted and interacted.
The goal of this research is to understand the relationship between participants’ affective state and their attentional breadth. There are competing theories regarding how different aspects of affective state (valence versus motivational intensity) might influence attentional breadth. A common approach to testing these models has been to attempt to experimentally induce different affective states. However, there are now multiple studies in the literature that have failed to replicate any influence of such experimentally induced affective state on attentional breadth. This could be because it is difficult to ethically experimentally induce a sufficiently potent and sustained affective state. An alternative approach is to utilise naturalistic variation in participants’ affective state at the time of the study – influenced by real-life events that have happened to them prior to participating. We will also measure trait affective tendencies to assess the selectivity of state versus trait associations with attentional breadth. We will collect two different types of affective ratings, those that predominately reflect the valence of one’s affective state, and those that predominately affect the motivational intensity of one’s affective state. This will allow us to assess the relative contribution of each to explaining variance in observed attentional breadth.
The paper is a qualitative literary study of how life choices and individualism intertwine in Robert Frost's 1916 poem "The Road Not Taken." The poem is considered one of Frost's most significant works. It uses symbolic division to question the problems of human choice, free will, and self-reflection, thereby dealing with the process of self-formation. Using thematic interpretation and close textual analysis, complemented by a systematic review of academic literature published since 1999, the study investigates how metaphor, imagery, tone, and structural ambiguity communicate the psychological, philosophical, and existential aspects of choice. The results show that Frost views life choices as ambiguous and consequential, with a focus on introspection, anticipation, and retrospective sense-making. The poem's main metaphor conveys the universality of the decision-making process and the individual responsibility taken in personal activity. A theme of individualism also develops, supported by lexical clues such as seldom and difference, which indicate the conflict between social norms and individual freedom. The reflection is inseparable from agency, and the analysis shows that people reconstruct the meaning of their decisions through memory and narrative. The comparative study also shows that the literary elements used by Frost, such as metaphor, ambiguity, and narrative point of view, shed light on both the cognitive and emotional aspects of decision-making, prompting the reader to engage in interpretation. In theory, this study broadens the application of literary, psychological, and philosophical theories to deepen understanding of autonomy, agency, and reflective cognition in poetry. In practice, the findings highlight the usefulness of the poem as an educational, counseling, and personal-development tool that fosters critical thinking about choice and responsibility. The weaknesses of the research are that it is qualitative and focused on textual analysis, and that the study lacks empirical evidence on reader responses, thereby indicating potential areas for future research that can utilize cross-cultural, longitudinal, or experimental research designs. On the whole, the paper has shown that The Road Not Taken has remained relevant in terms of its decision-making, individualism, and self-reflection, even in current human agency discourses.
This dataset contains multimodal neuroimaging and physiological data from a study investigating the effects of Targeted Memory Reactivation (TMR) during REM sleep on emotional reactivity. Participants encoded affective images paired with sounds, received auditory cues during subsequent REM sleep, and were rescanned 48 hours later during arousal rating tasks in an fMRI scanner. The dataset includes structural and functional MRI, polysomnographic recordings with EEG during sleep, heart rate measurements, and behavioral ratings across three sessions spanning two weeks.
English as a Lingua Franca (ELF) plays a vital role in global tourism, facilitating communication between international tourists and service providers but often causing miscommunication due to differing proficiencies, accents, and cultural norms. Such breakdowns can lead to service failures, underscoring the need for effective recovery. This qualitative case study explores ELF interactions in hotels, restaurants, travel agencies and tour operations across three destinations in Lombok, Indonesia. Data from observations, recordings, interviews and complaint documents were analysed using discourse and thematic methods. Four main breakdowns emerged: lexical misunderstandings, pronunciation issues, pragmatic failures and cultural mismatches, where lexical and pragmatic being most frequent. Recovery strategies included clarification, apology, non-verbal support, compensation and proactive adjustment, with cultural expectations mediating success. Relational and proactive approaches yielded the greatest satisfaction. The study integrates ELF service recovery theory, highlighting practical training in cross-cultural pragmatics and flexible recovery to improve multilingual service quality.
This article explores the dynamics of linguistic norms in the Russian literary language within the context of social media communication. The relevance of the study lies in the fact that digital platforms create new communicative conditions that reshape traditional language norms. The paper analyzes variability, simplification, and expressive tendencies in syntax, lexicon, and orthography. Furthermore, it demonstrates that linguistic norms in social media are not static but dynamic and context-dependent. The findings suggest that these processes reflect not the degradation of the literary language, but its adaptation to digital discourse.