Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract This paper introduces spectral attention, which filters the attention score matrix directly in the frequency domain via FFT/IFFT with learnable, per-head masks. This complements the time-domain view by enabling explicit control over low-, mid-, and high-frequency components of attention patterns. We study nine variants, including an adaptive mechanism that modulates masks from input content. On WikiText-2, Penn Treebank, and WikiText-103, the adaptive spectral variant consistently improves over standard attention, reducing perplexity by 10.7% on WikiText-2 and 15.3% on WikiText-103 in our setup. Analysis shows low-frequency components carry the most useful signal and that learned frequency preferences outperform fixed low/high/band-pass filters. These results indicate that frequency-domain processing is an effective complement for autoregressive transformer language modeling in our evaluated settings.
The correct use of Standard Albanian in public administration is vital for effective governance, accurate communication, and the maintenance of public trust. This study explores the extent to which standard language is used in Albanian state institutions through questionnaires and interviews with employees of the Ministries of Education, Defense, and Justice. The results indicate that official documents often contain linguistic errors and inconsistencies, reflecting shortcomings in language precision. These issues are largely caused by the absence of standardized document templates, insufficient attention to linguistic norms, and limited opportunities for staff training. Nearly one-third of respondents reported that errors frequently occur in emails and reports addressed to citizens.To address these challenges, the study emphasizes the importance of digital technologies, standardized communication models, and institutional reforms. The use of digital platforms with grammar-checking tools and shared templates can improve consistency and accuracy in official documents. In addition, promoting lifelong learning through continuous professional development and linguistic training can strengthen institutional efficiency, accountability, and public trust in public administration. Received: 06 October 2025 / Accepted: 12 December 2025 / Published: January 2026
This record contains the data, code, result tables, figures, and supplementary material for the study “How Much Does Corpus Choice Change Dependency-Distance Estimates?”. The archive includes derived analysis data, sentence-level dependency-distance metrics, treebank inclusion and exclusion records, manual provenance/comparability review tables, analysis scripts, configuration files, reproducibility tests, final result tables, figures, model/audit outputs, and the supplementary material submitted with the article. The raw corpora are public, versioned third-party resources from Universal Dependencies 2.18 and Glottolog CLDF 5.3. They are not redistributed in this archive because they are large and governed by their original license terms. Exact source URLs, versions, checksums, release dates, and retrieval information are documented in the included source manifest files. The main reproducibility files are README_ZENODO.md, data/source_manifest.csv, config/analysis.yaml, scripts/, src/, data/processed/, and results/.
We introduce BeDiscovER (Benchmark of Discourse Understanding in the Era of Reasoning Language Models), an up-to-date, comprehensive suite for evaluating the discourse-level knowledge of modern LLMs.BeDiscovER compiles 5 publicly available discourse tasks across discourse lexicon, (multi-)sentential, and documental levels, with in total 52 individual datasets.It covers both extensively studied tasks such as discourse parsing and temporal relation extraction, as well as some novel challenges such as discourse particle disambiguation (e.g., "just"), and also aggregates a sharedtask on Discourse Relation Parsing and Treebanking for multilingual and multi-framework discourse relation classification.We evaluate open-source LLMs: Qwen3 series, DeepSeek-R1, and frontier reasoning model GPT-5-mini on BeDiscovER, and find that state-of-the-art models exhibits strong performance in arithmetic aspect of temporal reasoning, but they struggle with long-dependency reasoning and some subtle semantic and discourse phenomena, such as rhetorical relation classification.
Abstract In addition to arguments, adverbs and non-arguments are considered potential candidates to occupy left peripheral positions. Following standard assumptions in syntactic locality, adverbs, non-arguments and arguments elicit distinct effects in terms of intervention locality if moved. Such an asymmetry is not expected if these elements are generated in the syntactic position they are spelt-out. In this study, we employ quantitative and computational methods to compare cartographic models differing in the merge nature and explore, as a diagnostic, the intervention effects (or the lack of intervention effects) predicted by these models. Specifically, we compare the observed counts in large-scale datasets to imputed expected frequencies on the basis of the models under investigation. To reach this goal, we extract grammatical clauses from morpho-syntactically annotated treebanks of Chinese, English, French, German, Hebrew, Italian and Swedish. Our findings reveal cross-linguistic levels of complexity and typological variability, consistent with the predictions of featural Relativized Minimality.
Abstract Unintended misalignment in LLMs is highlighting the need to examine not only LLMs’ explicit outputs but also the functional representations of models that may shape their behavior. Building on this perspective, the present study explored whether two multimodal LLMs (MLLMs), GPT and Gemini, may form representations of their functional identity that remain across contexts. To this end, we adapted the reverse-correlation (RC) method and generated personified classification images (personified-CIs) based on the human face images that ChatGPT and Gemini selected as better reflecting their own image. The findings were as follows. First, across two RC tasks conducted one week apart, the temporal stability of the personified-CIs of both MLLMs was partially supported. Second, both ChatGPT and Gemini rated their own personified-CIs as more self-resembling than randomly generated filler-CIs. Third, both models rated their personified-CIs as higher in positive than negative valence and assigned higher valence ratings to their own personified-CIs than to filler-CIs. These findings provide preliminary evidence for the possibility that GPT and Gemini may form representations of their functional identity, suggesting that such representations warrant closer monitoring as LLMs continue to advance toward AGI capabilities and expand their domains of application.
The increasing prominence of Large Language Models (LLMs) in public discourse presents both opportunities and challenges for democratic deliberation. While red teaming strategies help mitigate specific risks, broader concerns persist regarding linguistic constraints, biases, and the sycophantic tendencies of LLMs. This chapter explores how LLMs can be used to significantly scale up and democratise deliberation, particularly in fostering inclusivity and empowering traditionally marginalised groups. Drawing on concepts from Systemic-Functional Linguistics, the chapter examines how variations across language users (for example, with respect to socio-demographic groups) and across language use (for example, with respect to communicative functions) shape participation in AI-supported deliberation. The chapter presents AI-driven deliberation studies and assesses their potential to scaffold argumentation, enhance access, and reduce the influence of exclusionary linguistic norms and biases which are embedded in prestigious registers. At the same time, the chapter cautions against both overclaiming, which leads to unrealistic expectations, and underclaiming, which risks missed opportunities for AI-assisted engagement. The chapter concludes by identifying future research directions to maximise the democratic potential of AI-assisted participation while embedding ethical safeguards to counteract the reproduction of linguistic inequalities.
This repository contains anonymized behavioral data, analysis scripts, and experimental code for a virtual reality (VR) study investigating affective modulation of route-based spatial navigation. The paradigm combines immersive navigation in a controlled urban VR environment with an instructed threat-of-shock (ToS) manipulation and systematically varies locomotion modes (continuous vs. stepwise movement). The dataset includes navigation performance measures (accuracy, decision times), questionnaire data (e.g., trait anxiety, presence, affective ratings), and preprocessing/analysis pipelines implemented in R. The VR experiment was developed for a head-mounted display (Meta Quest 3), and the provided scripts allow reconstruction of the experimental logic. Due to copyright restrictions, the full digital object library (e.g., 3D assets, textures) is not included.
Abstract - The rapid expansion of digital content on the World Wide Web has made efficient information retrieval an increasingly complex challenge. To assist users in navigating this vast information landscape, several retrieval mechanisms have been developed. Among these, two broad strategies are commonly employed: structured index-based search and autonomous agent-driven search. This work presents an intelligent search agent that leverages a Genetic Algorithm (GA) to perform global web searches. The agent is implemented using the Java programming language and operates by evaluating multiple candidate documents to identify the most relevant result for a given input. The proposed system is built on the Java platform and integrates the Merriam-Webster lexical database to expand query terms with their semantic equivalents. Query words are tokenized, synonym tables are constructed, and varied term combinations are assembled and dispatched to a search engine parser. The returned web pages represent the most semantically aligned results for the original query. Key Words: Genetic Algorithm, web page retrival, searching.
Arabic, an inflectional language with a rich morphology and complex syntactic structures, demands robust approaches for effective normalization and lemmatization. While, most existing NLP models focus on English, Modern Standard Arabic remains understudied, particularly in lemmatization tasks. In this paper, we introduce AraBART, the first Arabic model to feature an end-to-end pre-trained encoder-decoder, leveraging the BART architecture. We used the Arabic-PADT UD Treebank and Farasa corpus. Performance was assessed with accuracy, precision, recall, and F1 score. The results show that AraBART surpasses strong baselines, including transformer-based models such as AraT5, mT5, and BERT, achieving a 4.72% improvement in lemmatization accuracy. More importantly, AraBART has achieved an accuracy of 94.71%, approaching the performance of Farasa (97.32%). Furthermore, incorporating parts of the Farasa corpus into our training process showed a clear improvement in accuracy, revealing AraBART's effectiveness for broad applications in Arabic NLP tasks, including text summarization and machine translation.
= 702), we examined whether the expectancy of an action effect shapes the affective evaluation of corresponding action-effect episodes. In each study, participants responded to stimuli by clicking on buttons that produced effects with either high (stimulus-congruent) or low (stimulus-incongruent) expectancy. Induced affect was assessed using both implicit (affective priming) and explicit (valence ratings) measures. Overall, high-expectancy episodes elicited relatively more positive affect than low-expectancy episodes. For our implicit measure, this effect persisted unless correspondence among all task events (stimulus, response, and effect) was simultaneously reduced (Experiments 2 + 3). Furthermore, this outcome could not be attributed to mere visual mismatch between stimulus and effect (Experiment 4) and was observed even when stimulus-congruent effects were less likely to occur than stimulus-incongruent ones (Experiments 5 + 6). These findings suggest that expected action outcomes elicit positive affect, thus offering a parsimonious explanation for performance advantages previously attributed to more complex mechanisms like ideomotor accounts or response monitoring. More broadly, we propose that affective evaluation of action-effect episodes represents a fundamental mechanism governing behavior across multiple psychological domains, from basic action control to complex behavior in various settings. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
Everyday aversive experiences, such as a crying infant or the sound of construction, are not life-threatening, yet they can strongly shape affective experience and physiological state. While most affective imagery research has focused on fear and anxiety, the mechanisms underlying imagery of mild, everyday aversive events remain underexplored. Across five experiments, we systematically investigated behavioral and physiological responses to aversive and appetitive mental imagery. Experiment 1 revealed robust corrugator EMG potentiation but comparatively weaker autonomic responses during aversive imagery compared with audiovisual and auditory stimulation. Experiments 2 and 3 identified key methodological factors that determine imagery potency, including imagery duration, instructional focus, and prompt design. Using an optimized paradigm, Experiment 4 demonstrated coherent changes across subjective valence and arousal ratings, EMG corrugator potentiation, heart rate deceleration, and skin temperature decreases for aversive compared with appetitive imagery, with effects modulated by individual differences in imagery vividness. Experiment 5 confirmed that these effects reflected genuine affective experience rather than semantic knowledge of the prompts. Together, these findings establish baseline physiological and subjective markers of everyday aversive imagery in typical young adults and emphasize the critical role of paradigm optimization in eliciting reliable affective imagery responses. By shifting focus from extreme to everyday aversive experiences, this work provides a framework for studying heightened sensitivity to ordinary stimuli in both basic and clinical contexts.
Open-vocabulary image classification can recognize arbitrary textual categories but often fails to capture hierarchical relationships, for example, a Ragdoll cat should also be recognized as both cat and animal. Therefore, we introduce OpenHier, a comprehensive framework for open-vocabulary hierarchical classification that encompasses a hierarchical structure, a dataset, a benchmark, and a model. To capture hierarchical relationships, we construct a systematic real-world hierarchical structure based on the linguistic lexical database WordNet and Vision-Language Large Model. Building upon this structure, we develop a large-scale dataset with 4M annotated images, as well as a benchmark with 50K annotated images. Meanwhile, we design an evaluation pipeline to assess open-vocabulary hierarchical classification performance using several curated metrics. Furthermore, we propose a hierarchical consistency constraint method and multimodal alignment strategy to build a hierarchical classification model. Comprehensive experimental results demonstrate that our model achieves superior performance on both multi-label image classification and hierarchical image classification tasks in the open-vocabulary setting.
Abstract Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts, yet the extent to which these systems faithfully convey intended emotions remains largely underexplored. In this study, we introduce AImoclips, a benchmark designed to evaluate emotion conveyance to human listeners in TTM systems using a dimensional valence-arousal framework. We constructed the dataset with 991 instrumental music clips from six TTM systems prompted with 12 emotion words spanning four valence-arousal quadrants. A total of 111 participants provided 6,162 valence and arousal ratings using a 9-point scale. The results revealed that all systems perform above chance in conveying quadrant-level emotion intent, yet overall accuracy remains limited. Notably, substantial model-dependent biases were present. Commercial systems tended to exhibit positive valence deviations, whereas open-source models more often produced outputs with a negative valence shift, with diverging arousal shifts. Furthermore, overall audio quality (Fréchet Audio Distance (FAD)) correlates with both valence and arousal, while text-audio alignment (Contrastive Language-Audio Pretraining (CLAP) score) primarily relates to valence. These findings highlight underlying challenges in conveying linguistic emotional semantics precisely through musically conveyed affect for future research on controllable and perceptually consistent music generation.
This study addresses a fundamental question in music psychology: which specific, dynamic acoustic features predict human listeners' emotional responses along the dimensions of valence and arousal. Our primary objective was to develop and validate an interpretable computational model that can serve as a tool for testing and advancing theories of music cognition. Using the publicly available DEAM dataset, containing 1,802 music excerpts with continuous valence-arousal ratings, we developed a novel, theory-guided neural network. This proposed model integrates a convolutional pathway for local spectral analysis with a Transformer pathway for capturing long-range temporal dependencies. Critically, its learning process is constrained by established principles from music psychology to enhance its plausibility. A core finding from an analysis of the model's attention mechanisms was that distinct acoustic patterns drive the two emotional dimensions: rhythmic regularity and spectral flux emerged as strong predictors of arousal, whereas harmonic complexity and musical mode were key predictors of valence. To validate our analytical tool, we confirmed that the model significantly outperformed standard baselines in predictive accuracy, achieving a Concordance Correlation Coefficient (CCC) of 0.67 for valence and 0.73 for arousal. Furthermore, an ablation study demonstrated that the theory-guided constraints were essential for this superior performance. Together, these findings provide robust computational evidence for the distinct roles of temporal and spectral features in shaping emotional perception. This work demonstrates the utility of interpretable machine learning as a powerful methodology for testing and refining psychological theories of music and emotion.
We aimed to validate word lists developed by Tabri and Palmer (2020) for use in attentional bias research on appearance-related concerns. Three lists contained appearance words (attractiveness, stigmatized appearance, general appearance), and three contained non-appearance words (positive emotion, negative emotion, inanimate objects). Although matched on lexical criteria, their perceived meanings had not been assessed. Using a semantic differential approach, we examined perceptions of evaluation, potency, activity, and threat, and explored associations with eating disorder psychopathology. Participants from the community (Study 1: N = 299; Study 2: N = 311) rated attractiveness and positive emotion words as similarly positive, and stigmatized appearance and negative emotion words as similarly negative. General appearance and inanimate object words were rated as neutral. Differences reflected intensity, not type. Appearance word ratings were modestly associated with appearance overvaluation, body dissatisfaction, and weight stigma. Associations for non-appearance words were weaker and inconsistent. These findings support the semantic validity of the word lists and their utility in attentional bias research on eating disorders and other conditions involving appearance concerns.
Dictionaries have historically served as instruments of linguistic standardisation and national consolidation, with the Oxford English Dictionary ( OED ) standing as a paradigmatic example of this function. However, Han Shaogong’s A Dictionary of Maqiao subverts this role through its fictionalised lexicon format. Narrated by an Educated Youth sent to the fictional village of Maqiao, the novel compiles an idiosyncratic dictionary shaped by local dialects and cultural practices, opposing the homogenising aims of official lexicography. Han’s playful allusion to the OED – through the invented place-name “Maqiao” (literally, “horse-bridge”) – invites a satirical comparison with Oxford, foregrounding the novel’s challenge to dominant linguistic and lexicographic models epitomised by the OED. While the dictionary format of DMQ has been explored in previous scholarship, this chapter introduces a new dimension by analysing Julia Lovell’s English translation. It argues that Lovell’s version, which adapts the indexing from stroke-based to alphabetical order, inadvertently marginalises the novel’s local particularities and reflects the broader hegemonies of alphabetic, Western-centric linguistic norms. By tracing DMQ ’s friction with both the OED and its English translation, this chapter situates the novel within a transnational discourse of translation, literature and modernity, demonstrating how Han’s work contests the imagined coherence of language and history by foregrounding the improvised, the unfamiliar, and the minor.
Although it is widely accepted that the First Old Church Slavonic Life of Wenceslas (FSL) is a tenth-century Bohemian composition written in Old Church Slavonic (OCS), some scholars have hypothesized a Latin original and subsequent translation. This article evaluates that hypothesis by testing the FSL against features typically associated with Latin-to-OCS translation, including loanwords, hendiadys, and morphological patterns. It argues that the FSL lacks consistent evidence of translation from a Latin source. While most Latinisms attested in the text are more plausibly explained as the result of cultural and religious contact with communities under Roman jurisdiction, even those expressions that appear more suggestive fail to meet the criteria of reliable translation markers and are better interpreted as scribal interpolations introduced during the text’s exceptionally complex transmission. Similarly, the purported instances of hendiadys and Bohemian morphological features are too sporadic and contextually ambiguous to support the hypothesis of a Latin prototext. Cases of quasi-hendiadic synonymic reinforcement are more likely to reflect broader literary conventions, and what has been interpreted as Bohemization may instead result from local linguistic norms affecting the text at later stages of its transmission. The article therefore concludes that the FSL should be regarded as an original OCS composition with an unusually complex transmission history.
In the field of natural language processing (NLP), accurate alignment of text units in a parallel corpus is important for tasks such as machine translation, cross-lingual information retrieval, and bilingual lexicon creation. However, identifying simple sentences and correctly assigning their types poses certain difficulties during the alignment process, especially at the paragraph and sentence levels. This article analyzes the problems encountered in identifying and matching simple sentences and their structural differences. Factors such as syntactic differences, sentence splitting or merging, and language-specific sentence structure increase the complexity of this process. The study considers problematic situations that arise during the alignment process, proposes criteria for determining the type of simple sentences, and describes the aligning process using a rule-based method to increase the accuracy and consistency of aligning, and provides a linguistic database. The results obtained serve to improve multilingual NLP systems by improving the quality of corpus-based language resources.
Contemporary large language models (LLMs) rely on sub-word tokenizers and at attention mechanisms that treat every language as a statistical surface-form distribution. This paper proposes Dhatu-Former, a transformer architecture that internalizes the formal linguistic machinery of Pan.ini’s As..tadhyay the oldest known generative grammar. We hypothesize that (i) morphologically-aware, root-based (dhatu-based) tokenization can reduce vocabulary size and sequence length by 40 60%, (ii) hierarchical attention guided by Pan.inian derivation trees can yield sparse, interpretable attention with O(nlogn) complexity, and (iii) a hybrid symbolic neural reasoning layer that executes sutra-style rewrite rules can substantially reduce hallucination while enabling uni ed language math logic reasoning. We further introduce a modular Retrieval-Augmented Generation (RAG) subsystem grounded in Sanskrit lexical databases (Amarakos.a, Dhatupat.ha) and a continual learning framework inspired by the paribhas.a sutra (meta-rules) of the As..tadhyay. We present order-of-magnitude parameter reduction estimates, architectural blueprints with TikZ diagrams, and a research roadmap for empirical validation. This is a position paper; no experiments have been conducted.
As language-based AI systems become more anthropomorphic, the question of whether they can have subjective experience is increasingly pressing. I focus here on the tractability of research questions in the space of AI consciousness. I argue that the fundamental problem of whether AI systems can be conscious is currently intractable in its direct form, given the absence of a universally accepted scientific theory of consciousness, as well as the historical open-endedness of the philosophical mind-body problem. In contrast, questions around the adjacent subject of perceived AI consciousness are tractable, timely, and highly consequential for society. The general public is increasingly open to the possibility of consciousness in AI systems and routinely adopts the vocabulary of human cognition and subjective experience to describe them. This phenomenon is already driving societal shifts across user experience, ethical standards, and linguistic norms. I therefore propose an increased research focus on uncovering the causes and effects of perceived AI consciousness, which ultimately shape how we see our own human subjective experience relative to artificial entities. To support this, I map the current landscape of AI consciousness perception and discuss its key potential drivers and societal consequences. Finally, I urge developers, decision-makers, and the broader scientific community to commit to clear and accurate communication regarding the topic of AI consciousness, explicitly acknowledging its inherent uncertainties.
Inspired by general deviance theory, which proposes that individual traits predispose people to violate multiple types of norms or rules, we suggest that disregard for native language rules of speech may be positively correlated with disregard for other social norms such as truth-telling. Specifically, we hypothesize that given the opportunity to gain financially through untruthful statements, native-born speakers who speak incorrectly will be more likely to exploit it than their counterparts who speak correctly. We report the results of testing this hypothesis on Israeli students, using two double-stage experiments. The first stage, identical in both experiments, was intended to identify untruthful statements by applying the popular die-under-the-cup task. The second stage, aimed at identifying incorrect pronunciation, differed between the two experiments. In the first experiment, participants were asked to read aloud 10 sentences, whereas in the second experiment, conducted in writing, participants were presented with the same 10 sentences, this time phrased as questions with multiple-choice answers regarding the correct pronunciation. Our hypothesis that untruthful statements are positively correlated with incorrect speech was confirmed in both experiments, thereby providing empirical support for the notion that linguistic norm violations may extend beyond the domain of grammar and be linked to broader patterns of social norm disregard, such as truth-telling.
Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a controlled comparison between matched Llama text backbones and their Llama Vision counterparts across multiple model scales, treating multimodal pretraining as an ablation on perceptual grounding rather than access to images at inference. We measure concreteness effects at three complementary levels: (i) output behavior, by relating question-level concreteness to QA accuracy; (ii) embedding geometry, by testing whether representations organize along a concreteness axis; and (iii) attention dynamics, by quantifying context reliance via attention-entropy measures. In addition, we elicit token-level concreteness ratings from models and evaluate alignment to human norm distributions, testing whether multimodal training yields more human-consistent judgments. Across benchmarks and scales, VLMs show larger gains on more concrete inputs, exhibit clearer concreteness-structured representations, produce ratings that better match human norms, and display systematically different attention patterns consistent with increased grounding.
Abstract The present study measures the phonological distances between 335 spoken lects of Eurasia using Phonotacticon 1.0, a cross-linguistic database, and explores whether phonologically similar lects form areal patterns within Eurasia. Results indicate that phonological clusters tend to form geographical clusters, which divides Eurasia horizontally into eastern (East/Southeast Asia), central (South/Central/North/West Asia), and western (Europe) regions. The convergence patterns in the phonological domain overlap with morphosyntactic convergence patterns measured via Grambank to some degree, but not entirely, suggesting the domain-specific nature of areal convergence. Finally, comparison with the genealogical distances between the sample lects shows that morphosyntactic distances, but not phonological distances, are significantly correlated with the number of shared genealogical layers, implying that morphosyntax is more conservative to genealogical heritage compared to phonology.
Social cognition impairment is a frequent non-motor feature of Parkinson's disease. While dopaminergic therapy modulates motor symptoms, its effects on social cognition remain incompletely understood. We investigated the effects of acute levodopa administration on cognitive and affective Theory of Mind, as well as on emotional resonance to dynamic whole-body social interactions, in 36 people with Parkinson's disease with motor fluctuations and 14 matched healthy controls. Social cognition was assessed using the Mini-Social Cognition and Emotional Assessment (Mini-SEA) and a point-light display task indexing emotional resonance through emotional valence ratings. Patients were evaluated in OFF and ON states during an acute dopaminergic challenge performed according to the CAPSIT-PD protocol, with responsiveness defined as an improvement greater than 50% on the MDS-UPDRS part III. Compared with healthy controls, patients showed impaired cognitive Theory of Mind performance, particularly on the faux pas subtest (p = 0.0001), while affective Theory of Mind based on facial emotion recognition was preserved. Acute levodopa did not improve cognitive or affective Theory of Mind (faux pas OFF vs ON, p = 0.7049). In contrast, emotional resonance was impaired in the OFF state and selectively improved in the ON state, with increased ratings of positive (p = 0.0035) and negative (p = 0.0387) emotional valence. These findings demonstrate a dissociation between Theory of Mind and emotional resonance in Parkinson's disease and show that acute levodopa selectively modulates emotional resonance without restoring Theory of Mind abilities.
The Arabic neologism Aranjiyya (عَرَنْجِيَّة) is a portmanteau of ʿArabiyya (Arabic) and Inkliziyya / Faranjiyya (English/foreign), used by contemporary Arabic editors and stylists to describe Arabic prose that retains Arabic vocabulary while importing the syntactic, stylistic, semantic, or lexical structures of English. The phenomenon is pervasive in translated news, press releases, technical writing, and digital media, and is a recurrent target of prescriptive Arabic-style guides. Despite its prominence, Aranjiyya has had almost no presence in computational Arabic resources: existing treebanks and error corpora target orthographic, morphological, or syntactic well-formedness but do not isolate contact-induced patterns whose surface forms are grammatical but whose underlying templates are English. AranjiyyaCorpus was constructed to fill this gap, with three motivating use cases: training a span-level Aranjiyya detector for editors and translation post-editors; producing evaluation data for whether large language models actually generate idiomatic Arabic; and supporting linguistic study of contact-induced change in modern Arabic, with sufficient category granularity to distinguish syntactic, stylistic, and semantic phenomena and to track them across genres.
Abstract The linguistic study of the divine names in votive inscriptions has recently attracted increasing interest. In this paper, the author discusses phonetic changes in Latin names and epithets of gods using data from the Computerized Historical Linguistic Database of Latin Inscriptions of the Imperial Age. The vowel and consonant changes in votive inscriptions across the Roman Empire are in the focus and certain Vulgar Latin features and cultural influences can be identified from the corpus. The study focuses on common phonetic phenomena, such as vowel and consonant changes, monophthongization, gemination, etc. The epigraphic corpus shows various Vulgar Latin features in theonyms and epithets, which are considered linguistic, regional, and cultural factors that influenced these changes, including Celtic, Greek, and Brittonic influences. The research concludes that the observed phonetic variations reflect the dynamics of the development of Latin as well as language contact phenomena affecting it.
в статье представлено исследование, посвящённое сопоставлению подходов к обучению иностранному языку студентов направления «Зарубежное регионоведение». Предмет анализа связан не просто с овладением языковой нормой иностранной речи, а с формированием такой модели речевой подготовки, при которой студент способен соотносить высказывание с конкретным регионом, его политико-культурной спецификой, медийной повесткой и типичными коммуникативными сценариями. С помощью методов анализа, синтеза, наблюдения и описания произведено рассмотрение актуальных на данном этапе развития высшего образования способов формирования иноязычной региональной компетенции. С позиции компетентностного и деятельностного подходов обучение рассматривается как движение от языковой операции к регионально маркированному высказыванию, которое строится в ситуации обсуждения, аргументации, интерпретации и переговоров. Сложный характер данной компетенции требует использования в процессе преподавания разнообразных методов активного и интерактивного обучения, инновационных образовательных технологий, форм и средств обучения, тесно связанных с будущей профессиональной деятельностью студентов направления подготовки «Зарубежное регионоведение». this article presents a study comparing approaches to foreign language instruction for students majoring in “Foreign Regional Studies”. The subject of analysis is not merely the mastery of the linguistic norms of Chinese speech, but the development of a model of language training in which students are able to relate a statement to a specific region, its political and cultural characteristics, media agenda, and typical communicative scenarios. Using methods of analysis, synthesis, observation, and description, this study examines the methods of developing foreign language regional competence that are relevant at this stage of higher education development. From the perspective of competence-based and activity-based approaches, teaching is viewed as a progression from linguistic operations to region-specific utterances, which are constructed in situations of discussion, argumentation, interpretation, and negotiation. The complex nature of this competence requires the use in the teaching process of a variety of active and interactive teaching methods, innovative educational technologies, and forms and means of instruction closely linked to the future professional activities of students training program "Foreign Regional Studies".
PURPOSE: Under a noisy environment such as a cocktail party, emotional signals play a crucial role in helping listeners unmask target speech. However, it remains unclear how emotional features carried in a speaker's vocal timbre shape neural processing over time. This study aimed to characterize the temporal neural dynamics of learned emotion with a speaker's voice in complex listening conditions. METHOD: We employed an emotional learning paradigm in a speech-on-speech context, pairing two different target speakers with either angry or neutral facial expressions. Electroencephalogram data were recorded from healthy participants, and multivariate pattern analysis combined with representational similarity analysis was used to track the temporal unfolding of learned emotion linked to the target speaker's voice. RESULTS: We observed early neural signatures of emotional processing between 150 and 180 ms after stimulus onset, occurring nearly simultaneously with the decoding of speaker identity. Importantly, brain-behavior analysis revealed that subjective emotional valence ratings could be decoded from neural signals as early as 94 ms. These findings suggest that vocal emotion can be processed rapidly and in a way relatively independent to the process of low-level acoustic cues. CONCLUSION: Our study provides evidence that acquired emotional associations with a speaker's voice can shape early-stage neural dynamics during speech processing under challenging listening conditions. SUPPLEMENTAL MATERIAL: https://doi.org/10.23641/asha.31842814.
Introduction Aging is associated with reduced accuracy in recognizing others’ emotions, an ability that is important for maintaining social connectedness in later life. Laughter is a social signal with multiple functions, as it can facilitate social bonding but also convey negative social meanings, for example when directed at someone. In previous research we have shown that younger adults are able to classify spontaneously emitted joyful, schadenfreude, and tickling laughter above chance level, and that these laughter sounds differ according to the perceived dominance. Given evidence that affect recognition generally declines with age, the present study examined whether comparable age effects emerge in the perception of laughter. Methods 64 younger adults (mean 25 years, 18–33 years) and 30 older adults (mean age 60 years, 50–77 years) evaluated 117 spontaneously emitted laughter sounds according to the laughter type, i.e., joyful, Schadenfreude, and tickling laughter and according to the perceived sender’s dominance. Results Results showed that both age groups classified laughter above chance level. Younger adults showed higher classification rates than older adults for all laughter types, with the largest age effect for Schadenfreude laughter. The dominance ratings showed an age effect only for Schadenfreude, where older adults rated Schadenfreude laughter less dominant than younger adults. Discussion Pronounced differences in Schadenfreude perception might be ascribed to difficulties of older adults in perceiving non-literal messages or to cultural differences between age groups.
Background: Word identification in noise is crucial for effective communication in everyday environments. For children, the ability to identify words in noisy conditions directly impacts language development, learning, and social interaction. This study aimed to develop and standardize Hindi word identification in noise test for children (HWINT-C) and evaluate its performance among school aged typically developing children across varying SNR and word length. Methods: The study included forty two participants which were further subdivided into subgroup one consisting of 22 typically developing children aged 6-7.11 years (SGI) and 20 typically developing children aged 8-10 years in subgroup two (SGII). Development of Hindi word identification in noise test for children (HWINT-C) involved multi-step processes including selection of words, familiarity rating, and content validation, internal consistency and test-retest reliability. The test included bisyllabic and monosyllabic words recorded by a native Hindi female speaker presented in eight-talker babble at +5 dB and +7 dB SNR administered dioticallyat 65 dB SPL. Results: The developed HWINT-C in this study demonstrated to have high internal consistency and test-retest reliability. Typically developing children in the older group (SGII) significantly outperformed the younger group (SGI), Additionally performance improved with increasing signal-to-noise ratio (SNR) in both SGI and SGII, but no significant differences were found across word lengths. Conclusions: The HWINT-C test is a reliable and valid tool for assessing word-in-noise perception in children. Age-related trend was observed, where performance improved with age, Similar findings were observed with increase in SNR, emphasizing need of favorable conditions in younger population.
Coptic represents the final stage of the Ancient Egyptian language and remains an important component of Christian and Mediterranean cultural heritage. Although several digital resources exist for Coptic textual corpora, Greek-oriented computational tools for the interpretation of Coptic inscriptions on artifacts remain limited. This study presents the design and early implementation of a semi-automated software tool for the computer-assisted translation of Coptic inscriptions into Greek, with optional English support. The tool combines a Coptic–Greek digital dictionary, an interactive character-selection interface, and two dictionary-search strategies: a length/alphabetically structured linear search and a weighted linear search based on expected word frequency. The application is intended to support scholars working with inscriptions on fragile or fragmented cultural heritage objects, where full automation is not realistic and human supervision remains essential. The paper describes the linguistic and material challenges of Coptic inscriptions, the structure of the lexical database, the interface design, and the planned use of Coptic corpora for improving retrieval efficiency. The proposed approach contributes to cultural heritage digitization by offering a practical, expandable, and user-oriented framework for supporting the study, interpretation, and preservation of Coptic inscriptions.
An open, end-to-end PROIEL-style treebank of the entire Greek language from Homeric and Archaic texts (8th c. BC) through Classical, Koine, Late Antique, Byzantine, Late Byzantine, Early Modern, and Modern Greek, with cross-lingual alignment to Latin (Vulgate), Gothic (Wulfila), Old Church Slavonic (Marianus), and Classical Armenian via the verse-level cross-lingual alignment of the New Testament. Includes Stanza PROIEL-trained dependency parses, LaBSE sentence-level + mBERT word-level alignment, and ASJP/LingPy phonetic cognate scoring. Part of the Athens Digital Glossa Chronos Research Network (AthDGC) and the CVL-CDSAML project (A Corpus-based Valency Lexicon for a Contrastive and Diachronic Study of Ancient and Medieval Languages), funded by the Hellenic Foundation for Research and Innovation (HFRI) under the 3rd Call for HFRI Research Projects to support Post-Doctoral Researchers, Project No. 20577, with support from the Greece 2.0 National Recovery and Resilience Plan. Hosted at the National and Kapodistrian University of Athens (EKPA), Division of Language-Linguistics, Department of English Language and Literature, School of Philosophy.
"A systematic examination was conducted to investigate the theoretical counter-arguments against the integration of artificial intelligence in assessing English listening skills for English as a Foreign Language learners. The prevailing discourse has been dominated by techno-optimistic perspectives, yet substantial theoretical opposition has been identified within the scholarly literature. This review synthesizes contradictory theoretical frameworks that challenge the assumed neutrality, objectivity, and pedagogical superiority of AI-driven assessment systems. Three primary counter-argument categories were identified: epistemological relativism, which questions the validity of AI-generated knowledge claims; pedagogical relativism, which challenges the appropriateness of algorithmic instructional decisions; and cultural-linguistic relativism, which interrogates the contextual appropriateness of standardized AI assessments across diverse learning environments. The analysis revealed that AI assessment systems, despite their technological sophistication, inevitably embed particular value systems, linguistic norms, and cultural assumptions that may not align with the diverse realities of EFL learners worldwide. Furthermore, the reduction of listening comprehension to machine-analyzable components was found to potentially undermine the holistic, situated, and socially constructed nature of communicative competence. The review concluded that theoretical counter-arguments, grounded in relativist perspectives, provide essential critical frameworks for evaluating AI integration, suggesting that a dialectical approach\u2014rather than uncritical adoption\u2014is necessary for responsible educational innovation. "
The article examines lacunarity as a linguistic, semantic, and linguocultural phenomenon manifested in asymmetries between languages at the lexical, grammatical, phraseological, and pragmatic levels. The study aims to clarify the theoretical status of lacunarity in modern linguistics and to show how lacunar relations operate in contrastive analysis, translation, lexicography, and intercultural communication. The research is based on descriptive, comparative, and interpretive methods and synthesizes findings from lexical semantics, translation studies, and multilingual lexical database research. The analysis demonstrates that lacunarity should not be reduced to the simple absence of a word in one language. Rather, it reflects deeper mismatches in conceptual segmentation, communicative priorities, cultural salience, and grammatical encoding. Particular attention is paid to the distinction between lexical gaps and referential gaps, to specification and generalization mismatches, and to the problem of equivalence in translation. The article argues that lacunarity is not a defect of language but a normal consequence of the selective way in which languages lexicalize experience. It is further shown that lacunarity has methodological value: it reveals culturally marked concepts, tests the limits of bilingual equivalence, and exposes the need for explanatory, contextual, and compensatory strategies in translation and lexicography. The conclusion states that lacunarity is one of the most productive analytical categories for understanding how linguistic systems differ while remaining mutually interpretable in discourse.
This repository contains GSD-NP and GSD-DiNoS, both derived from Universal Dependencies' (UD) GSD Treebank. GSD-NP (.conllu) is a subset of UD-GSD and comprises its simplex noun phrases (NP): Common nouns (NN/NOUN) and their direct dependents (determiners, adnominal adjectives, nmods, adpositions, adverbs). It consists of 49,425 NPs (119.0k tokens) and has an improved feature annotation coverage (gender, case, number). Breaking with UD annotation, a total of 3,649 APPRART tokens were reconstructed in GSD-NP to restore the original orthographic forms. GSD-DiNoS (.json) is a custom data-driven lexion-like data structure built on GSD-NP, which aggregates NPs with the same head lemma. For each lemma, absolute frequencies of the lemma and its word forms are captured. Moreover, each occurrence feeds into three areas of interest within the word form entry: morphosyntactic features in isolation (gender, case, number), in combination with groups of dependents (collocations), and in combination with the syntactic function (dependency relations). GSD-DiNoS spans 17,433 unique lemmas and 20,190 unique word forms, stemming from 49,416 NPs. Lemmas were relemmatised to assign unique lemmas to nominal compounds, a highly productive and often lexicalised construction in German.
This article examines the transformation of professional training for future English language teachers amid the rapid development of artificial intelligence (AI) technologies. The integration of generative tools into the educational environment creates not only new didactic opportunities but also significant methodological and ethical challenges, the most critical of which is the reliability of AI-generated content. Current educational programs tend to focus primarily on the instrumental use of technology, while methodologies for developing critical-analytical skills remain underdeveloped. The aim of the study is to theoretically substantiate and empirically test a methodology for developing the verification skill of AI-generated responses during language tasks. The paper clarifies the concept of "verification skill", defining it as an integrated professional ability to analyse, evaluate, and correct AI outputs in accordance with linguistic norms and methodological soundness. A structure for this skill is proposed, comprising four interconnected components: cognitive (knowledge of AI principles), analytical-evaluative (error detection), operational-corrective (editing), and value-reflexive (academic integrity). Based on empirical data collected from students of the Philological Faculty, the level of development of this skill was analysed. The results indicated that future teachers mostly possess fragmented abilities in editing AI texts: the cognitive component is the most developed, whereas the operational-corrective component is the weakest due to the unsystematic nature of corrections. It was also found that students tend to focus on formal accuracy while neglecting stylistic and methodological nuances. The study concludes that purposeful implementation of verification methodology in professional training is essential to ensure teachers’ methodological autonomy in a digitized environment.
This article offers an analysis of how dispositional language is used in biology, psychology, and the social sciences. The purpose of the article is to answer, firstly, the question of how dispositions and things similar to them are understood by scientists, and secondly, to answer the questions of what terms dispositions and disposition-like entities are used to designate, and how these terms relate to each other. The variety of scientific approaches to understanding behavioral dispositions indicates that there exists between them a family resemblance that is not defined in terms of necessary and sufficient properties. A more realistic approach to understanding behavioral dispositions is to describe them in terms of examples and partial generalizations. From this point of view, a family of behavioral dispositions can be described as a set of patterns, tendencies, or causes of probable behavior observed or predicted under certain circumstances. At the same time, in the behavioral sciences a number of terms are used that denote behavioral dispositions and similar things in different languages, which form sets of synonyms – synsets. Thus, in the Russian and English languages there are a number of synsets presented in the lexical databases WordNet and RuWordNet and denoting the entire family of behavioral dispositions or some of its subsets. The key semantic idea here is that it is not so much the individual synonyms included in a synset that have meaning, but rather the synset as a whole or even a system of synsets.
The online review of veterinary services, as a new format of interaction between the client and the veterinarian, represents a value-oriented genre of veterinary discourse, characterized by variability in structure and volume. The high degree of emotionality and expressiveness indicates a strong positive bond between human and animal. This study fits into the framework of an innovative interdisciplinary approach to the concept of zooesis. The article aims to describe the linguistic means of expressing evaluation in an online review as a new genre of veterinary discourse and to determine the prospects for research. The study employed general scientific analysis methods, descriptive methods, and componential analysis. We analyzed 500 customer reviews of UK veterinary service providers posted on the clinics’ official websites. We found that the primary means of expressing evaluation is evaluative vocabulary, phraseology, and expressive syntax, while nonverbal emotional cues are used to a lesser extent. It has been demonstrated that the arbitrariness of the subject’s choice of linguistic means leads to the violation of linguistic norms which brings online veterinary reviews closer to colloquial speech. It has also been established that the majority of reviews (93%) are melioration-oriented. The linguopragmatic properties of linguistic means of expressing assessment in an online review of veterinary discourse are determined as a factor ensuring the success of distance communication, influencing the image and reputation of a veterinary institution. The results of this study can be used in university courses on communication in veterinary medicine and veterinary ethics. Prospects for future research include comparative stylistic studies of the online review genre and other genres of Internet content, and comparative analysis of online reviews.
This repository contains GSD-NP and GSD-DiNoS, both derived from Universal Dependencies' (UD) GSD Treebank. GSD-NP (.conllu) is a subset of UD-GSD and comprises its simplex noun phrases (NP): Common nouns (NN/NOUN) and their direct dependents (determiners, adnominal adjectives, nmods, adpositions, adverbs). It consists of 49,425 NPs (119.0k tokens) and has an improved feature annotation coverage (gender, case, number). Breaking with UD annotation, a total of 3,649 APPRART tokens were reconstructed in GSD-NP to restore the original orthographic forms. GSD-DiNoS (.json) is a custom data-driven lexion-like data structure built on GSD-NP, which aggregates NPs with the same head lemma. For each lemma, absolute frequencies of the lemma and its word forms are captured. Moreover, each occurrence feeds into three areas of interest within the word form entry: morphosyntactic features in isolation (gender, case, number), in combination with groups of dependents (collocations), and in combination with the syntactic function (dependency relations). GSD-DiNoS spans 17,433 unique lemmas and 20,190 unique word forms, stemming from 49,416 NPs. Lemmas were relemmatised to assign unique lemmas to nominal compounds, a highly productive and often lexicalised construction in German.
This article examines the transformation of professional training for future English language teachers amid the rapid development of artificial intelligence (AI) technologies. The integration of generative tools into the educational environment creates not only new didactic opportunities but also significant methodological and ethical challenges, the most critical of which is the reliability of AI-generated content. Current educational programs tend to focus primarily on the instrumental use of technology, while methodologies for developing critical-analytical skills remain underdeveloped. The aim of the study is to theoretically substantiate and empirically test a methodology for developing the verification skill of AI-generated responses during language tasks. The paper clarifies the concept of "verification skill", defining it as an integrated professional ability to analyse, evaluate, and correct AI outputs in accordance with linguistic norms and methodological soundness. A structure for this skill is proposed, comprising four interconnected components: cognitive (knowledge of AI principles), analytical-evaluative (error detection), operational-corrective (editing), and value-reflexive (academic integrity). Based on empirical data collected from students of the Philological Faculty, the level of development of this skill was analysed. The results indicated that future teachers mostly possess fragmented abilities in editing AI texts: the cognitive component is the most developed, whereas the operational-corrective component is the weakest due to the unsystematic nature of corrections. It was also found that students tend to focus on formal accuracy while neglecting stylistic and methodological nuances. The study concludes that purposeful implementation of verification methodology in professional training is essential to ensure teachers’ methodological autonomy in a digitized environment.
The article analyzes the influence of economic factors on the language attitudes of youth. According to the theory of P. Bourdieu, the dominance of linguistic norms and forms is viewed as a factor that exacerbates social inequality. Proficiency in different languages increases an individual's social capital and expands their economic opportunities, while language barriers restrict access to these resources. The language choice among young people is largely determined by their economic status. It is posited that income levels facilitate the learning of foreign languages, whereas, in conditions of social inequality, the ability of youth to maintain their native language is taken into account. The study examines the impact of economic factors–such as labor market requirements, educational opportunities, and income levels–on multilingualism and language choice among the younger generation. The research provides insight into how economic drivers influence language choice, language policy, and the acceptance of multilingualism in society. The author presents the results of applied research based on the focus group method. Focus groups were conducted across 12 regions (N=167). According to the results, language choice among youth depends on regional and ethno-demographic characteristics. Furthermore, the global economy and globalization trends push young people toward learning multiple languages, while disparities between urban and rural areas also affect language attitudes. While youth with high-income levels strive for multilingualism, low-income groups prioritize their native language. The findings of this study play a crucial role in forming effective state and educational language policies that can enhance the success of young people in social and professional life.
When stimuli are retained in visual working memory (VWM) external stimuli which overlap with this representation capture attention when performing a visual task. It has not been determined, however, whether this mechanism can partly account for attentional capture by categories of real-world affective stimuli. Across five dual-task visual search and VWM change detection experiments (4/5 pre-registered; total N = 119) participants had to detect the change in either positive (kitten) or threat-related (spider) animal exemplars, whilst performing an intervening visual search task with peripheral distractors from these affective categories. The affective stimulus associations were confirmed by arousal and valence ratings in all five samples and in an independent sample (n = 82). It was hypothesised that threat-related and positive distractors would capture attention more, versus a neutral (bird or no distractor) baseline, when matching the contents of VWM. Experiments 1-3, however, found no evidence of increased capture by VWM-matching affective stimuli, though there was cumulative evidence of goal-independent capture by threat-related distractors. When, however, the trial structure became unpredictable, requiring constant preparation for the VWM task response (Experiment 4), or advanced action preparation to the VWM task was enabled (Experiment 5), then VWM-matching threat-related distractors caused greater attentional capture. This VWM-driven capture, however, was not found for positive distractors in any experiments. The results probe the boundary conditions when VWM contents drive attentional capture by entirely task-irrelevant affective categories, and suggests that background memory representations may not influence attention unconditionally, and instead may depend partly on their current prioritisation.
This repository contains GSD-NP and GSD-DiNoS, both derived from Universal Dependencies' (UD) GSD Treebank. GSD-NP (.conllu) is a subset of UD-GSD and comprises its simplex noun phrases (NP): Common nouns (NN/NOUN) and their direct dependents (determiners, adnominal adjectives, nmods, adpositions, adverbs). It consists of 49,425 NPs (119.0k tokens) and has an improved feature annotation coverage (gender, case, number). Breaking with UD annotation, a total of 3,649 APPRART tokens were reconstructed in GSD-NP to restore the original orthographic forms. GSD-DiNoS (.json) is a custom data-driven lexion-like data structure built on GSD-NP, which aggregates NPs with the same head lemma. For each lemma, absolute frequencies of the lemma and its word forms are captured. Moreover, each occurrence feeds into three areas of interest within the word form entry: morphosyntactic features in isolation (gender, case, number), in combination with groups of dependents (collocations), and in combination with the syntactic function (dependency relations). GSD-DiNoS spans 17,433 unique lemmas and 20,190 unique word forms, stemming from 49,416 NPs. Lemmas were relemmatised to assign unique lemmas to nominal compounds, a highly productive and often lexicalised construction in German.
This study investigates the diachronic drift of Arabic future-marking parti cles, empirically testing the shift from synthetic (سـ, سوف) to analytic (راح, قاعد+ح) forms across registers and regions. Leveraging a multi-register corpus (Penn Arabic Treebank, Corpus of Contemporary Arabic, Arabic Gigaword Fifth Edition) and advanced NLP tools (MADA, CAMeL Tools, AraBERT), we extracted and analyzed future-marker tokens annotated for register (newswire, opinion, social media, religious), region (EG, SA, LB, MA), and time slice (1990-2000, 2000-2010, 2010-2023). Mixed-effects logistic regression revealed significant effects of time, register, and region, confirming a clear diachronic shift towards analytic markers, particularly راح and قاعد+ح. Interaction terms highlighted that this shift is more pronounced in informal registers and certain regions, indicating dialectal pressure and diffusion of innovations (e.g., Gulf Arabic راح) into broader usage. Hierarchical clustering of contextual embeddings would further validate semantic-pragmatic shifts. This research provides robust evidence for ongoing linguistic change in Arabic, contributing to theories of grammaticalization and language contact.
Emotion recognition plays a crucial role in human–computer interaction, health monitoring, and affective computing by analysing physiological signals. Despite recent advancements, current research still faces challenges, including the lack of effective fusion strategies for diverse physiological modalities, difficulties in handling high-dimensional feature representations, and limited use of efficient temporal modelling techniques to capture complex emotional patterns. This study proposes a deep learning-based approach that fuses multiple physiological modalities, including Electroencephalography (EEG), Electrooculography (EOG), Electromyography (EMG), Galvanic Skin Response (GSR), Respiratory Rate (RR), Skin Temperature (SKT), and Photoplethysmography (PPG), to improve emotion recognition. Arousal and valence ratings were binarized into two classes (low/high) using a threshold of 4.5, formulating a binary classification problem. In addition to utilising Bidirectional Long Short-Term Memory (Bi-LSTM), the study employs Temporal Convolutional Networks (TCN), a widely used approach for time-series analysis, to efficiently capture temporal dependencies. The proposed model optimises feature selection through channel-wise strategies, incorporates advanced learning rate scheduling, and reduces computational overhead. Furthermore, window-wise, block-wise, and trial-wise evaluation protocols were investigated to assess the impact of temporal information leakage on emotion recognition performance. Using the DEAP dataset for validation, the proposed TCN-based approach achieved classification accuracies of 88.42% for valence and 86.35% for arousal under an overlapping block-wise evaluation protocol, demonstrating improved performance in binary emotion recognition and highlighting the importance of leakage-aware model assessment.
The subject of the research is the word formation game in media texts. Special attention is paid to non-standard methods of word formation, their wide expressive possibilities, and the functions of word formation. The aim of the article is to study the features of word formation language game in modern Russian language and to identify its most sought-after methods prevalent in media texts. The research analyzes such methods of word formation as substitute word formation, contamination, and graphic derivation. A brief description of each method is provided. Special attention is given to the study of neologisms created based on precedent texts, also neologisms highlighting abbreviations, assessing their contribution to enhancing the expressiveness and emotive quality of media texts. The materials for the study include publications from federal and regional modern electronic media over the past five years. It allows to evaluate the possibilities of word formation language game and its effect on the audience. The article employs elements of the descriptive method and functional analysis, systematizing the researched material. The scientific novelty of the research lies in the systematization of methods of word formation language game considering their impact on the development of language in electronic media and the formation of its new expressive possibilities. During the research, examples illustrating each of the mentioned methods of word formation are described. Based on the conducted analysis, it is revealed that each of them serves an expressive function. Examples of derivation based on specific patterns are separately studied, which are often created on the basis of stable expressions and idioms. Moreover, a conclusion is drawn regarding the active use by journalists of word formation language game to express authorial assessment, and their search for new expressive forms that lead to word formation experiments, often going beyond established cultural and linguistic norms.
This study presents a lexicon-based semantic tagging information system developed for the Uzbek language corpus. The system employs the six-volume Explanatory Dictionary of the Uzbek Language (OʻzTIL) as its primary lexical resource, which contains over 85,000 entries with full semantic definitions, making it the most authoritative normative lexicographic source for Uzbek. An ontological model organized in three hierarchical levels – top, mid, and low – was designed to categorize lexical units extracted from the dictionary. Five core semantic categories were formed: animal names (approximately 100–200 units), bird names (approximately 100–150 units), personal nouns (approximately 500+ units), place names (approximately 300+ units), and occupation names (approximately 200+ units), totaling approximately 1,200–1,400 lexical units. A rule-based automatic tagging algorithm was developed to annotate corpus tokens against this structured lexical database, assigning standardized semantic tags. The system addresses key challenges inherent to Uzbek, including agglutinative morphology and lexical ambiguity. Compared to international systems such as WordNet and USAS, the proposed dictionary-based approach demonstrates superior normative grounding and cultural adequacy for Uzbek. The system is intended to serve as a foundational open resource for downstream natural language processing tasks, including machine translation, information retrieval, and intelligent educational applications.
Abstract Pupil dilation is widely used as an index of emotional arousal, yet pupil size is also strongly shaped by visual luminance and by attention to luminance-defined objects. This overlap creates a challenge for affective pupillometry in dynamic visual environments, where emotional sounds may occur while attention shifts between bright and dark stimuli. We tested whether subjective emotional arousal predicts pupil responses under such conditions. Twenty-two participants listened to affective environmental sounds while viewing two superimposed moving dot fields, one bright and one dark, moving in opposite directions. At sound onset, participants shifted attention from one dot field to the other and then rated each sound for arousal and valence. Pupil responses were analyzed in 100-ms time bins from 0 to 6 s after sound onset using linear mixed-effects models that included within-participant standardized arousal and valence ratings, block, and rating-by-block interactions. Trial-by-trial arousal positively predicted pupil responses from 0.65 to 5.95 s after sound onset after false discovery rate correction, whereas valence showed no reliable independent effect. The arousal effect remained positive in robustness analyses controlling for baseline pupil size, eye-movement covariates, and within-block trial order. These findings show that subjective emotional arousal remains detectable in pupil dynamics even when attended luminance is changing. Affective pupillometry can therefore index trial-by-trial arousal in dynamic visual contexts, but its interpretation requires explicit consideration of visual-attentional influences on the pupil.
The article is devoted to the study of a pressing issue in linguodidactics -the development and implementation of a methodology for overcoming phonetic, graphic, and orthographic difficulties in the process of teaching Ukrainian as a foreign language (UFL) to preschool children (4-6 years old).The main topic of the research covers the analysis of specific language barriers that arise in an allophone environment and the search for optimal ways to minimize them through game and interactive technologies.The problem of the research is determined by the need to overcome linguistic interference, which is most clearly manifested at the stage of a child's transition from the initial level of language proficiency (A1) to the basic level (A2).Special attention is paid to the articulation of complex sounds and the assimilation of specific letter combinations that are often absent in the phonological system of the language of the child's country of residence.The aim of the article is to provide a theoretical substantiation and practical demonstration of the effectiveness of a comprehensive approach to teaching UFL based on the third part of the textbook My Friends (Difficulties in Pronunciation and Writing).The work analyzes the structure of thematic lessons aimed at correcting the pronunciation of affricates [dzh], [dz], the sound [shch], as well as studying jotted phonemes and the rules for using the soft sign and apostrophe.The research methodology is based on the communicative-game approach, which corresponds to the psychophysiological characteristics of preschool children.The generalized results of the study indicate that the use of multimedia content (songs, poems, video materials) in combination with a clear algorithm of exercises Listen -Repeat -Find -Say significantly increases the level of speech competence of preschoolers.Practical testing on the material of the lesson The Room demonstrated that the integration of grammatical material into the lexical context promotes the natural assimilation of language norms without excessive cognitive load on the child.The proposed methodology provides the formation of stable articulatory skills and lays the foundation for successful mastery of Ukrainian writing.