Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
<h3>Introduction</h3> Ancient Chinese WordNet <a href="../../../LDC2026L03">(LDC2026L03)</a> was developed by <a href="https://www.njnu.edu.cn/">Nanjing Normal University</a> and contains lexical and semantic information for Ancient Chinese vocabulary dating back to the Pre-Qin period (before 221 BCE). The WordNet comprises 38,781 word forms and 55,100 senses, each manually linked to a corresponding synset in <a href="https://wordnet.princeton.edu/">Princeton WordNet 1.6</a>. The Ancient Chinese WordNet (ACWN) project began in 2012 with the goal of creating a structured lexical database to support linguistic research and natural language processing applications involving historical Chinese language materials. ACWN organizes vocabulary using WordNet's noun, verb, adjective, and adverb hierarchies and provides WordNet definitions, semantic relations, and categorization for each sense. <h3>Data</h3> Ancient Chinese WordNet contains 55,100 records, where each record represents a single Ancient Chinese lexical item mapped to one WordNet synset. It follows WordNet 1.6 organizational structure, including 22 noun categories, 15 verb categories, and additional adjective and adverb categories. Each entry includes the following fields: <ul> <li>ID - The serial number of the ACWN entry</li> <li>Word - Ancient Chinese word form</li> <li>wn_offset - 8-digit WordNet 1.6 synset offset with trailing POS (n/v/a/s/r)</li> <li>senseid - Sense number for this word form (ordinal among that word's senses)</li> <li>pos - Part of speech (noun (n), verb (v), adj (a/s), adv (r))</li> <li>wn_category - Numeric code for the WordNet 1.6 lexicographer file (category)</li> <li>wn_synset - Synset headword(s) in WordNet 1.6</li> <li>wn_definition - WordNet gloss for the synset</li> <li>wn_similar to - Synset with similar meaning</li> <li>wn_pertainym - Pertainym synset offset(s)</li> <li>wn_attribute - Attribute synset offset(s)</li> <li>wn_hypernym - Hypernym synset offset(s)</li> <li>wn_hyponym - Hyponym synset offset(s)</li> </ul> The data is presented in UTF-8 encoded CSV and XLSX formats. <h3>Updates</h3> No updates at this time.
In the field of natural language processing (NLP), accurate alignment of text units in a parallel corpus is important for tasks such as machine translation, cross-lingual information retrieval, and bilingual lexicon creation. However, identifying simple sentences and correctly assigning their types poses certain difficulties during the alignment process, especially at the paragraph and sentence levels. This article analyzes the problems encountered in identifying and matching simple sentences and their structural differences. Factors such as syntactic differences, sentence splitting or merging, and language-specific sentence structure increase the complexity of this process. The study considers problematic situations that arise during the alignment process, proposes criteria for determining the type of simple sentences, and describes the aligning process using a rule-based method to increase the accuracy and consistency of aligning, and provides a linguistic database. The results obtained serve to improve multilingual NLP systems by improving the quality of corpus-based language resources.
Contemporary large language models (LLMs) rely on sub-word tokenizers and at attention mechanisms that treat every language as a statistical surface-form distribution. This paper proposes Dhatu-Former, a transformer architecture that internalizes the formal linguistic machinery of Pan.ini’s As..tadhyay the oldest known generative grammar. We hypothesize that (i) morphologically-aware, root-based (dhatu-based) tokenization can reduce vocabulary size and sequence length by 40 60%, (ii) hierarchical attention guided by Pan.inian derivation trees can yield sparse, interpretable attention with O(nlogn) complexity, and (iii) a hybrid symbolic neural reasoning layer that executes sutra-style rewrite rules can substantially reduce hallucination while enabling uni ed language math logic reasoning. We further introduce a modular Retrieval-Augmented Generation (RAG) subsystem grounded in Sanskrit lexical databases (Amarakos.a, Dhatupat.ha) and a continual learning framework inspired by the paribhas.a sutra (meta-rules) of the As..tadhyay. We present order-of-magnitude parameter reduction estimates, architectural blueprints with TikZ diagrams, and a research roadmap for empirical validation. This is a position paper; no experiments have been conducted.
Do vision--language models (VLMs) develop more human-like sensitivity to linguistic concreteness than text-only large language models (LLMs) when both are evaluated with text-only prompts? We study this question with a controlled comparison between matched Llama text backbones and their Llama Vision counterparts across multiple model scales, treating multimodal pretraining as an ablation on perceptual grounding rather than access to images at inference. We measure concreteness effects at three complementary levels: (i) output behavior, by relating question-level concreteness to QA accuracy; (ii) embedding geometry, by testing whether representations organize along a concreteness axis; and (iii) attention dynamics, by quantifying context reliance via attention-entropy measures. In addition, we elicit token-level concreteness ratings from models and evaluate alignment to human norm distributions, testing whether multimodal training yields more human-consistent judgments. Across benchmarks and scales, VLMs show larger gains on more concrete inputs, exhibit clearer concreteness-structured representations, produce ratings that better match human norms, and display systematically different attention patterns consistent with increased grounding.
Large Language Models (LLMs) are increasingly used as research tools to facilitate the fast and automated extraction of text features. In psychological studies, they have been used to quantify the degree to which verbal stimulus materials reflect certain psychological constructs. However, the application of LLMs entails a high degree of flexibility regarding prompt design (e.g., instruction details and examples) and model specification (e.g., model family, size, and configuration), which can produce divergent results and threaten the robustness and generalizability of conclusions. To navigate the multiverse of possible choices, we develop a structured workflow for evaluating the quality of LLM feature extraction across diverse model and prompt specifications. Motivated by generalizability theory, the workflow distinguishes between construct variance across stimulus items, method variance due to model and prompt choices, and error variance across repeated iterations. To guide researchers through the planning, execution, and reporting of LLM simulation studies, we introduce an adapted version of the ADEMP template (Aims, Data-generating mechanism, Estimands and targets, Methods, Performance measures), originally developed for methodological simulation research. The template supports two complementary validation strategies: variance decomposition for studying consistency across LLM specifications and external validation against human gold-standard ratings. In a pre-registered case study using locally runnable, open-weight LLMs, we illustrate the workflow by examining the influence of model choice, response format, and prompt examples on the quality of valence and arousal ratings for multi-word expressions. We additionally assess the efficacy of aggregating repeated, stochastic LLM ratings to improve feature extraction quality.
Social cognition impairment is a frequent non-motor feature of Parkinson's disease. While dopaminergic therapy modulates motor symptoms, its effects on social cognition remain incompletely understood. We investigated the effects of acute levodopa administration on cognitive and affective Theory of Mind, as well as on emotional resonance to dynamic whole-body social interactions, in 36 people with Parkinson's disease with motor fluctuations and 14 matched healthy controls. Social cognition was assessed using the Mini-Social Cognition and Emotional Assessment (Mini-SEA) and a point-light display task indexing emotional resonance through emotional valence ratings. Patients were evaluated in OFF and ON states during an acute dopaminergic challenge performed according to the CAPSIT-PD protocol, with responsiveness defined as an improvement greater than 50% on the MDS-UPDRS part III. Compared with healthy controls, patients showed impaired cognitive Theory of Mind performance, particularly on the faux pas subtest (p = 0.0001), while affective Theory of Mind based on facial emotion recognition was preserved. Acute levodopa did not improve cognitive or affective Theory of Mind (faux pas OFF vs ON, p = 0.7049). In contrast, emotional resonance was impaired in the OFF state and selectively improved in the ON state, with increased ratings of positive (p = 0.0035) and negative (p = 0.0387) emotional valence. These findings demonstrate a dissociation between Theory of Mind and emotional resonance in Parkinson's disease and show that acute levodopa selectively modulates emotional resonance without restoring Theory of Mind abilities.
This research explores the historical emergence of linguistic terminology in three languages—English, Uzbek, and Karakalpak—with special attention to the role of Latin, Greek, and Arabic heritage. It traces how borrowed concepts were nativized and localized in each linguistic setting. By juxtaposing five evolutionary stages in English with analogous processes in Uzbek and Karakalpak, the paper illustrates the interplay between international scholarly traditions and indigenous linguistic norms. The conclusions highlight both universal tendencies and language-specific particularities in the growth of terminological systems.
Introduction Aging is associated with reduced accuracy in recognizing others’ emotions, an ability that is important for maintaining social connectedness in later life. Laughter is a social signal with multiple functions, as it can facilitate social bonding but also convey negative social meanings, for example when directed at someone. In previous research we have shown that younger adults are able to classify spontaneously emitted joyful, schadenfreude, and tickling laughter above chance level, and that these laughter sounds differ according to the perceived dominance. Given evidence that affect recognition generally declines with age, the present study examined whether comparable age effects emerge in the perception of laughter. Methods 64 younger adults (mean 25 years, 18–33 years) and 30 older adults (mean age 60 years, 50–77 years) evaluated 117 spontaneously emitted laughter sounds according to the laughter type, i.e., joyful, Schadenfreude, and tickling laughter and according to the perceived sender’s dominance. Results Results showed that both age groups classified laughter above chance level. Younger adults showed higher classification rates than older adults for all laughter types, with the largest age effect for Schadenfreude laughter. The dominance ratings showed an age effect only for Schadenfreude, where older adults rated Schadenfreude laughter less dominant than younger adults. Discussion Pronounced differences in Schadenfreude perception might be ascribed to difficulties of older adults in perceiving non-literal messages or to cultural differences between age groups.
The intricate relationship between truth and language has long fascinated philosophers, linguists, and scholars across disciplines. In this work, Prof. Dr. Yoesoep Edhie Rachmad, Ph.D., DBA., embarks on a profound exploration of how language shapes, conveys, and sometimes distorts truth. Published in 2000 under The United Nations and The Education Training Centre, this book critically examines the philosophical foundations of linguistic representation and its implications for human understanding. Through an analysis of classical and contemporary theories, the work delves into the roles of meaning, interpretation, and the social constructs that govern communication. By questioning whether absolute truth can ever be expressed without ambiguity, this book invites readers to reconsider their assumptions about knowledge, semantics, and reality itself. Language is the primary medium through which humans express ideas, communicate beliefs, and establish shared realities. However, can language truly capture the essence of truth? This book is born from a deep intellectual curiosity regarding the limitations and potentials of linguistic structures in truth-seeking endeavors. With rapid advancements in technology, media, and cross-cultural exchanges, the question of how truth is conveyed in different languages and frameworks becomes increasingly urgent. Addressing these concerns, this book seeks to bridge the philosophical and practical aspects of truth in language. Understanding truth within the linguistic paradigm requires an examination of key philosophical theories, such as the correspondence, coherence, and pragmatic theories of truth. This book dissects the works of foundational thinkers, including Wittgenstein, Austin, Quine, and Derrida, to present an encompassing view of how meaning is formed and interpreted. Concepts such as semiotics, hermeneutics, and speech act theory play central roles in deciphering the mechanisms through which language represents reality. The complexities of truth in language manifest in various real-world phenomena, including political discourse, media manipulation, translation challenges, and legal interpretation. This book investigates how different languages construct reality in unique ways, leading to variations in perception and understanding. It also considers the implications of artificial intelligence in language processing and whether machines can ever truly comprehend truth. By integrating classical and modern linguistic theories, this book establishes a framework for analyzing truth in communication. It presents a structured approach to evaluating how meaning is derived, how context shapes interpretation, and how linguistic structures influence perception. The core principle is that truth is not merely an objective reality but is also shaped by the languages and systems within which it is communicated. Key indicators of linguistic truth include coherence, consistency, factual alignment, and pragmatic effectiveness. This book identifies operational variables such as cultural context, linguistic ambiguity, speaker intent, and audience interpretation as fundamental factors in understanding how truth is conveyed through language. Numerous elements influence the relationship between language and truth, including cognitive biases, historical contexts, power structures, and evolving linguistic norms. By analyzing these factors, this book provides insights into why different cultures and societies perceive truth differently and how language can both reveal and obscure reality. Applying the theoretical insights from this book, readers are guided through strategies for enhancing clarity, reducing misinterpretation, and fostering more precise communication in various fields, from academia to politics. The book also addresses how language policies, media regulations, and ethical considerations shape the pursuit of truth in public discourse. While language serves as a bridge to truth, it is also a source of distortion, manipulation, and misunderstanding. This book highlights the challenges posed by misinformation, propaganda, and ideological biases, while also recognizing the supporting role of linguistic diversity and education in fostering more nuanced truth-seeking. Truth and language are inseparable in human thought and communication. Through this in-depth exploration, the book underscores the importance of linguistic awareness in navigating the complexities of truth. Whether in philosophy, politics, law, or everyday conversation, understanding how language constructs and conveys truth is vital for clearer and more meaningful human interactions.
Vision-language models (VLMs) show promise as tools for inferring affect from visual stimuli at scale; it is not yet clear how closely their outputs align with human affective ratings. We benchmarked nine VLMs, ranging from state-of-the-art proprietary models to open-source models, on three psycho-metrically validated affective image datasets: the International Affective Picture System, the Nencki Affective Picture System, and the Library of AI-Generated Affective Images. The models performed two tasks in the zero-shot setting: (i) top-emotion classification (selecting the strongest discrete emotion elicited by an image) and (ii) continuous prediction of human ratings on 1-7/9 Likert scales for discrete emotion categories and affective dimensions. We also evaluated the impact of rater-conditioned prompting on the LAI-GAI dataset using de-identified participant metadata. The results show good performance in discrete emotion classification, with accuracies typically ranging from 60% to 80% on six-emotion labels and from 60% to 75% on a more challenging 12-category task. The predictions of anger and surprise had the lowest accuracy in all datasets. For continuous rating prediction, models showed moderate to strong alignment with humans (r > 0.75) but also exhibited consistent biases, notably weaker performance on arousal, and a tendency to overestimate response strength. Rater-conditioned prompting resulted in only small, inconsistent changes in predictions. Overall, VLMs capture broad affective trends but lack the nuance found in validated psychological ratings, highlighting their potential and current limitations for affective computing and mental health-related applications.
This study reveals a critical paradox in social media privacy communication: Although platforms like Meta (Instagram and Facebook), TikTok, and X have evolved their policies in an effort towards simpler, standardized disclosures, the language remains cognitively inaccessible to their core adolescent audience. Our analysis demonstrates that these disclosures, benchmarked against the developmental norms of 13–17‐year‐olds, are written at a university‐level complexity, calling into question the validity of informed consent for minors. We use a triangulated method to assess the accessibility of platform policies for teens. Structural mapping shows consistent topic coverage, but readability indices indicate a college‐level reading requirement. Lexical analysis confirms high rates of difficult words, exceeding the threshold for adolescent understanding. Our findings lead to a sobering conclusion: The prevailing model of using a single, text‐based privacy policy is caught in an inherent tension between legal completeness and adolescent comprehension, making it fundamentally unworkable. This research provides evidence that calls for the need for a redesign of privacy communication for minors or a reconsideration of the current minimum age for digital consent.
в статье представлено исследование, посвящённое сопоставлению подходов к обучению иностранному языку студентов направления «Зарубежное регионоведение». Предмет анализа связан не просто с овладением языковой нормой иностранной речи, а с формированием такой модели речевой подготовки, при которой студент способен соотносить высказывание с конкретным регионом, его политико-культурной спецификой, медийной повесткой и типичными коммуникативными сценариями. С помощью методов анализа, синтеза, наблюдения и описания произведено рассмотрение актуальных на данном этапе развития высшего образования способов формирования иноязычной региональной компетенции. С позиции компетентностного и деятельностного подходов обучение рассматривается как движение от языковой операции к регионально маркированному высказыванию, которое строится в ситуации обсуждения, аргументации, интерпретации и переговоров. Сложный характер данной компетенции требует использования в процессе преподавания разнообразных методов активного и интерактивного обучения, инновационных образовательных технологий, форм и средств обучения, тесно связанных с будущей профессиональной деятельностью студентов направления подготовки «Зарубежное регионоведение». this article presents a study comparing approaches to foreign language instruction for students majoring in “Foreign Regional Studies”. The subject of analysis is not merely the mastery of the linguistic norms of Chinese speech, but the development of a model of language training in which students are able to relate a statement to a specific region, its political and cultural characteristics, media agenda, and typical communicative scenarios. Using methods of analysis, synthesis, observation, and description, this study examines the methods of developing foreign language regional competence that are relevant at this stage of higher education development. From the perspective of competence-based and activity-based approaches, teaching is viewed as a progression from linguistic operations to region-specific utterances, which are constructed in situations of discussion, argumentation, interpretation, and negotiation. The complex nature of this competence requires the use in the teaching process of a variety of active and interactive teaching methods, innovative educational technologies, and forms and means of instruction closely linked to the future professional activities of students training program "Foreign Regional Studies".
This article offers an analysis of how dispositional language is used in biology, psychology, and the social sciences. The purpose of the article is to answer, firstly, the question of how dispositions and things similar to them are understood by scientists, and secondly, to answer the questions of what terms dispositions and disposition-like entities are used to designate, and how these terms relate to each other. The variety of scientific approaches to understanding behavioral dispositions indicates that there exists between them a family resemblance that is not defined in terms of necessary and sufficient properties. A more realistic approach to understanding behavioral dispositions is to describe them in terms of examples and partial generalizations. From this point of view, a family of behavioral dispositions can be described as a set of patterns, tendencies, or causes of probable behavior observed or predicted under certain circumstances. At the same time, in the behavioral sciences a number of terms are used that denote behavioral dispositions and similar things in different languages, which form sets of synonyms – synsets. Thus, in the Russian and English languages there are a number of synsets presented in the lexical databases WordNet and RuWordNet and denoting the entire family of behavioral dispositions or some of its subsets. The key semantic idea here is that it is not so much the individual synonyms included in a synset that have meaning, but rather the synset as a whole or even a system of synsets.
The paper presents a prototype of a web-app designed to automatically generate verb valency lexica based on the Universal Dependencies (UD) treebanks.It offers an overview of the structure of the app, its core functionality, and functional extensions designed to handle treebank-specific features.Besides, the paper highlights the limitations of the prototype and the potential of its further development.
The rapid development of large language model technology has evolved machine translation from a low-level tool into a cultural transmission vehicle with semantic understanding capabilities, shifting the relationship between artificial intelligence and human translators from one of substitution to one of collaboration.Employing Translator Behavior Criticism theory and comparative analysis, this study systematically analyzes the behavioral characteristics of student translators and multi-model machine translators across the two dimensions of "truth-seeking" and "utility-attaining," revealing the differential patterns between human and machine translators in three aspects: semantic fidelity, cultural adaptability, and audience orientation.The findings indicate: 1) Student translators demonstrate stronger subjectivity in terms of cultural awareness and ideological expression, enabling a deeper grasp of the philosophical connotations and value orientation of terminology; 2) Machine translators hold significant advantages in lexical innovation and adaptation to linguistic norms, yet exhibit notable limitations in understanding complex rhetorical structures and cultural metaphors; 3) Humanmachine collaborative pathways can achieve a more optimal balance of tension between preserving Chinese characteristics and achieving international accessibility, forming a bidirectional enhancement effect characterized by "complementarity between truth-seeking and innovation, and integration of utility-attaining and flexibility"; 4) A collaborative translation system requires the construction of a three-tier progressive mechanism of "multi-model inspiration-in-depth student revision-expert feedback optimization" to realize the organic unity of cultural confidence and international communication.
The article analyzes the influence of economic factors on the language attitudes of youth. According to the theory of P. Bourdieu, the dominance of linguistic norms and forms is viewed as a factor that exacerbates social inequality. Proficiency in different languages increases an individual's social capital and expands their economic opportunities, while language barriers restrict access to these resources. The language choice among young people is largely determined by their economic status. It is posited that income levels facilitate the learning of foreign languages, whereas, in conditions of social inequality, the ability of youth to maintain their native language is taken into account. The study examines the impact of economic factors–such as labor market requirements, educational opportunities, and income levels–on multilingualism and language choice among the younger generation. The research provides insight into how economic drivers influence language choice, language policy, and the acceptance of multilingualism in society. The author presents the results of applied research based on the focus group method. Focus groups were conducted across 12 regions (N=167). According to the results, language choice among youth depends on regional and ethno-demographic characteristics. Furthermore, the global economy and globalization trends push young people toward learning multiple languages, while disparities between urban and rural areas also affect language attitudes. While youth with high-income levels strive for multilingualism, low-income groups prioritize their native language. The findings of this study play a crucial role in forming effective state and educational language policies that can enhance the success of young people in social and professional life.
This repository contains GSD-NP and GSD-DiNoS, both derived from Universal Dependencies' (UD) GSD Treebank. GSD-NP (.conllu) is a subset of UD-GSD and comprises its simplex noun phrases (NP): Common nouns (NN/NOUN) and their direct dependents (determiners, adnominal adjectives, nmods, adpositions, adverbs). It consists of 49,425 NPs (119.0k tokens) and has an improved feature annotation coverage (gender, case, number). Breaking with UD annotation, a total of 3,649 APPRART tokens were reconstructed in GSD-NP to restore the original orthographic forms. GSD-DiNoS (.json) is a custom data-driven lexion-like data structure built on GSD-NP, which aggregates NPs with the same head lemma. For each lemma, absolute frequencies of the lemma and its word forms are captured. Moreover, each occurrence feeds into three areas of interest within the word form entry: morphosyntactic features in isolation (gender, case, number), in combination with groups of dependents (collocations), and in combination with the syntactic function (dependency relations). GSD-DiNoS spans 17,433 unique lemmas and 20,190 unique word forms, stemming from 49,416 NPs. Lemmas were relemmatised to assign unique lemmas to nominal compounds, a highly productive and often lexicalised construction in German.
Humans are inherently social beings, and social cues such as faces and voices guide attention and behavior. Auditory perception, especially binaural hearing, is essential for social cognition, enabling sound localization and speech comprehension in noisy environments. Deficits in auditory processing can impair social functioning, and conditions such as social anxiety are linked to reduced social functioning. Since social functioning is closely linked to overall well-being, improving social behavior represents a key objective in psychological research. Virtual reality (VR) is increasingly used to study social behavior due to its flexibility and ecological validity. However, users often report limited social presence, reducing the effectiveness of VR-based interventions especially for social anxiety. One reason may be the dominance of visual over auditory realism: audio is often presented in mono or stereo, reducing naturalness and presence. Binaural auralizations, which provide realistic, externalized spatial audio, may enhance presence and support virtual social interactions. This thesis pursues four main research objectives: identifying suitable behavioral and subjective measures for evaluating binaural realism; assessing immersion, realism, and audio quality across auralization techniques; comparing synthetic and natural speech in a socially stressful VR scenario; and examining effects of binaural audio on affect, presence, and attention under varying social stress levels. Study 1 examined how the virtual visual scene and measurement method affect localization and distance perception of physical sound sources. Across two experiments (N=60), audiovisual incongruence reduced localization accuracy but did not affect presence or realism. Distance estimation was influences by the interaction of task and scene: overestimation increased when using a placement task in a reduced-visibility scene. Study 2 compared localization accuracy for loudspeakers and four virtual audio renderings using a placement task and a gaze-based paradigm (N=49). Binaural renderings produced slightly lower localization accuracy but similar ratings of social presence and realism. A simple generic rendering performed as well as more complex ones. Only the anchor condition lacked externalization and was inferior across measures. Social presence and subjective realism were strongly correlated. Study 3 compared AI-generated text-to-speech with natural human speech in the Trier Social Stress Test (N=40). Both conditions elicited substantial stress responses and produced similar presence and affect ratings, demonstrating the practicality of synthetic speech in virtual social interactions. Study 4 investigated audiovisual realism in a virtual social stress scenario (N=78). A high-stress group showed stronger physiological and subjective stress responses than a low-stress group. Binaural audio increased perceived realism and externalization but did not affect social presence, stress responses, or gaze behavior. High arousal across all groups may have masked audio effects. Across all 4 studies, social anxiety did not consistently affect auditory perception or presence but influenced affective states and subjective evaluations of the interaction. Overall, the findings highlight the importance of VR-specific auditory perception and the role of acoustic immersion. Auditory realism enhances social and physical presence, though its impact varies by context. It appears most effective in low- to moderate-arousal scenarios and may be less critical in highly affective VR applications such as anxiety treatments. Practical advancesn such as TTS integration and simplified binaural rendering methods can support the broader use of realistic audiovisual VR environments in psychological research.
This study investigates the diachronic drift of Arabic future-marking parti cles, empirically testing the shift from synthetic (سـ, سوف) to analytic (راح, قاعد+ح) forms across registers and regions. Leveraging a multi-register corpus (Penn Arabic Treebank, Corpus of Contemporary Arabic, Arabic Gigaword Fifth Edition) and advanced NLP tools (MADA, CAMeL Tools, AraBERT), we extracted and analyzed future-marker tokens annotated for register (newswire, opinion, social media, religious), region (EG, SA, LB, MA), and time slice (1990-2000, 2000-2010, 2010-2023). Mixed-effects logistic regression revealed significant effects of time, register, and region, confirming a clear diachronic shift towards analytic markers, particularly راح and قاعد+ح. Interaction terms highlighted that this shift is more pronounced in informal registers and certain regions, indicating dialectal pressure and diffusion of innovations (e.g., Gulf Arabic راح) into broader usage. Hierarchical clustering of contextual embeddings would further validate semantic-pragmatic shifts. This research provides robust evidence for ongoing linguistic change in Arabic, contributing to theories of grammaticalization and language contact.
Abstract: The Emperor Marcus Aurelius and the former slave Epictetus represent the social poles of the Roman Empire, yet both are cornerstones of late Stoic thought. This study employs digital humanities tools to investigate how their disparate life experiences and professional roles produced divergent philosophical "signatures" in their extant literature. By analyzing the Lemmatized Ancient Greek Texts (LAGT) corpus, we identify a distinct linguistic polarity: Marcus Aurelius demonstrates a significant preference for physical and cosmological terminology, reflecting a Stoicism centered on the providential order of the universe. Conversely, Epictetus’s lexicon shifts toward terms of ethical practice, pedagogy, and the transformation of the moral will. While Marcus Aurelius employs a more poetically diverse and intellectually wide-ranging vocabulary, Epictetus utilizes a more repetitive, concentrated technical vocabulary suited for the classroom. Despite these differences, a high degree of overlap reveals a "common core" of Stoic concepts shared by both authors, such as the nature of impressions and the primacy of the divine. These findings quantitatively highlight the adaptability of Stoicism, illustrating how a robust philosophical core was reframed to serve both the private reflections of a struggling ruler and the public exhortations of a committed teacher. Technical Context & Methodology This research integrates philology with a computational pipeline to analyze late Stoic literature. The following technical components are included in this repository: Computational Environment: All analyses were performed using Python 3.11. The pipeline utilizes Pandas and PyArrow for high-speed data processing, and the Classical Language Toolkit (CLTK) for part-of-speech tagging and grammatical filtering. Corpus Data: The primary linguistic data was extracted from the Lemmatized Ancient Greek Texts (LAGT) v4.1 dataset, which provides advanced lemmatization via the GLAUx treebank and GreCy models. Lexicographical Mapping: English definitions were integrated using the LSJ Dictionary (JSON v1.0.0). A custom normalization pipeline was used to standardize lemmata into Normalization Form Canonical Composition (NFC). Lexical Metrics: Vocabulary richness was assessed using Type-Token Ratio (TTR), Guiraud’s Index (R) to compensate for corpus size differences, and the percentage of hapax legomena (terms appearing only once). Generative AI Integration: A Gemma-3-27b-it model was utilized for the thematic classification and translation of 5,371 sentences. Sentences were tagged into the traditional Stoic tripartite division—Logic, Physics, or Ethics—based on the framework established by Pierre Hadot. Visualizations: The included scripts generate Lexical Volcano Plots (mapping total relative frequency against authorial skew) and Weighted Word Clouds that distinguish between author-specific signatures and the "Shared Stoic Core". Files included in this record: Supplementary File S1: Complete Python computational pipeline, README, and requirements.txt. Supplementary File S2: stoic_master_comparison.tsv containing comprehensive lemma frequencies and delta-RF values. Supplementary File S3: Statistical visualizations, including KDE overlap plots and delta-RF histograms. Supplementary File S4: Thematic analysis CSV containing 5,371 sentences with original Greek, English translations, and AI-generated thematic tags.
Author: Denny van Gulik Methodology: The 80/20 Matrix Status: Version 1.0 (Linguistic Corpus) 1. Abstract (Exposé) This work presents a comprehensive decoding of the Voynich Manuscript (MS 408). Moving beyond traditional cryptographic attempts, this research approaches the codex from a technical and structural perspective. The manuscript is identified as a functional pharmaceutical and balneological manual of a late medieval scholarly brotherhood, likely operating within a courtly or monastic context (Palar). The core of this discovery is the 80/20 Matrix: 80% Phonetically Deformed Latin: Technical terms of medieval botany and medicine, obscured through systematic phonetic shifts and the specific EVA character set. 20% Balkan Regionalisms: Use of regional terminology (e.g., Amum for water, Otlar for herbs, Pala for court/palace) serving as bridge vocabulary. Statistical validity is maintained across all 246 pages, identifying complex processes of thermal extraction (Pokedum), honey-based preservation (Melle), and advanced hydrotherapeutic systems. 2. Methodological Transparency (Authorship & AI Usage) Important Note on Research Genesis: The discovery of the 80/20 Matrix and the linguistic identification of the Balkan-Latin hybrid system is the exclusive intellectual property and original work of Denny van Gulik. Artificial Intelligence (specifically the Google Gemini model) was utilized strictly as a digital research assistant and scaling tool. Its role was limited to: Formatting manually decoded data into scientific tables and HTML. Cross-referencing author-identified word stems with linguistic databases. Translating research notes into academic English to facilitate international peer review. The logic, intuition, and systematic pattern recognition are entirely human-led. This project is not a result of "AI hallucination" but a rigorous analysis of the codex as a logistical document. 3. Project Roadmap & Updates Current Version (v1): Focuses on the textual corpus, the 80/20 linguistic matrix, and the primary glossary. Upcoming Version 2.0: Will feature fully integrated high-resolution folio images and direct visual cross-references for every analyzed page. English Edition: A full, 246-page English translation of the entire study is currently in progress to ensure accessibility for the global scientific community. 4. Keywords Voynich Manuscript, MS 408, 80/20 Matrix, Medieval Medicine, Balkan Linguistics, Codicology, Balneology, Historical Pharmacy. Deutsche Zusammenfassung: Dieses Projekt präsentiert die vollständige Dekodierung des Voynich-Manuskripts mittels der 80/20-Matrix (deformiertes Latein & Balkan-Regionalismen). Es identifiziert das Werk als pharmazeutisches Handbuch einer spätmittelalterlichen Bruderschaft. Version 1.0 sichert die linguistische Priorität; Version 2.0 mit Bildreferenzen sowie eine vollständige englische Übersetzung folgen in Kürze.
BACKGROUND: This pilot randomized controlled trial evaluated the effectiveness of an artificial intelligence (AI)–assisted solo workflow for intraoral photography training. The study examined whether real‑time AI feedback could enhance photographic quality, procedural efficiency, learner self‑efficacy, and patient comfort compared with conventional approaches. METHODS: Fifty-four first‑year dental students were randomly assigned to one of three groups: assistant‑supported workflow (four‑handed technique, control), solo workflow without AI support, and solo workflow with AI‑driven real‑time feedback. All participants performed standardized intraoral photography tasks. The primary outcome was a composite photographic quality score derived from expert ratings of three standardized intraoral views (frontal intercuspal, frontal open-bite, and lateral intercuspal), each rated on a 0–10 scale (total range 0–30). Data were analyzed using ANOVA; mean differences (MD) with 95% confidence intervals (CI) were calculated. RESULTS: Inter‑rater reliability for expert image ratings was good (ICC = 0.84, 95% CI: 0.72 to 0.90). The AI-supported solo group achieved the highest composite quality scores (18.2 ± 2.7). This was significantly superior to the unassisted solo group (15.8 ± 3.3), with a mean difference (MD) of 2.4 points (95% CI: 0.45 to 4.35; p = 0.027) and a large effect size (Cohen’s d = 0.80). Compared to the assistant-supported group (17.1 ± 2.1), the difference was not statistically significant (MD = 1.1; 95% CI: -0.65 to 2.85; p = 0.28). Secondary outcomes, including task completion time (F(2,51) = 1.25, p = 0.30), self‑efficacy (all p > 0.40), and patient‑reported comfort (χ²(4, N = 54) = 5.2, p = 0.27), showed no significant between‑group differences. CONCLUSION: In this single‑centre pilot trial, an AI‑assisted solo workflow enabled novice dental students to achieve higher intraoral photographic quality than unguided solo operation, with performance broadly comparable to a conventional four‑handed assistant‑supported workflow and without detectable compromises in efficiency, self‑efficacy, or patient‑reported comfort. These preliminary findings may serve as a valuable adjunct for autonomous skill acquisition, warranting further validation in larger, multi-institutional cohorts. CLINICAL TRIAL NUMBER: Not applicable. This study evaluated an educational training intervention rather than a clinical treatment, and prospective trial registration was not required under institutional policy at the time of initiation. Ethical approval was obtained from Shanghai Ninth People’s Hospital Ethics Committee (SH9H-2022-T30-1).
This study presents a lexicon-based semantic tagging information system developed for the Uzbek language corpus. The system employs the six-volume Explanatory Dictionary of the Uzbek Language (OʻzTIL) as its primary lexical resource, which contains over 85,000 entries with full semantic definitions, making it the most authoritative normative lexicographic source for Uzbek. An ontological model organized in three hierarchical levels – top, mid, and low – was designed to categorize lexical units extracted from the dictionary. Five core semantic categories were formed: animal names (approximately 100–200 units), bird names (approximately 100–150 units), personal nouns (approximately 500+ units), place names (approximately 300+ units), and occupation names (approximately 200+ units), totaling approximately 1,200–1,400 lexical units. A rule-based automatic tagging algorithm was developed to annotate corpus tokens against this structured lexical database, assigning standardized semantic tags. The system addresses key challenges inherent to Uzbek, including agglutinative morphology and lexical ambiguity. Compared to international systems such as WordNet and USAS, the proposed dictionary-based approach demonstrates superior normative grounding and cultural adequacy for Uzbek. The system is intended to serve as a foundational open resource for downstream natural language processing tasks, including machine translation, information retrieval, and intelligent educational applications.
Paper 6 (Silva 2026) introduced BPE Mean Vocabulary Morpheme Length (VMML) as a writing system classifier and showed that the Voynich Manuscript occupies a discriminant zone (VMML = 5.918, 95% CI 5.77-6.05) above all 15 tested alphabetic natural languages. This paper (v2.5) expands to 71 corpora across 40+ languages and reports six extended analyses: (1) Alphabetic ceiling confirmed at 5.76; (2) Tagalog (VMML=5.914) is the sole natural-language entry into the Voynich CI, but BC=0.202 distinguishes it from Voynich (BC=0.361); (3) Romanization inflates VMML by 2.4-5.3 units (methodological confound). Extended analyses: (4) Currier A vs B: delta VMML=+1.27, delta CBMI=+0.16 bits -- two quantifiably distinct writing registers; (5) BC coherent across all 7 manuscript sections (CV=6.7%) -- single writing system confirmed; (6) 3D discriminant (VMML x BC x CBMI): Voynich isolated, nearest natural-language neighbor Irish at distance 0.17; (7) Six named hoax mechanisms (monoalphabetic, Vigenere/barbavara, Vigenere/Italian-Knowles 2026, null insertion, syllabic compression, vocabulary shuffle) each fail all three criteria simultaneously; (8) BC orthogonal to all classical textual metrics (|r| < 0.23 vs entropy, TTR, hapax, Zipf) -- genuinely new structural dimension. All code and six extension scripts publicly available in companion repository. v2.3 (2026-06-08): Section 5.9 added - per-folio Currier A/B reanalysis using the Gaskell and Bowern (2022) canonical corpus (36,361 tokens, min_freq=5 BPE). Cross-boundary mutual information (CBMI) identified as primary discriminant: CBMI_A = 1.97 bits vs CBMI_B = 1.51 bits, Cohen d = -1.01, permutation p less than 0.001 (n = 10,000 shuffles, Bonferroni-corrected). CBMI survives within-quire control (pooled nA=46, nB=33; permutation p = 0.0008; Fisher combined within-quire p = 0.001), ruling out manuscript section as a confound. All three metrics (BC, BPE-ratio, CBMI) show A greater than B direction. Fisher combined full-corpus: chi-squared(6) = 40.66, p less than 0.000002. Section 5.1 corrected: direction is A greater than B on BC and CBMI. Finding is orthogonal to Parisel (2026) vowel-selection model. Conclusion 12 added. v2.4 (2026-06-09): §5.10 added — Currier-preserving null model (n = 200 iterations, size-matched) quantifying each metric's section-discrimination sensitivity independently of dialect. Key result: CBMI is the weakest section discriminant (mean |z| = 1.20 across six sections), confirming that the large CBMI A/B gap (§5.9) is not a section-composition artifact. STTR@100 is the strongest section discriminant (mean |z| = 4.75). Herbal section shows anomalously low vocabulary diversity (STTR z = -13.9); Stars shows anomalously high unique vocabulary (Hapax@500 z = +4.4). Demonstrates two independent organizational layers: CBMI tracks dialect, STTR tracks content domain. Conclusion #13 added. v2.5 (2026-06-10): Corpus expanded from 55 to 71 corpora across 40+ languages. §5.11 adds five medieval European corpora in native script via Universal Dependencies treebanks (Gothic transliteration, Old Church Slavonic, Old East Slavic, Ancient Greek PROIEL and Perseus; VMML 3.54-5.18 — all below alphabetic ceiling of 5.748). §5.12 adds 11 Australian Aboriginal language corpora via BibleNLP/eBible (Pama-Nyungan Western Desert, Ngumpin-Yapa, Arandic; Yolngu; Gunwinyguan; Daly; VMML 6.09-8.00 — predominantly above the Voynich zone). Warlpiri (VMML 5.851) is the sole near-entry on VMML but fails BC (0.233) and CBMI (0.244); 3D normalized distance from Voynich = 0.746 (vs. Irish = 0.200, the nearest neighbor from §5.4). Voynich zone is now charted on both sides: fusional alphabetic below (VMML 3.5-5.75), agglutinative-to-polysynthetic above (VMML 6.0-8.0). Voynich occupies a structural configuration not replicated by any of the 71 corpora tested. To our knowledge, this is the first systematic BPE profiling of Pama-Nyungan languages in the computational linguistics literature. Conclusions #14 and #15 added. v2.6 (2026-06-12): Section 5.10.1 adds a prose-only robustness check for the Section 5.10 Currier-preserving null model. Excluding all label, circular and radial loci (8.7% of tokens), every headline deviation survives essentially unchanged: Herbal STTR z = -13.5, Balneological z = -10.7, Stars Hapax z = +4.3; the sensitivity ranking is unchanged with CBMI last in both conditions. A mean-vs-median distributional note (both summaries rank lexical-diversity metrics first, boundary metrics last) and a coverage note (Astro/Zodiac folios carry no Currier tags and are outside any Currier-preserving design) are added. Erratum: Section 5.10 folio count corrected to 226 parsed / 196 Currier-labeled.
Abstract Introduction Sleep supports emotion regulation by preferentially consolidating emotional memories while attenuating reactivity. We have shown that dream recall plays an active role by increasing negative over neutral memories and reducing reactivity. In women, fluctuating reproductive hormones across the menstrual cycle influence sleep features implicated in emotional memory, yet whether menstrual phases influence how dreams shape emotional processing remains unknown. This study investigates how dreams shape sleep-dependent emotional processing across the menstrual cycle in naturally cycling women. Methods 128 women (Mage = 32.85 ±11.93 years) completed up to four visits across verified menstrual phases (menses, late-follicular, mid-luteal, late-luteal). At each visit, participants performed the Emotional Picture Task with negative and neutral IAPS images in the evening (Test 1) and the next morning (Test 2). Participants rated old/new, arousal, and valence of images shown at each test. Dream reports were collected upon waking prior to Test 2. Linear mixed-effects models tested main and interaction effects of menstrual phase and dream recall. Results The menstrual cycle altered how dreaming shaped overnight emotional memory. Dream recall typically benefited the emotional trade-off effect —favoring consolidation of negative relative to neutral images (Δd′; t(410)=1.95, p=0.05)—but this pattern reversed during the late-luteal phase (dream × menstrual cycle: t(381)=-2.29, p=0.02). Dreaming showed independent effects on emotional reactivity. Higher valence and arousal ratings for negative images during Test 1 predicted greater dream recall (valence: t(344)=2.05, p=0.04; arousal: t(327)=2.04, p=0.04). Additionally, the more negatively participants rated the images at Test 1, the more negative their dreams tended to be (t(166)=-2.11, p=0.04). Dream recall was linked to reduced next-morning emotional reactivity (valence: t(413)=-2.89, p=0.004; arousal: t(413)=-2.65, p=0.01), with stronger reductions following more negative dreams (β=0.15, t(182)=2.86, p=0.005). Conclusion Menstrual cycle phase influenced how dreams shaped overnight emotional memory. Negative waking experiences increased dream recall and shaped dream content—and recalling dreams, especially negative ones, reduced emotional reactivity and typically strengthened emotional memory—but this benefit disappeared in the late-luteal phase when there are declining reproductive hormones. These findings suggest a novel interaction between the menstrual cycle and dreaming, showing that hormonal fluctuations reshape how sleep and dreams regulate emotional experience and memory. Support (if any) RF1AG061355 (Baker/Mednick)
This study examines the current state and future prospects of Urdu digital translation within the broader historical and technological development of machine translation. It begins by outlining the evolution of digital translation systems and reviewing their application across major world languages, followed by a critical analysis of existing Urdu translation platforms such as Google Translate, Bing Translate, ChatGPT, and other AI-based tools. The research identifies key linguistic and technical challenges that affect Urdu translation quality, including script directionality, morphological and syntactic complexity, polysemy, idiomatic expressions, cultural references, tokenization and parsing difficulties, and Unicode compatibility issues. By situating Urdu within the framework of Artificial Intelligence (AI) and Natural Language Processing (NLP), the study highlights the need for language-specific AI models, large-scale corpora, annotated treebanks, and domain-sensitive lexical resources to improve translation accuracy and contextual coherence. It further explores the applicability of advanced language models such as BERT, LLaMA, and generative AI systems in enhancing Urdu machine translation. In response to the identified limitations, the research proposes a corpus-driven, AI-integrated Urdu translation web application framework designed to provide context-aware, stylistically appropriate, and semantically accurate translations. The study contributes both analytically and practically by offering a comprehensive evaluation of Urdu digital translation and presenting a scalable model aimed at strengthening Urdu’s position in the global digital and AI-driven linguistic landscape.
Heritage grammars tend to undergo structural change owing to their severely constrained input conditions and /or transfer effects from the L2 (Polinsky, 2018). This study uses adjectives in Tamil (Dravidian) to show a systematic difference between rule-governed, structural aspects of grammar and those that require case-by-case lexical learning. The former remains stable and the latter undergoes change in heritage Tamil. The empirical domain of adjectives in Tamil is novel and particularly informative, as the derived nature of these adjectives helps us identify areas of grammatical stability and those of change when the context of acquisition diverges from the norm, i.e, heritage grammars. The paper has two major aims: (i) to provide an explanation of adjectives in standard Tamil, and (ii) to inquire into how the derivation of adjectives fares in the context of heritage Tamil. (i) is addressed by showing that adjectives in Tamil are not an independent category in the lexicon, but the derivational component recognises them as a distinct category. We then proceed to question (ii): With respect to heritage Tamil, two domains — one requiring intensive learning, and not requiring learning — are identified. The paper provides novel empirical evidence to demonstrate their stability or variation in heritage grammars.
ABSTRACT Classical fear conditioning describes how neutral cues acquire a threat value, yet how learned associations are retrieved and generalised across similar stimuli specifically is an ongoing debate. We combined behavioural ratings, physiological measures, and fMRI in a two-day classical fear conditioning paradigm to characterize acquisition, retrieval, and generalisation across modalities. Twenty-five healthy participants completed acquisition trials on Day 1 and retrieval and generalisation trials on Day 2 using conditioned (CS+, CS-) and graded generalisation stimuli (GS). Outcomes included trial-wise US expectancy ratings, pre/post fear and arousal ratings, skin conductance responses (SCR), pupil dilation, and ROI-based fMRI (amygdala, hippocampus, insula, periaqueductal gray (PAG), locus coeruleus). Acquisition yielded robust CS+/CS-discrimination in behavioural ratings and increased BOLD responses in bilateral insula and PAG. During retrieval, US-expectancy ratings indicated early retrieval of CS contingency. The fMRI results showed greater BOLD activity during CS+ presentations than during CS-presentations in bilateral hippocampus, left insula and right PAG. Additionally, hippocampus-insula coupling increased. Critically, parametric modulation during retrieval revealed that trial-wise mean US-expectancy modulated BOLD responses in left insula and right PAG, with a trend in left hippocampus. Across generalisation, US-expectancy and pupil dilation responses followed graded profiles, which could be explained by a Gaussian model, whereas SCR generalised, but was not captured by a Gaussian model. Parametric modulation by US-expectancy correlated with BOLD activity in left PAG, with a trend in right hippocampus. Stimulus identity explained variance in bilateral insula and left PAG. Findings converge on a hippocampus-insula-PAG network that retrieves learned predictions, and scales defensive output according to similarity-based threat probability, linking subjective, physiological, and neural outcomes.
The limitless semantic potencies of communication is within the framework of the language conventional semantics, which imposes a number of restrictions, including on the explication of emotional experiences by the speaker. The latter either chooses a read y-made preset formula, or directs communicative efforts to search for and objectify emotional and semantic shades of meaning with an uncodified form of verbalization. If the form of expression of an emotional experience is new, atypical, unconventional, we should talk about the representation of diffuse emotive semantics, approaching the actual emotional experience. Diffusivity (fuzziness, vagueness, multiple inconsistencies, ambiguity) is an immanent property of semantics that corresponds to both the natur e of the sign and the environment in which the sign acts. The assumption is that, depending on the characteristics of the discourse and the genre characteristics of the elements included in it, artistic communication was considered. The author of a work of art must go beyond the linguistic prescription, which allows him to have the desired effect on the addressee. The analysis of a dramatic work shows a variety of forms of explication of diffuse emotivity, when emotional experiences become a discursive and genre-forming category: the expression of complex vague emotions allows creating an image of a multifaceted and interesting character. The texts of modern plays allow tracing a similar trend towards diversifying the form of expression of diffuse emotivity, but the emotional tonality is less diverse: negative emotional experiences set the emotional dominant, therefore, the ‘consolidation’ of the emotive occurs rather than through the vector of mixing positive and negative assessments, but the intensification and concretization of the negative evaluative component. The author also postulates that language always approximately describes emotions, but in artistic communication such approximativeness is expressed in conscious and creative imitation, which transforms and develops the linguistic norm.
This study examines the discursive construction of sexism in Moroccan football fandom through a digital ethnography of online posts and stadium banners. Drawing on Facebook posts and widely circulated Ultras banners, the analysis explores how gendered exclusion is produced and normalized in contemporary fan communities. Using Teun A. van Dijk's Critical Discourse Analysis (CDA), the study examines lexical choices, syntactic patterns, and rhetorical devices—such as epiphora, metaphor, and hyperbole—that portray women as biologically unfit, morally loose, or out of place in stadiums. At the meso-level, the analysis reveals shared social cognitions that position women as an out-group whose presence threatens the imagined authenticity of male fandom. At the macro-level, informed by feminist theory, the findings show how these discourses reproduce broader patriarchal norms in Moroccan society, including gendered gatekeeping of public space, moral policing of women's bodies, and the use of female kinship “sisters” as tools for male-to-male humiliation. The findings demonstrate that sexist fan discourse operates as a patterned ideological practice that contributes to the exclusion of women from Moroccan football fandom and public life.
This paper presents a small-scale dependency treebank for Tunisian Arabic (TADT) developed within the Universal Dependencies framework, addressing the scarcity of linguistic resources for the Arabic varieties.The approach employs domain adaptation, leveraging a machine learning model (UDPipe 1.0) trained on Algerian Arabic data to annotate 100 Tunisian Arabic social media comments, followed by manual correction.This pilot study evaluates the feasibility of using machine learning-assisted annotation to scale resource development for spoken Arabic and identifies key challenges in cross-dialectal transfer for improving annotation quality and efficiency.This work contributes to more inclusive and fair representation of Arabic linguistic varieties in academic research and NLP applications.
Conventional Kansei Engineering (KE) models treat morphological components as equally weighted predictors of affective response, implicitly assuming that fixation volume indexes perceptual significance. This study proposes an attention-weighted KE framework integrating AOI-based eye-tracking evidence into morphological variable weighting prior to form–emotion modelling. Ten Ming-style chair samples were encoded into 30 binary morphological parameters; affective ratings were collected via semantic differential survey (N = 389) and element-level fixation distributions via eye-tracking (N = 30), from which coefficient-of-variation-derived weights were assigned. A critical dissociation emerged: the backrest dominated fixation (50.01%) yet received the lowest perceptual weight, while the head-rail and front arm support received the highest — demonstrating that fixation dominance and perceptual weighting are decoupled constructs. Attention-weighted models achieved 83.3% predictive consistency across five of six Kansei dimensions (p &gt;.05). Findings offer a quantitative method for integrating visual attention into KE modelling, with implications for perception-informed design.
BACKGROUND: Major depressive disorder (MDD) is characterized by pervasive cognitive-emotional biases, yet the spatiotemporal dynamics of prefrontal involvement during emotional association remain poorly understood. This study aimed to delineate temporal deviations and spatial reweighting of prefrontal activation in MDD and evaluate their relationship to behavioral bias and clinical symptom burden. METHODS: Sixty-nine participants (47 MDD, 22 healthy controls) completed the Free Association Semantic Task (FAST) during functional near-infrared spectroscopy (fNIRS) recording. Participants generated 10 associations in response to neutral, positive, or negative cue words. Emotional valence ratings were derived for each association, and group differences were assessed using independent-samples t-tests. Prefrontal hemodynamic responses were preprocessed and compared across six regions of interest using Bonferroni-corrected analyses. Spearman correlations examined brain-behavior relationships between valence trajectories, prefrontal activation, and clinical measures. RESULTS: MDD patients exhibited consistent negative valence drift under neutral and positive cues, but not under negative cues. fNIRS revealed distinct temporal deviations, characterized by insufficient early prefrontal activation followed by delayed compensatory recruitment, and spatial deviations, with exaggerated reliance on medial prefrontal cortex (mPFC) and attenuated dorsolateral prefrontal cortex (dlPFC) modulation. These spatiotemporal patterns persisted even in the absence of behavioral differences under negative cues. Brain-behavior analyses showed that stronger late negative associations correlated with higher depressive severity and insomnia, whereas increased mPFC and temporal activation reflected compensatory attempts at regulation. CONCLUSIONS: MDD is marked by disrupted spatiotemporal prefrontal signatures, including delayed and prolonged activation and spatial imbalance favoring mPFC over dlPFC. These deviations provide mechanistic insight into depressive cognitive bias and nominate temporal-spatial prefrontal dynamics as ecological biomarkers with potential utility for precision psychiatry. CLINICAL TRIAL NUMBER: Not applicable.
KIParla is a large, modular corpus of spontaneous spoken Italian originally transcribed in ELAN using Jefferson-style conventions. While this representation preserves fine-grained interactional detail and time alignment, it limits interoperability, large-scale querying, and computational reuse. This paper presents the design and implementation of a pseudo-tokenized, verticalized pivot format developed to support validation, maintenance, and infrastructural integration without sacrificing descriptive richness. The proposed format makes explicit the analytical units implicit in Jefferson transcription—transcription units, spans, and tokens—and enforces well-formedness constraints at character, span, and unit levels. Overlap, the most complex relational phenomenon, is resolved through a graph-based algorithm that derives temporal overlap events from alignment data and deterministically matches them to textual spans. Each token is represented as a structured record enriched with lexical, prosodic, interactional, and alignment features, anchored through explicit character offsets. The vertical format functions as a maintained pivot representation from which alternative formats, including ELAN files and UD-compatible treebank representations, can be reproducibly derived. This architecture enables large-scale lemmatization, part-of-speech tagging, syntactic annotation, and cross-layer querying, while supporting version-controlled, DevOps-inspired workflows for sustainable corpus growth. The KIParla pivot format thus reconciles interaction-oriented transcription practices with computational standards and provides a model for reuse-oriented spoken-language data engineering.
The Universal Dependencies (UD) project has grown rapidly through semi-automatic conversion of existing treebanks, but ensuring the quality of converted annotations remains a challenge. Manual verification does not scale, and existing automatic methods cannot distinguish conversion errors from inherent annotation complexity. We present a method that addresses this gap by training two parsers on a token-aligned parallel corpus: one on the original annotation and one on its UD conversion. By requiring correct predictions from the parser trained on the original annotation, our approach isolates errors specifically introduced during conversion while filtering out cases where the construction is simply difficult to parse. We demonstrate the effectiveness of this method on the SynTagRus corpus and its UD counterpart. To enable a direct comparison, we created a fully token-aligned version of the two corpora, resolving differences in tokenization and ellipsis representation. We also proposed a simple method for aligning syntactic relations across the two corpora, addressing the fact that relations involving the same token do not always correspond due to differences in annotation schemes. Our analysis identified several hundred errors in the test set. These comprise six distinct types of conversion errors, three of which persist in the current converter, and ten groups of annotation inconsistencies between the old and new corpus parts. Our method offers a practical, scalable tool for conversion error detection and is applicable to any language pair with aligned original and converted annotations.
Urban mental health burdens are increasing, prompting interest in how nearby green spaces aid emotional restoration. Bamboo-dominant green spaces are widespread in East Asia, but evidence connecting their management and structural features to restorative experiences is limited. This study conducted a controlled photo-exposure experiment in Ya’an, China, to examine how bamboo space typology and structural attributes relate to visual attention, affective responses, and short-term physiological recovery. One hundred and twenty participants viewed 50 photographs representing five bamboo space types (ecological conservation, productive–economic, protective–greenbelt, landscape–recreational, and understory–composite). Each image was linked to a matched field plot, enabling integration of structural indicators with eye tracking, EEG β/α, and repeated ratings of relaxation, pleasure, and preference. Results showed that landscape–recreational spaces received the highest affective ratings, while understory–composite spaces had longer fixations, indicating higher visual processing demands. Vertical stratification and groundcover coverage were robust predictors of affect beyond typology. Eye-movement metrics did not mediate structure–affect associations, and EEG β/α, as an auxiliary and context-dependent indicator under brief photo-based exposure, showed limited sensitivity. These findings offer insights into structural elements that can inform the design and management of bamboo green spaces for improved emotional restoration.
Pronouns indicate significant importance in both pedagogical and communicative contexts, as they shape the way individuals are addressed and understood in social interactions and educational settings. Beyond the traditional pronouns “he” and “she,” the American Psychological Association endorses the scholarly use of the singular pronoun “they,” recognizing its relevance in promoting inclusive language practices. In addition, the popularity of neopronouns continues to rise, providing non-binary individuals with a broader range of linguistic options to express their identities. Despite this growing recognition, there remains a dearth of empirical research that systematically investigates the awareness, knowledge, and preferences regarding pronoun use among non-binary populations. Addressing this gap, the present quantitative inquiry examined the level of awareness, knowledge, and preference of pronouns among non-binary college students at a state university in the Philippines. The study involved 80 participants, including 20 lesbians, 20 gays, 20 bisexual males, and 20 bisexual females, selected through criterion sampling, who responded to a four-part researcher-developed survey questionnaire. The results indicate that, overall, non-binary college students are aware of the categories of pronouns (M=2.48, SD=0.39 for traditional pronouns; M=2.79, SD=0.25 for gender-neutral pronouns; and M=2.64, SD=0.32 for neopronouns) and knowledgeable about them (M=2.49, SD=0.31 for traditional pronouns; M=2.68, SD=0.23 for gender-neutral pronouns; and M=2.47, SD=0.32 for neopronouns). However, despite this awareness and knowledge, participants expressed a preference for using traditional pronouns (“he” and “she”) when being referred to. These findings underscore the persistence of traditional linguistic norms in educational settings and highlight the potential influence of formal language instruction on pronoun preference. Empirically, this study contributes to the limited body of research on non-binary pronoun use speficically in the Philippines, providing a foundational dataset that can inform inclusive language policies, pedagogical strategies, and future sociolinguistic investigations. Its significance lies not only in documenting patterns of pronoun awareness and preference but also in offering evidence-based insights for educators, policymakers, and advocates seeking to foster more inclusive and affirming learning environments.
This paper quantifies how parameter-efficient fine-tuning narrows the performance gap between compact transformers and massive Large Language Models (LLMs) on sentiment analysis while reducing memory, energy, and financial budgets. Experiments use the publicly released BERT-base-uncased checkpoint (12 layers, 110 M parameters), the Stanford Sentiment Treebank-2 benchmark (67,349 movie phrases with binary polarity labels), PyTorch and Hugging Face Transformers 4.41, Weights & Biases for experiment tracking, and an NVIDIA A100 GPU configured for mixed-precision computation. Low-Rank Adaptation (LoRA) adapters (rank 16) are injected into the query/key/value projections, so only 0.54 % of weights are updated. The model was trained for five epochs with AdamW, cosine scheduling, batch size 64, and optional 4-bit post-training quantization. Accuracy and macro-averaged F1 are logged every 25 steps. The LoRA-tuned model achieves 91% accuracy and 0.91 F1 on the SST-2 test set, an improvement of 40 percentage points over the off-the-shelf checkpoint and comparable to GPT-4o-mini (93%), while using fewer than 1⁄1500 of its parameters. Validation loss plateaued without overfitting, and 4-bit quantization compressed the model to 27 MB with <0.5 point accuracy loss. Energy profiling shows a 73% reduction in GPU consumption compared with full-parameter fine-tuning. Purpose-built adapters and quantization can unlock high-quality NLP on edge devices. Future work should extend the protocol to multilingual corpora, streaming inference, federated learning, and on-device continual adaptation to preserve accuracy under concept drift and safeguard user privacy.
Endangered Hausa dialects spoken in Jigawa State and selected communities in Kano, Katsina, and Yobe States represent an irreplaceable repository of cultural and linguistic heritage, yet these varieties are diminishing at an accelerating rate under the pressure of urbanisation, formal education in Standard Hausa, and mass media. This paper reports the principal findings of a linguistic documentation project that conducted the first systematic digital preservation of nine endangered Hausa dialect varieties across the four-state study region. Using a linguistic documentation research design, the project engaged 120 native speaker participants distributed across three age cohorts and nine dialect communities. Data were collected through 84 hours of audio and video recording, semi-structured interviews, participant observation, GPS-based dialect mapping, and systematic metadata documentation. Linguistic analysis produced annotated transcriptions of all recordings, a lexical database of 1,850 entries including 342 dialect-only lexical items, a phonological description of 24 unique tone contrasts and 18 distinct vowel patterns, 216 documented proverbs, and 68 dialect-specific kinship terms. A digital archive of 120 gigabytes was created following UNESCO and Endangered Language Documentation Programme standards. Language vitality assessment using the UNESCO framework revealed that six of the nine documented varieties score at or below the 3.0 endangered threshold, with Guri dialect critically endangered at 2.1. Mean young adult dialect use rates average 36 percent of elderly speaker rates, confirming severe intergenerational transmission failure. The study proposes a three-phase Digital Preservation Framework and offers evidence-based recommendations for language policy, curriculum design, and institutional capacity building in Nigerian Hausa dialect documentation.
Currently, East Asia, especially South Korea, is facing social problems such as population decline and regional inequality. To address these challenges, tourism has been leveraged as a means of economic revitalization, especially in fishing villages that are economically disadvantaged. This study examined the authentic food experiences of tourists who visited fishing villages. Tourists’ food experiences, dimensions of destination food image, and destination loyalty were assessed. In September 2024, 448 responses from Korean tourists were collected and analyzed using confirmatory factor analysis and structural equation modeling to test 15 hypotheses. Local food authenticity showed a significant effect on all dimensions of destination food image. Of the five dimensions of destination food image in this study, food taste, health and hygiene, and unique cultural experiences significantly influenced destination loyalty. In addition, geographic differences moderated the relationship between local marine food authenticity and the perceived food image of the destination. Tourists in the southern coastal regions reported the highest destination food image ratings, driven by authentic local cuisine, while those in the western regions reported the lowest. This study offers both practical and theoretical implications related to sustainable coastal tourism.
Neural DNA (NDNA) uses a compact learned genome to grow network topology through type-based compatibility rules. In prior work (Sudarshan, 2026), we showed that genomes of 226 to 374 parameters could control connectivity in networks with up to 2.2 million connections. Here we test whether this approach scales to real language models. We apply NDNA to GPT-2 Small (124M parameters), where a genome of 354 parameters controls 35.4 million connections in the attention output projections and feed-forward first layers across all 12 transformer layers, a compression ratio of 99,970:1. The genome and model weights are co-trained from scratch on OpenWebText. The genome discovers a striking stratification: layers 5 through 12 converge to 100% connectivity while layers 1 through 4 are progressively pruned, with layer 4 retaining only 7.7% of connections. Despite permanently disabling one-third of all masked connections, the genome-wired model beats GPT-2's published numbers on WikiText-103 perplexity (36.0 vs 37.5), Penn Treebank perplexity (59.4 vs 65.9), and LAMBADA perplexity (22.2 vs 35.1), while reaching 92% of GPT-2 on HellaSwag and 94% on Children's Book Test. Training reveals interpretable dynamics: the genome over-activates early (83% density at iteration 200), prunes aggressively (layers 1 through 4 go to 0% by iteration 800), then partially resurrects layer 1 (from 0% to 98% hard density by iteration 26,000). These results demonstrate that NDNA's developmental framework scales from toy tasks to production-scale language models, achieving 12x higher compression than any prior genome experiment while producing competitive benchmark performance.
Species-level tree identification is a fundamental task in forest monitoring, biodiversity assessment, and climate-smart ecosystem modeling. Close-range laser scanning technologies have become indispensable tools for forest mapping because they provide high-resolution, three-dimensional structural data at the individual-tree level. However, species-level identification remains a major challenge in global environmental monitoring and AI-driven ecological assessment owing to high species diversity, structural plasticity, and variability across sensing platforms. Here, we propose the cognition-inspired multimodal attention fusion network (CI-MAFusion), a dual-branch deep learning framework that integrates point cloud data with multi-view imagery. Guided by expert dendrological reasoning and cognitive neuroscience principles, CI-MAFusion incorporates a structural branch based on an improved graph attention-based point network for encoding 3D morphological patterns and a visual branch that processes standardized multi-view projections to extract textural features. A cross-gate attention mechanism adaptively fuses structural and visual features. Each branch uses an enhanced convolutional block attention module to highlight salient features, analogous to selective attention in the human visual system. We tested CI-MAFusion using Global LiDAR TreeBank, which contains 12,057 trees from 36 species across four continents and six Köppen climate zones. The model achieved 87.50% overall accuracy at the genus level and 86.12% at the species level, outperforming unimodal and existing fusion approaches by up to 8.1%. Additionally, it further achieved > 90% overall accuracy across regions and > 80% across climate zones, with attention visualizations highlighting biologically diagnostic features such as crown contours, bark textures, and branch junctions. This cognitively inspired architecture improves generalization and advances AI-based systems toward robust recognition of biological structures in complex environments.
I. Core Concept: Language as a "Spiritual Commons" (精神公地)• The Shared Semantic Space: Language is not private property; it is a shared environment. Entering a community means entering its "Spiritual Commons."• Commons Literacy (公地素養): The ability and willingness of an individual to maintain the functionality of this shared space through effective communication, respect for linguistic norms, and the rejection of "communicative bullying."II. Research Branch A: The Ethics of Global Communication• The "Natural Selection" of Lingua Franca: Analyzing why English became the global commons—not through coercion, but through the "Civilizational Excellence" and economic utility of the English-speaking world.• The Responsibility of Public Figures: Researching the "Moral Decay" inherent in public figures who use "Shameful Refusal to Communicate" (e.g., using a language the audience cannot understand to exert power or exclusion).• The Credit of the National Commons: How individual behavior in international spaces mortgages or enhances the reputation (credit) of their entire ethnic/linguistic group.III. Research Branch B: Cognitive Diversity and "Thinking Inertia"• Breaking Mental Sets: Investigating how learning a second language acts as a "Symmetry Breaking" event in the brain, allowing individuals to escape the "Thinking Inertia" (思維定勢) of their mother tongue.• The Creativity of "Weakness": A study on how the struggle of learning a new language leads to professional-level creative output (fables, lyrics) by forcing a simplified, more profound engagement with universal concepts.• AI vs. Human Literacy: Why AI translation can solve the "Vexation" of data transfer but cannot replace the "Commons Literacy" required for high-level creative and moral leadership.IV. Research Branch C: Comparative Civilizational Performance• Linguistic Proficiency and Economic Development: Analyzing the correlation between a nation's "Commons Literacy" (openness to the global lingua franca) and its economic/scientific success.• The "Japan-Korea Paradox": Studying how highly developed nations maintain a strong "Private Mother Tongue" while effectively interfacing with the "Global Commons."V. Proposed Moral Maxims for the LCL Field1. Transparency: Public speech in a shared space must be accessible to that space's participants.2. Symmetry: One must tolerate the same level of linguistic/cultural critique they dish out.3. Elevation: To promote one’s mother tongue, one must elevate their civilization's quality, not shame others for their ignorance.
Virtual agents powered by large language models are increasingly deployed in digital mental health services, yet the influence of avatar appearance on users' emotional, cognitive, and physiological responses remains insufficiently understood. This study was conducted between March and April 2024 and examined how three avatar designs-animal-like, human-like, and object-like-shape affective experience, user evaluation, autonomic activity, and attentional allocation during virtual doctor interactions. Forty-two participants completed a within-subjects experiment involving self-reported affect ratings, multidimensional user-experience assessments, heart rate variability (HRV) measures, and eye-tracking indicators. The avatar type did not yield statistically significant differences in changes in positive or negative affect across conditions. However, physiological data revealed clear divergences. The animal-like avatar elicited the strongest parasympathetic activation, reflected by significant increases in the root mean square of successive differences (RMSSD) and high-frequency (HF) power, whereas the object-like avatar produced a sympathetic-dominant response. Across six user-experience dimensions, the animal-like avatar consistently received the highest evaluations. Eye-tracking results showed faster first fixation and a longer face-directed fixation duration for the animal-like avatar, indicating stronger social attention. The human-like avatar demonstrated slightly delayed initial fixation, consistent with subtle yet nonsignificant uncanny-valley tendencies. These findings underscore the critical role of avatar visual design in shaping emotional safety, engagement, and social processing in virtual mental-health interactions.
Our brain maps the space immediately surrounding the body, the peripersonal space (PPS), to sharpen sensory-motor coordination whenever an object enters it. Within PPS, past research demonstrated how several factors influence motor readiness: from stimulus characteristics, such as body-object distance and stimulus semantics, to personality traits like anxiety. However, most paradigms infer PPS boundaries using tactile and visual stimuli, with auditory cues often playing an ancillary role, rather than directly measuring response modulation to stimuli entering the PPS. Here, we measured anticipatory postural adjustments, as direct physiological indices of motor planning, to examine whether semantic content and individual suggestibility modulate responses to looming sounds stopping within PPS. Thirty-three adults heard affective semantic sounds (positive - applause, negative - dentist drill) or neutral non-semantic (pink noise) sounds halting at five within-PPS distances while we recorded muscle activation timing, distance estimates, affective ratings, and sensory suggestibility. Motor responses were faster as sounds stopped nearer the body but systematically delayed and less precise for semantic sounds compared to non-semantic sounds, with higher suggestibility predicting longer and more variable latencies, particularly for non-semantic sounds. These findings demonstrate that semantic evaluation imposes measurable processing costs that systematically delay motor preparation and increase perceived distance. Individual suggestibility produces freeze-like response patterns that amplify motor uncertainty, specifically when semantic context is absent, revealing distinct cognitive and trait-based mechanisms that jointly regulate defensive behaviour within PPS.
As of 2025, more than 5.2 billion people in the world use social media, which is about 63.9% of the world’s population, with a growth rate of 4.1% over the past 12 months. The most popular platforms are Facebook, Instagram, TikTok, Twitter, and WhatsApp. The average time spent on social media is about 2 hours and 26 minutes per day, and the average user has access to seven different platforms. Speech on social media is based on the same language norms (lexical, spelling, grammar, syntax) as live speech. The purpose of the article is to provide an extended analysis of lexical innovations in the language space under the influence of social media and digital communication tools. The object of this study is the modern vocabulary of several languages used within social platforms (Twitter, TikTok, Facebook, Instagram). Particular attention is paid to modern English, which is the most widespread language in communication practice – approximately 1.5 billion people speak English, and 52% of the world’s most popular websites contain English-language content. The article uses scientific and linguistic analysis to investigate the peculiarities of the transformative impact of social media communication on language at all structural and functional levels: lexical, phonetic, grammatical, syntactic and graphic. The article analyzes the characteristic lexical changes by groups – memes, neologisms, abbreviations and acronyms, phraseological units, hashtags. The functions of different categories of lexical innovations of social networks are determined, in particular: hashtags form the basis for unimpeded communication in an intercultural context, neologisms are means of constructing the identity of certain social groups, memes have the functionality of entertainment and information, disseminating precedent information in the format of textual and graphic expression. The negative aspects of the impact of social networks on language are identified: excessive simplification of language and loss of its individual nuances, the emergence of inaccuracies and grammatical errors due to the spontaneous nature of communication on social networks, as well as potential negative consequences for mental health. The study proves that the modern space of innovative language practices reflects new concepts of social media communication culture, interactive upgrading and visualization, which transforms religious and cultural aspects and promotes sustainable language changes.