Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backchannels - in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.
The paper proposes annotation guidelines for syntactic dependencies that span across speaker turns - including collaborative coconstructions proper, wh-question answers, and backchannels - in spoken language treebanks within the Universal Dependencies framework. Two representations are proposed: a speaker-based representation following the segmentation into speech turns, and a dependency-based representation with dependencies across speech turns. New propositions are also put forward to distinguish between reformulations and repairs, and to promote elements in unfinished phrases.
What constrains working memory capacity? Classic theories place visual working memory close to perceptual systems, with fixed limits. Yet, emerging evidence shows that visual working memory capacity is increased for real-world objects compared to simple or abstract stimuli. The present study demonstrates that this memory advantage arises from semantic understanding of real-world objects – contrary to classic perceptual accounts of this cognitive system. Using counterfeit objects generated by generative adversarial networks that match real objects in terms of object form and visual similarity, we show that improvements in behavioral performance and increases in neural delay activity emerge solely for semantically meaningful, real objects. Correlation analyses indicate that subjective familiarity ratings predict memory for real objects, whereas stimulus colourfulness predicts memory for artificial objects, suggesting distinct mechanisms support memory for different stimulus types. Thus, conceptual knowledge exerts strong effects on visual working memory, significantly extending current theories that emphasize low-level perceptual features.
This article focuses on improving the language skills of future professionals in Ukrainian language classes tailored to their specific fields. It has been established that Ukrainian language instruction tailored to specific professional fields is a key tool for enhancing language skills, facilitating the development of professional competence, adherence to linguistic norms, mastery of a business style, and effective communication. It is emphasised that the teaching profession places very high demands on specialists: a teacher must be a unique, vibrant personality, a repository of diverse and profound knowledge, and possess a sufficiently high level of cultural refinement. Therefore, future specialists must have a perfect command of Ukrainian linguistic etiquette, as the teacher’s personal example will help students develop the skills of linguistic communication. It has been established that the main areas for improving communication skills include: adherence to linguistic norms; mastery of specialist terminology; development of professional communication; observance of communication etiquette; and the development of logical and clear expression. The following methods are used in practical sessions: business games and role-play scenarios; text analysis; discussions and debates; drafting and editing documents; and the use of multimedia resources. The Ukrainian language for professional purposes is an important component of professional training, ensuring the high level of intellectual and linguistic proficiency required for a successful career. To successfully deliver the curriculum, a primary school teacher must not only have a thorough understanding of language teaching methodology, but also serve as a model of standard speech. This entails impeccable literary pronunciation, the correct use of vocabulary and grammatical structures, as well as skilful control of intonation during spoken communication and reading. In other words, teachers must pay equal attention to both the content and the form of their speech. Unlike professionals in other fields, teachers use language as a public tool that serves as a benchmark for their pupils.
The authors see the purpose of the study as a comparative analysis of political discourse on the example of D. Trump’s speeches during his election campaigns in 2016 and 2024. The scientific novelty lies in the confirmation of the concept of using language as a weapon, which acts in Trump’s speeches as a tool to manipulate and control people through different discursive means and in different periods of time. Comparing the changes that discourse has undergone over time allowed the authors to analyse the priorities that reflect the general political tendency for positive emotional encouragement rather than threats and aggression. The relevance of this article is determined by the shift that has occurred in political discourse, which requires a rethinking of how the political actors select linguistic norms and how this selection will affect the formation of modern political language.
Background: Dense cross-layer connectivity can shorten gradient paths and promote feature reuse, potentially improving optimization under fixed training budgets. Objective: We test whether concatenation-based dense historical connectivity improves decoder-only autoregressive language modeling under controlled comparison protocols. Methods: We compare a standard Transformer decoder and a dense decoder on Penn Treebank and WikiText-2 under two fairness regimes: (i) a same training recipe setting with a fixed baseline and a bounded dense architectural search, and (ii) a same parameter budget setting where the dense model is resized to not exceed the baseline parameter count. Results: Dense connectivity does not consistently reduce test perplexity; on WikiText-2, the baseline remains better in both regimes, while gains on Penn Treebank are small and regime-dependent. Ablations within the dense family show that depth and feed-forward capacity are the most reliable drivers of perplexity improvements. Conclusions: Probes and attention diagnostics do not reveal a clear advantage for dense connectivity in our limited probe set, while Zipf–RQA analysis of long-form generations reveals systematic structural differences between baseline and dense outputs. Specifically, Zipf–RQA is used here as a descriptive structural probe rather than a performance metric.
Cet article présente la dernière version du treebank Rhapsodie, un corpus de français parlé multi-genres annoté en syntaxe et prosodie. Les deux principales innovations sont une annotation morphosyntaxique réellement basée sur la version orale du corpus (ce qui est prononcé) et non sur sa transcription orthographique et une intégration de l’ensemble des niveaux d’annotations, syntaxe, prosodie et métadonnées, dans une même structure, permettant ainsi des requêtes croisées.
Abstract This study examines clitic placement and the use of contracted forms in 13th-century medieval Spanish, focusing on the works of Gonzalo de Berceo. The results indicate that, although proclisis was the predominant norm, unexpected enclitic usages appear, especially in initial position and after metrical caesuras. These occurrences suggest that Spanish at the time was in a transitional stage regarding the fixation of clitic position. Metrical patterns and parallel syntactic structures also influenced clitic distribution, revealing an interaction between linguistic norms and stylistic constraints. Furthermore, the analysis identifies contracted forms affecting not only third-person pronouns but also first- and second-person forms, although their usage was not fully standardized. These findings support the view that the evolution of clitics in Spanish was a gradual process, shaped by phonological, syntactic, and discourse-related factors, with poetic composition offering a space in which grammatical rules and artistic expression coexisted and interacted.
This chapter examines the notion of racialised languaging, which emphasises that languaging practices are never assessed independently of the bodies, identities, and social positions of their speakers. It demonstrates how language is evaluated not only in terms of what is said but also through the racialised perceptions of who is speaking and how society chooses to listen. The chapter argues that accents, dialects, and speech patterns associated with racialised communities are often constructed as inferior, humorous, deficient, or even criminal, while similar features in white speakers are normalized or excused. By centring languaging as a site of racial meaning-making, the chapter exposes the ways in which communication is entangled with race, racism, and embodied identities. Racialised languaging is further situated within the broader colonial matrix of power, where Western linguistic norms and white racial identities are privileged over non-Western languages and non-White speakers.
This study focused on structural and semantic changes in Kazakh borrowed vocabulary adaption. The study examined how imported concepts were absorbed into Kazakh, altered by the national linguistic system, and contributed to current terminology. Various linguistic methodologies were used, including historical and contemporary text analysis, structural and comparative term analysis, and hybrid word classification. The study covered both present and historical English, allowing vocabulary changes to be tracked. The findings showed that complicated historical and cultural processes in active intercultural exchanges caused Kazakh terminology hybridisation. The Greco-Latin, Persian, and Arabic languages enriched Kazakh lexicon with science, religion, culture, and daily life concepts. Phonetic, morphological, and visual changes were made to borrowed terminology to make them useful and conform to Kazakh linguistic norms. The study showed that hybrid terms are crucial to borrowing integration.
This study investigates the integration of dialectal features into official Czech toponyms, with a particular focus on street names and non settlement names. It explores how dialectal elements persist in official naming practices despite standardization efforts, especially in regions with strong dialect traditions. The authors analyze the linguistic and administrative processes behind toponym standardization and highlight the discrepancies between different mapping platforms – namely the state-run Geoprohlížeč and the commercial mapy.com. While street names are regulated and recorded in the national database (RÚIAN), non-settlement names lack centralized oversight, resulting in greater variability. The paper identifies specific phonological and morphological dialectal features that appear in official names, often due to the direct adoption of local spoken forms. The authors argue for a balanced approach to standardization that respects both linguistic norms and regional identity, emphasizing the cultural and communicative significance of toponyms in public space.
This chapter explores how artificial intelligence (AI) tools mediate the identity development, emotional labor, and academic adaptation of international students in U.S. higher education. Framed through the lenses of intersectionality, resilience, and self-authorship, the chapter draws on duoethnography to examine how AI is used not merely as a technical aid but as a scaffold for rewriting the self in unfamiliar academic terrain. While AI offers immediate access to academic conventions, its reliance on dominant linguistic norms often flattens cultural expression and obscures opportunities for deeper growth. Through personal narrative, peer reflection, and theoretical analysis, this chapter interrogates what is gained and what is lost when AI supplements or replaces human-centered support systems. It argues that international student engagement with AI reveals a broader story about survival, belonging, and identity negotiation in an increasingly technologized and globalized university landscape.
This paper presents a direct framework for sequence models with hidden states on closed subgroups of U(d). We use a minimal axiomatic setup and derive recurrent and transformer templates from a shared skeleton in which subgroup choice acts as a drop-in replacement for state space, tangent projection, and update map. We then specialize to O(d) and evaluate orthogonal-state RNN and transformer models on Tiny Shakespeare and Penn Treebank under parameter-matched settings. We also report a general linear-mixing extension in tangent space, which applies across subgroup choices and improves finite-budget performance in the current O(d) experiments.
This article explores the pivotal role of information technology in the study of endangered languages, examining its various applications, benefits, and challenges. From digital archives and linguistic databases to computational analysis and online collaboration platforms, IT offers a plethora of tools and resources that facilitate language documentation, analysis, and revitalization efforts. Moreover, IT enables greater accessibility and dissemination of linguistic data, fostering collaboration among researchers, communities, and stakeholders across geographical and cultural boundaries.
This study examined how pre-listening information influences music appreciation among 107 Japanese junior and senior high school students. Two songs were used: The Italian Sogno, where musical tone aligns with lyrics, and the German Im wunderschönen Monat Mai, where they do not align. Participants were assigned to three groups differing in the amount of prior information: none (“No Information Group”), brief lyric explanations (“Lyrics Explanation Group”), and detailed explanations including lyrics, background, and acoustic features (“Lyrics and Background Explanation Group”). When lyrics and tone were incongruent, the No Information Group’s emotional valence ratings aligned more with the tone than did those with prior information; this effect was absent in the congruent condition. Open-ended responses showed the No Information Group focused on surface features like the languages of lyrics rather than thematic content. These findings highlight the educational value of emphasizing lyric understanding in Japanese music education.
Abstract Research on how non-natives process and learn binomials ( black and white ) is limited. The present study addresses this gap using online (eye-tracking) and offline (familiarity rating) tasks. Sixty non-native speakers of English (L1 = Arabic) read six stories seeded with 21 novel binomials in three conditions: one exposure, six exposures, and no exposure (i.e., only in post-test) in a counter-balanced design. Each item was also presented in the reversed order ( white and black ). The non-natives read the stories as their eye movements were monitored and answered comprehension questions. In addition to the novel binomials, 12 existing binomials (congruent with Arabic) were included in the passages as a baseline for comparison. After completing the reading task, the participants completed an offline rating task as a measure of declarative knowledge of the binomial configuration (i.e., word order). All items were rated twice, once in the forward direction and once in the reversed direction. Online results showed that non-natives were not sensitive to the configuration of existing binomials, and there was limited evidence of any sensitivity to novel binomials. Offline, non-natives showed sensitivity to the configuration restrictions of existing binomials but not novel ones.
In Spanish linguistics today, it is widely recognized that Spanish corresponds to the image of a pluricentric language. Different normative centers coexist, which are perceived as such by the speakers. However, the degree of recognition, status, and prestige of these linguistic norms varies greatly. This article examines the extent to which the pluricentrism of Spanish is represented in textbooks for Spanish as a foreign language. As a case study, three textbooks – Puente Nuevo, ¡Adelante!, and A_tope.com – used in high schools in Basel, Switzerland, are analyzed both quantitatively and qualitatively in terms of their pluricentric character. The study examines the textbooks as a whole (thematic priorities, topics of units, maps) and selected linguistic phenomena, namely: the forms of address (morphosyntax), seseo (pronunciation), and the use of regional vocabulary. At various levels, it can be demonstrated that the textbooks are based on a Eurocentric view of language, which gives Castilian Spanish a superior role to that of other language norms, without this being explicitly stated. The article concludes with some practical recommendations for a more pluricentric approach to teaching Spanish.
We describe THIVLVC, a two-stage system for the EvaLatin 2026 Dependency Parsing task. Given a Latin sentence, we retrieve structurally similar entries from the CIRCSE treebank using sentence length and POS n-gram similarity, then prompt a large language model to refine the baseline parse from UDPipe using the retrieved examples and UD annotation guidelines. We submit two configurations: one without retrieval and one with retrieval (RAG). On poetry (Seneca), THIVLVC improves CLAS by +17 points over the UDPipe baseline; on prose (Thomas Aquinas), the gain is +1.5 CLAS. A double-blind error analysis of 300 divergences between our system and the gold standard reveals that, among unanimous annotator decisions, 53.3% favour THIVLVC, showing annotation inconsistencies both within and across treebanks.
Reviewer assignment is increasingly critical yet challenging in the LLM era, where rapid topic shifts render many pre-2023 benchmarks outdated and where proxy signals poorly reflect true reviewer familiarity. We address this evaluation bottleneck by introducing LR-bench, a high-fidelity, up-to-date benchmark curated from 2024-2025 AI/NLP manuscripts with five-level self-assessed familiarity ratings collected via a large-scale email survey, yielding 1055 expert-annotated paper-reviewer-score annotations. We further propose RATE, a reviewer-centric ranking framework that distills each reviewer's recent publications into compact keyword-based profiles and fine-tunes an embedding model with weak preference supervision constructed from heuristic retrieval signals, enabling matching each manuscript against a reviewer profile directly. Across LR-bench and the CMU gold-standard dataset, our approach consistently achieves state-of-the-art performance, outperforming strong embedding baselines by a clear margin. We release LR-bench at https://huggingface.co/datasets/Gnociew/LR-bench, and a GitHub repository at https://github.com/Gnociew/RATE-Reviewer-Assign.
The current paper investigates the translation of proverbs in Abai Kunanbaev’s “Words of Edification” into Russian and English languages, focusing on the strategies used to convey semantic accuracy, cultural meaning and pragmatic intent. In this study, we have used Molina and Albir’s (2002) translational classification to analyze the selected proverbs in Kazakh langauge. As a result, we have identified some applied translation techniques in rendering the source proverbs. These findings indicate that word-for-word translation is the most frequently used strategy in the indirect version. As to the Russian translation, more often employed strategies are established equivalent, modulation and adaptation which align with Russian cultural and linguistic norms. Under certain circumstances, the indirect translation shows semantic shifts, metaphorical loss of meaning because of mediating language rather than the source culture. Yet, applying modulation and borrowing strategies enable to maintain several cultural components and metaphorical traits.
The rise of Artificial Intelligence (AI)-based tools is transforming language education, offering adaptive innovations for both language learners and instructors. However, concerns remain about their ability to represent linguistic diversity. Shaped by the data they process and the priorities of their creators, AI systems risk reinforcing dominant linguistic norms. This study explores these issues using Austrian Standard German (ASG), a distinct variety of German, as a case study. By analysing the behaviour of four AI tools—two models of a chatbot, a text-feedback system, and a grammar-correction tool—we assessed whether they recognised ASG as a legitimate standard or altered it to align with German Standard German (GSG). The evaluation, informed by structured testing, revealed significant shortcomings in how these systems handle linguistic variation. Our findings underscore the risks of erasing linguistic particularities and emphasise the need for AI tools to serve as a means of fostering linguistic identity—particularly in educational contexts.
This chapter explores the uneven geographies of access, participation, and belonging in global higher education, focusing on how language shapes the lived experiences of international students and faculty. Drawing on classroom-based narratives from India and Oman, it examines how linguistic norms, institutional expectations, and cultural assumptions determine inclusion and exclusion in academic spaces. While student mobility is often framed as a success of globalization, the chapter argues that it remains embedded in hierarchies of language, identity, and geography. Multilingual learners from rural or non-elite backgrounds, and faculty from “non-native” English-speaking contexts, often face marginalization due to misalignment with dominant academic norms. Using Bourdieu's linguistic capital, postcolonial critiques, and critical internationalization studies, the chapter calls for moving beyond tokenistic diversity. It advocates for multilingual pedagogies, translanguaging, and ethical student mobility to build a more equitable and culturally responsive global education landscape.
This article analyzes the systemic crisis of Arabic culture from the 13th to 18th centuries and the pivotal role of Arab Christians in initiating the Arab Renaissance. The author explores the causes of stagnation in Muslim society, primarily the institutional dominance of taqlīd (imitation) and the widening gap between sacralized linguistic norms (fuṣḥa) and living speech. Central to the study is the scholarly contribution of the Maronite scholar Ibn Farhat, whose work bridged Western rationalism and Eastern tradition. Special focus is placed on his treatise «Baḥṯ аl-maṭālib wа ḥаṯṯ аl-ṭālib», which simplified Arabic pedagogy and integrated the language into the daily and liturgical practices of Christian communities. The paper emphasizes that the transition from Karshuni script to classical Arabic, alongside Ibn Farhat’s reforms, provided the ideological foundation for overcoming cultural isolation. It concludes that Lebanon’s Christian intellectuals, educated through European models like the Pontifical Maronite College, were the primary catalysts for modernizing Arabic philology and precursors to the 19th-century Enlightenment.
Multilingualism is defined as a mode of communication in contemporary world. The multilingualism teaches us the important values to understand the context. This study analyzes dual point of view about the multilingualism: the foreign languages that appear in it, i.e. explicit multilingualism and the universal aspect or hidden languages that are indirectly described, i.e. implicit multilingualism. Thismay comprise linguistic norms, reader and text interaction, among others. The aim of this study is to highlight the impacts of elements of multilingualism used in Amélie Nothomb’s novels. It focuses essentially on the works of the contemporary Francophone writer, notably, Amélie Nothomb. She articulates the enriching elements of multilingualism in French and Japanese languages through herwritings. Her breakthrough works mainly articulate the diversity of multilingualism and also the essential meaning of understanding the different elements or expressions related to the French and Japanese language through the richness of culture from a geographical point of view and also the other elements. These elements are articulated about expressions which show the impact ofmultilingualism in her writings that refer either to French, Japanese, or other languages.
ABSTRACT This study investigates the perceptions of Americanisms among three generations of Nigerians. While prior research has provided quantitative evidence for American influence in contemporary Nigerian English, the role of language beliefs and ideologies in mediating such changes remains underexplored. Developing a sociolinguistic perspective of mobile linguistic resources, this study construes an individual's linguistic repertoire as an identity‐construction resource, agentively mobilised across geographical, social and digital spaces. Interview data indicate that younger speakers orient towards multiple linguistic norms, while older speakers remain critical of Americanisms and favour British norms. Reading task results further indicate that American realisations are most frequent among younger speakers. The study demonstrates that multinormativity extends beyond linguistic production to speakers’ evaluative orientations and perceived repertoires. This finding advances the sociolinguistics of mobility and World Englishes research by showing that shifting language ideologies – rather than usage patterns alone – constitute a key mechanism driving linguistic change in postcolonial varieties.
Child-directed fingerspelling is an approach used by Deaf parents for communication, language, and literacy development. This study reports on findings from a qualitative intrinsic case study aimed at understanding how Deaf parents use fingerspelling with their young children. The research questions were: (1) What are the cultural beliefs of Deaf parents regarding fingerspelling with young children? (2) What are their patterns of use of child-directed fingerspelling in natural settings? Twenty-one Deaf families with 27 deaf children ages 5 years and under were interviewed via recorded Zoom meetings conducted in American Sign Language. Data were analyzed using grounded theory to develop a new theoretical contribution with the core category: Deaf families socialize their children into Deaf visual-linguistic norms through fingerspelling. This new theoretical insight aligns with Holcomb's Deaf epistemological framework (2010) and Ochs and Schieffelin's (2008, 2011) language socialization theory. Limitations and recommendations for future research are also included.
AI-mediated communication refers to communicative processes in which artificial intelligence systems actively generate, interpret, modify, or facilitate language. With the rapid advancement of language technologies such as large language models, conversational agents, speech recognition systems and machine translation tools, AI has evolved from a supportive tool into an active mediator of human communication. This paper conceptualizes AI-mediated communication as a form of language technology that reshapes traditional models of interaction, meaning-making and authorship. The study examines the technological foundations of AI-driven language processing and their broader socio-cultural, pedagogical and ethical implications. It explores how AI mediation influences linguistic norms, accessibility and power relations, while also raising critical concerns related to bias, surveillance and linguistic homogenization. By situating AI-mediated communication within contemporary digital culture, the paper argues that understanding AI as a language technology is essential for critically engaging with evolving communicative practices and for developing responsible, inclusive and ethically grounded AI-driven communication systems.
The English language, traditionally viewed as monolithic and monocentric, now evolves into diverse varieties embedded in unique communities of practice. The expansion of online communication has further informed this evolution, giving rise to virtual spaces where members negotiate shared meanings, identities, and linguistic norms. This paper explores language use within an online networking business community of practice in the Philippines. Drawing on theories of Communities of Practice (COP) and Cultural Models, it examines how English is positioned in relation to the members’ cultural models and how community-specific jargons function as markers of identities. Through in-depth interviews and analysis of actual online conversations and social media posts, the study reveals that group members use English and local languages in dynamic linguistic strategies, such as code-switching and translingual practices, to construct and perform their identities within the group. The findings further reveal the role of English as a tool for empowerment and a marker of identity in digital spaces, hence emphasizing the need for a more inclusive understanding of language use in contemporary online communities.
This article examines the role and significance of language corpora and linguistic databases in linguistic expertise. It discusses the possibilities of conducting objective semantic, pragmatic, and stylistic analyses of texts through corpus-based methods. The study also highlights the contribution of linguistic databases and artificial intelligence technologies to improving the accuracy, reliability, and efficiency of expert conclusions. Furthermore, the relevance of developing specialized corpora and databases for forensic linguistics in Uzbekistan is substantiated.
We revisit punctuation-aware tree binarization for constituency parsing and ask whether dependency-induced headedness improves binary parser supervision. Although learned heads substantially outperform rule-based heads in intrinsic head prediction, they do not yield consistent parsing gains after debinarization. In particular, punctuation-conditioned evaluation shows that learned headedness underperforms rule-based binarization in macro-average punctuation-sensitive $F_1$, despite a small overall gain on CTB. Similar instability appears under cross-treebank transfer. These results suggest that \ycc{linguistically grounded} headedness is not necessarily parser-optimal when used as a binarization control signal. The paper presents a negative result: better head prediction does not imply better punctuation-sensitive constituency parsing.
This study investigates the linguistic and communicative functions of abbreviations in English and Karakalpak advertising discourse, focusing on how these compressed forms contribute to message efficiency, stylistic expression, and cultural positioning. Although abbreviations are widely used across global advertising, their structural patterns and pragmatic roles vary according to linguistic norms and audience expectations. Therefore, the research employs a mixed qualitative methodology integrating structural analysis, discourse interpretation, and comparative linguistics. The results demonstrate that English advertising makes extensive and creative use of acronyms, initialisms, blends, and hybrid forms to construct modern, technologically oriented, and globally recognizable brand identities. In contrast, Karakalpak advertising relies more on functional initialisms and borrowed English abbreviations, reflecting both local communicative preferences and growing global influence. The discussion interprets these findings within broader socio-cultural and economic contexts, revealing that abbreviation usage serves as a marker of globalization, cultural continuity, and linguistic innovation. Ultimately, the study contributes to a deeper understanding of how abbreviated forms shape contemporary advertising communication in multilingual environments.
Pre-trained language models (PLMs) achieve high accuracy on standard benchmarks for sentiment analysis. However, this performance can hide systematic weaknesses in determining the sentiment of negated sentences, for example when the phrase “not good” is still classified as positive. In this study, we use sentiment classification of English movie reviews in the Stanford Sentiment Treebank 2 (SST-2) as a case study to specifically examine and improve how BERT handles negated sentences. We perform a brief additional fine-tuning of the existing BERT model on a small, automatically constructed set of lexicon-based counterfactual examples that target simple lexical negation. Experimental results on carefully paired original-negated sentences show that this procedure substantially reduces prediction errors on negated inputs while leaving overall performance on SST-2 almost unchanged.
According to an act-based conception of propositions, propositions are types of cognitive or linguistic acts. Such accounts are advertised as having major metaphysical and epistemological advantages over traditional platonic accounts. However, existing versions of such accounts appeal to platonic properties and relations in order to account for the contents expressed by predicates, reintroducing many of the problems they aim to solve. Characterizing both that a is F and that it's F as different types of ``assertibles'' (the former can be asserted full-stop and the latter can be asserted of things), the issue can be seen as a limitation of existing act-based approaches: they apply only to a restricted class of assertibles. In this paper, I show how adopting a normative functionalist approach to linguistic meaning enables one to generalize the act-based approach to all assertibles such that no appeal to extrinsic properties and relations is needed. I show, further, how this radicalized act-based account provides the resources for a satisfactory account of our knowledge of objective states of affairs, properties, and relations (which I dub ``instantiables'') in terms of our mastery of linguistic norms.
Social media has significantly reshaped the linguistic practices of South India, giving rise to new forms of vocabulary, syntax, and hybridized expressions. This qualitative study examines the role of platforms such as WhatsApp, Instagram, and Twitter in facilitating a blend of regional languages like Kannada, Tamil, Telugu, and Malayalam with English, creating a unique digital vernacular. By employing interviews, focus group discussions, and content analysis of over 500 social media posts, the research delves into the sociolinguistic dynamics that drive these transformations. Findings reveal that digital communication is not only a space for creativity and identity assertion but also a reflection of socio-cultural shifts. The study identifies key patterns of code-switching, platformspecific adaptations, and the emergence of new linguistic norms, while also exploring generational perspectives on language evolution. This research underscores the importance of understanding digital language as a living, adaptive phenomenon that mirrors the region’s linguistic and cultural diversity. The implications extend to areas such as language education, digital literacy, and cultural preservation, offering a nuanced perspective on how social media redefines communication in multilingual contexts.
Headedness is widely used as an organizing device in syntactic analysis, yet constituency treebanks rarely encode it explicitly and most processing pipelines recover it procedurally via percolation rules. We treat this notion of constituent headedness as an explicit representational layer and learn it as a supervised prediction task over aligned constituency and dependency annotations, inducing supervision by defining each constituent head as the dependency span head. On aligned English and Chinese data, the resulting models achieve near-ceiling intrinsic accuracy and substantially outperform Collins-style rule-based percolation. Predicted heads yield comparable parsing accuracy under head-driven binarization, consistent with the induced binary training targets being largely equivalent across head choices, while increasing the fidelity of deterministic constituency-to-dependency conversion and transferring across resources and languages under simple label-mapping interfaces.
This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard.We focus on the variety spoken in Turkey (Uzunköprü) and ask how well a dependency parser trained on the existing Pomak Universal Dependencies treebank, which was built primarily from the variety that is spoken in Greece, transfers across dialects.We run two experimental phases.First, we train a parser on the Greek-variety UD data and evaluate zero-shot transfer to Turkish-variety Pomak, quantifying the impact of phonological and morphosyntactic differences.Second, we introduce a new manually annotated Turkish-variety Pomak corpus of 650 sentences and show that, despite its small size, targeted fine-tuning substantially improves accuracy; performance is further boosted by cross-variety transfer learning that combines the two dialects.
Replication package for "CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models" (DMLR 2026). The Contextual Emotional Inference (CEI) Benchmark is a dataset of 300 expert-authored scenarios for evaluating how well language models interpret pragmatically complex utterances in social contexts. Each scenario presents a communicative exchange involving indirect speech (sarcasm, mixed signals, strategic politeness, passive aggression, or deflection) where the speaker's literal words diverge from their actual emotional state. Three trained annotators independently labeled every scenario using Plutchik's 8 basic emotions and Valence-Arousal-Dominance ratings. This archive contains: • data/human-gold/ — 5 merged annotation CSVs (300 scenarios, 3 annotators each) • scripts/ — Pipeline, analysis, and HuggingFace upload scripts • config/ — Model definitions and pricing configuration • papers/dmlr2026/ — Paper source (LaTeX), bibliography, and figures • reports/dmlr2026/ — Baseline results (JSON) and LaTeX tables • LICENSE (MIT for code) and README.md The dataset is released under CC-BY-4.0. Code is released under MIT. GitHub: https://github.com/jon-chun/cei-tom-dataset-base HuggingFace: https://huggingface.co/datasets/jonc/cei-benchmark
Replication package for "CEI: A Benchmark for Evaluating Pragmatic Reasoning in Language Models" (DMLR 2026). The Contextual Emotional Inference (CEI) Benchmark is a dataset of 300 expert-authored scenarios for evaluating how well language models interpret pragmatically complex utterances in social contexts. Each scenario presents a communicative exchange involving indirect speech (sarcasm, mixed signals, strategic politeness, passive aggression, or deflection) where the speaker's literal words diverge from their actual emotional state. Three trained annotators independently labeled every scenario using Plutchik's 8 basic emotions and Valence-Arousal-Dominance ratings. This archive contains: • data/human-gold/ — 5 merged annotation CSVs (300 scenarios, 3 annotators each) • scripts/ — Pipeline, analysis, and HuggingFace upload scripts • config/ — Model definitions and pricing configuration • papers/dmlr2026/ — Paper source (LaTeX), bibliography, and figures • reports/dmlr2026/ — Baseline results (JSON) and LaTeX tables • LICENSE (MIT for code) and README.md The dataset is released under CC-BY-4.0. Code is released under MIT. GitHub: https://github.com/jon-chun/cei-tom-dataset-base HuggingFace: https://huggingface.co/datasets/jonc/cei-benchmark
In a multilingual and multicultural society like Malaysia, the spelling of names serves as a personal identifier and as a reflection of sociocultural and linguistic norms. While prior studies have examined spelling variation in educational and digital contexts, less is known about public perceptions of spelling variations of names. Against this backdrop, this study investigates how Malaysians perceive the acceptability of spelling variations. Using a mixed-methods design, 355 participants responded to a questionnaire comprising Likert-scale evaluations of ten real names, along with open-ended questions. Quantitative analysis revealed significant variability in participants’ acceptability ratings, though names that retained phonological clarity (e.g., Adrianah) were generally more accepted than others (e.g., Frrdy). No significant differences were found across gender, age, ethnicity, or professional background. Thematic analysis of qualitative responses highlighted two key influences on naming decisions: individual sociocultural factors (e.g., religious beliefs, family tradition) and sociolinguistic-aesthetic factors (e.g., pronunciation, trends, media influence). These findings suggest a growing tolerance towards spelling variations, possibly indicative the value of distinguishing one’s identity through names.
<p>The purpose of this study is to identify the specific features of applying folk pedagogy in developing communicative competence among students of nonlinguistic specializations and to assess its effectiveness in foreign language teaching. The methodology involves student surveys, educator questionnaires and interviews, a pedagogical experiment within educational institutions, and a Strengths, Weaknesses, Opportunities, and Threats (SWOT) analysis of the data obtained. Both quantitative and qualitative data collection methods are employed, enabling a comprehensive evaluation of the proposed approach. The main findings demonstrate a positive impact of folk pedagogy on the acquisition of linguistic norms, increased student motivation, and the development of communicative skills. The proposed approaches may contribute to improving the quality of language training and expanding the range of methodological tools available in foreign language education. The practical value of the study lies in the potential to integrate folk pedagogy into the foreign language learning process, which may enhance material acquisition and promote deeper cultural understanding. The recommendations offered could be used to improve higher education curricula. The application of folk pedagogy supports a more engaging, interactive, and natural learning experience, aligning with current educational trends.</p>
Abstract This paper introduces spectral attention, which filters the attention score matrix directly in the frequency domain via FFT/IFFT with learnable, per-head masks. This complements the time-domain view by enabling explicit control over low-, mid-, and high-frequency components of attention patterns. We study nine variants, including an adaptive mechanism that modulates masks from input content. On WikiText-2, Penn Treebank, and WikiText-103, the adaptive spectral variant consistently improves over standard attention, reducing perplexity by 10.7% on WikiText-2 and 15.3% on WikiText-103 in our setup. Analysis shows low-frequency components carry the most useful signal and that learned frequency preferences outperform fixed low/high/band-pass filters. These results indicate that frequency-domain processing is an effective complement for autoregressive transformer language modeling in our evaluated settings.
The correct use of Standard Albanian in public administration is vital for effective governance, accurate communication, and the maintenance of public trust. This study explores the extent to which standard language is used in Albanian state institutions through questionnaires and interviews with employees of the Ministries of Education, Defense, and Justice. The results indicate that official documents often contain linguistic errors and inconsistencies, reflecting shortcomings in language precision. These issues are largely caused by the absence of standardized document templates, insufficient attention to linguistic norms, and limited opportunities for staff training. Nearly one-third of respondents reported that errors frequently occur in emails and reports addressed to citizens.To address these challenges, the study emphasizes the importance of digital technologies, standardized communication models, and institutional reforms. The use of digital platforms with grammar-checking tools and shared templates can improve consistency and accuracy in official documents. In addition, promoting lifelong learning through continuous professional development and linguistic training can strengthen institutional efficiency, accountability, and public trust in public administration. Received: 06 October 2025 / Accepted: 12 December 2025 / Published: January 2026
Background: Aging is characterized by a decrease in olfactory, attentional, memory, language, and visuospatial/executive abilities. In this context, our study aimed to evaluate the potential effects of Rosmarinus officinalis L. (rosemary) and Carum carvi L. (caraway) essential oils (EOs) on aging. First, we assessed, in 402 participants, the age-related changes in olfactory functions (odor threshold, discrimination, and identification), gustatory perceptions (sweet, sour, salty, and bitter taste), cognitive functions (focusing on attention, memory, language, and visuospatial/executive functions), and their possible correlations with aging. To achieve this, olfactory function, gustatory perception, and cognitive abilities were evaluated in healthy participants across different age groups. Then, to evaluate the age-related decrease in trigeminal function (59 participants), we used rosemary and caraway EOs that contain carvone, limonene, and 1,8-cineole, all of which are considered typical trigeminal stimuli. Methods: Olfactory function was assessed with the Sniffin’ Sticks test, gustatory function by the Taste Strips test, and rosemary and caraway EOs by the ratings of odor pleasantness, intensity, and familiarity using a labeled hedonic Likert-type scale. Results: Olfactory function could be a potential early indicator of attentional, memory, language, and visuospatial/executive dysfunctions. Our data indicated that rosemary and caraway EOs were perceived without any significant decrease in odor pleasantness, intensity, and familiarity ratings in relation to aging. Conclusion: Our results suggest the potential bioactive effects of rosemary and caraway natural EOs as a new strategy to promote healthy aging.
In this paper, we combine methods used in penalized generalized empirical likelihood (GEL) frameworks with feature-extraction techniques that project textual data into numerical spaces. We develop recent penalization techniques used in GEL when the number of features gets very large compared to the sample size. We relate this approach to the Maximum Entropy (MaxEnt) principle used in several tasks of Natural Language Processing (NLP), in particular in Part Of Speech (POS) Tagging. Since the features belong to a large dimensional space, we propose a penalization method based on the dual representation of the original problem: this yields an explicit approximation of conditional probabilities of tags given the context. This method considerably reduces computational costs. As a byproduct, for different GEL methods, we obtain the corresponding POS-Tagging classifiers generalizing the MaxEnt method. We apply it successfully to the Penn-Treebank corpus with an error rate less than 5%.
This paper presents new resources and baselines for Dependency Parsing in Pomak, an endangered Eastern South Slavic language with substantial dialectal variation and no widely adopted standard. We focus on the variety spoken in Turkey (Uzunköprü) and ask how well a dependency parser trained on the existing Pomak Universal Dependencies treebank, which was built primarily from the variety that is spoken in Greece, transfers across dialects. We run two experimental phases. First, we train a parser on the Greek-variety UD data and evaluate zero-shot transfer to Turkish-variety Pomak, quantifying the impact of phonological and morphosyntactic differences. Second, we introduce a new manually annotated Turkish-variety Pomak corpus of 650 sentences and show that, despite its small size, targeted fine-tuning substantially improves accuracy; performance is further boosted by cross-variety transfer learning that combines the two dialects.
CHILDES is a paramount resource for language acquisition studies -- yet computational tools for analyzing its syntactic structure remain limited. Leveraging the recent release of the UD-English-CHILDES treebank with gold-standard Universal Dependencies (UD) annotations, we train a state-of-the-art dependency parser specifically tailored to CHILDES. The parser more accurately captures syntactic patterns in child--adult interactions, outperforming widely used off-the-shelf English parsers, including SpaCy and Stanza. Alongside the parser, we also release a Part-of-Speech tagger and an utterance-level construction tagger, which together form the open-source Syntactic Parsing Toolkit for Child--Adult InTeractions (CAIT). Through a detailed error analysis and a case study tracking the distribution of syntactic constructions across developmental time in CHILDES, we demonstrate the practical utility of the toolkit for large-scale, reproducible research on language acquisition.
This record contains the data, code, result tables, figures, and supplementary material for the study “How Much Does Corpus Choice Change Dependency-Distance Estimates?”. The archive includes derived analysis data, sentence-level dependency-distance metrics, treebank inclusion and exclusion records, manual provenance/comparability review tables, analysis scripts, configuration files, reproducibility tests, final result tables, figures, model/audit outputs, and the supplementary material submitted with the article. The raw corpora are public, versioned third-party resources from Universal Dependencies 2.18 and Glottolog CLDF 5.3. They are not redistributed in this archive because they are large and governed by their original license terms. Exact source URLs, versions, checksums, release dates, and retrieval information are documented in the included source manifest files. The main reproducibility files are README_ZENODO.md, data/source_manifest.csv, config/analysis.yaml, scripts/, src/, data/processed/, and results/.
Emotional memories persist within individuals and over generations. Past research has examined functions of parent-child memory sharing, but little work has assessed how the emotional qualities of memories are transmitted from parent to child and whether emotion transmission relates to memory content transmission. An understanding of how emotional memories are transmitted may be particularly important during adolescence, a developmental period marked by heightened sensitivity to emotional information, increasing independence from caregivers, and the emergence of mental health symptoms. The current study investigated the intergenerational transmission of parents’ autobiographical emotional memories to their teen offspring in healthy dyads using behavioral measures and natural language processing tools. We found that emotional valence and arousal linked to parents’ memories are transmitted from parent to teen. Greater parent-teen agreement in valence ratings was associated with greater parent-teen overlap in both subjective vividness ratings and objective memory content of individual memories. Subjective memory transmission was modestly related to lower mental health symptoms in teens after accounting for parent symptoms. These findings demonstrate that parents’ emotional memories are transmitted to their teens and provide preliminary evidence that autobiographical emotional memory transmission from parents could be a protective factor for mental health in adolescence.
Emotional memories persist within individuals and over generations. Past research has examined functions of parent-child memory sharing, but little work has assessed how the emotional qualities of memories are transmitted from parent to child and whether emotion transmission relates to memory content transmission. An understanding of how emotional memories are transmitted may be particularly important during adolescence, a developmental period marked by heightened sensitivity to emotional information, increasing independence from caregivers, and the emergence of mental health symptoms. The current study investigated the intergenerational transmission of parents’ autobiographical emotional memories to their teen offspring in healthy dyads using behavioral measures and natural language processing tools. We found that emotional valence and arousal linked to parents’ memories are transmitted from parent to teen. Greater parent-teen agreement in valence ratings was associated with greater parent-teen overlap in both subjective vividness ratings and objective memory content of individual memories. Subjective memory transmission was modestly related to lower mental health symptoms in teens after accounting for parent symptoms. These findings demonstrate that parents’ emotional memories are transmitted to their teens and provide preliminary evidence that autobiographical emotional memory transmission from parents could be a protective factor for mental health in adolescence.
This research explores the historical emergence of linguistic terminology in three languages—English, Uzbek, and Karakalpak—with special attention to the role of Latin, Greek, and Arabic heritage. It traces how borrowed concepts were nativized and localized in each linguistic setting. By juxtaposing five evolutionary stages in English with analogous processes in Uzbek and Karakalpak, the paper illustrates the interplay between international scholarly traditions and indigenous linguistic norms. The conclusions highlight both universal tendencies and language-specific particularities in the growth of terminological systems.