Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
This study examines the lexico-semantic features of Nigerian English as deployed in Bayo Adebowale’s Lonely Days, with a view to demonstrating how language reflects cultural identity and sociolinguistic realities in Nigerian contact literature. Anchored on the theoretical insights of Adegbija (1989) Lexico – semantic Nigerianisms, the study adopts a qualitative textual analysis of selected excerpts from the novel. It explores how meaning is constructed through lexical choices and semantic processes that deviate from Standard English norms but remain contextually meaningful within the Nigerian setting. The analysis reveals that Adebowale extensively employs lexico-semantic strategies such as transfer, coinages/neologisms, semantic shift or extension and borrowing from indigenous Yoruba language. These features function not merely as linguistic deviations but as deliberate linguistic resources that encode cultural values, social structures and indigenous worldviews. Expressions such as “a man of timber and caliber,” “bad market” and “face-me-I-face-you” illustrate semantic transfer and localisation, while items like “wrapper,” “amala” and “juju” exemplify semantic extension and cultural embedding. Furthermore, the incorporation of Yoruba lexical items and expressions enhances authenticity, reinforces setting and foregrounds the oral and performative tradition inherent in African literature. The study concludes that the lexico-semantic features identified in Lonely Days substantiate the existence and legitimacy of Nigerian English as a distinct variety. They also demonstrate how Nigerian writers creatively appropriate the English language to reflect indigenous experiences, thereby expanding its expressive capacity. Ultimately, the paper affirms that lexico-semantics is central to understanding the intersection of language, culture and identity in Nigerian literary discourse. Key words: Lexico-semantics, Nigerian English, Lonely Days
This is the Adult–Child Touching Hands Dataset (AC-THD): a dataset of 59 image stimuli depicting hand-to-hand touch between an adult and a child, systematically categorized into three emotional valence classes: negative (n = 15), neutral (n = 29), and positive (n = 15). The tool emphasizes the central role of the hands in conveying emotional information within an adult–childinterpersonal context.The images and the hand-to-hand positions were developed by four trained psychologists, and were first assigned into three emotional valence categories: negative (e.g., one person tightly pinching or forcefully gripping the other person’s hand); neutral (e.g., the two individuals’ fingers lightly overlapping or their wrists touching), and positive (e.g., one person gently stroking the back of the other person’s hand or the two individuals interlocking their hands). Images were initially assigned provisional alphabetical labels, which were subsequently refined through a percentile-based procedure in accordance with validation data. The AC-THD acquisition phase involved a mother (Caucasian, aged 48, right-handed) and her child (Caucasian male, aged 10, right-handed). Images were acquired by a professional photographer using a Canon EOS 1100D digital SLR camera with a 12.2 megapixel APS-C CMOS sensor and a resolution of 4272 × 2848 pixels, equipped with an EF-S 18–55 mm f/3.5–5.6 lens. The built-in flash (Guide Number 9.2 at ISO 100) was manually triggered for each shot to standardize illumination and minimize shadows. The camera was positioned on a tripod to maintain a fixed 90° angle. All images were captured in color in sRGB mode, with framing restricted to the hands and forearms. In the validation study, images were rated by 321 participants on valence, using a numeric rating scale ranging from 0 to 10, where 0 indicates a strongly negative image, 5 a neutral image, and 10 a strongly positive image, with anchors at 0 (“strongly negative”) and 10 (“strongly positive”). Mean emotional valence ratings were calculated for each image, along with standard deviations (SDs), mode, minimum and maximum values, and percentiles. Stimuli were finally classified as negative, neutral, or positive based on the 25th and 75th percentiles of the distribution of mean valence ratings. Accordingly, images with mean valence scores ≤ 4.25 were classified as negative, those with mean valence scores > 4.25 and < 7.23 as neutral, and those with mean valence scores ≥ 7.23 as positive. Therefore, the final dataset is comprised of: N = 15 images in the negative category; N = 29 images in the neutral category; N = 15 images in the positive category. Study validation of the dataset also examined potential differences in emotional valence ratings as a function of participants’ socioeconomic status (SES) and gender, which are reported as Supplementary Materials. These images can be used across a range of research fields, including emotion elicitation, neuroscience, and psychophysiological studies, as well as for assessing affective responses in individuals with histories of supportive or adverse tactile caregiving, given their potential to evoke autobiographical memories.
This is a retrained Slovenian model for the Trankit v1.1.1 library for multilingual natural language processing (https://pypi.org/project/trankit/), trained on the concatenation of the SSJ UD treebank of written Slovenian (featuring fiction, non-fiction, periodicals and Wikipedia texts) and the SST UD treebank of spoken Slovenian (featuring transcriptions of spontaneous speech in various settings). It is able to predict sentence segmentation, tokenization, lemmatization, language-specific morphological annotation (MULTEXT-East morphosyntactic tags), as well as universal part-of-speech tagging, morphological features, and dependency parses in accordance with the Universal Dependencies annotation scheme (https://universaldependencies.org/). In comparison to its counterpart models trained on SSJ (http://hdl.handle.net/11356/1963) or SST datasets only, this model yields a significantly better performance on spoken transcripts and an identical state-of-the-art performance on written texts. The model can therefore be recommended as the default, 'universal' Trankit model for processing Slovenian, regardless of the data type. To utilize this model, please follow the instructions provided in our github repository (https://github.com/clarinsi/trankit-train) or refer to the Trankit documentation (https://trankit.readthedocs.io/en/latest/training.html#loading). This ZIP file contains models for both xlm-roberta-large (which delivers better performance but requires more hardware resources) and xlm-roberta-base. Version 1.3 was trained on the same data as version 1.2, except that spoken SST data (UD v2.15) was augmented by colloquial (non-standardized) transcriptions of spoken Slovenian alongside the standardized ones. The resulting model achieves state-of-the-art performance on both standardized (e.g. "včasih govorimo takole") and colloquial speech transcriptions (e.g. "včas govorimo tkole"), without affecting the performance on written data.
Generic nouns such as Sache and Ding pose a challenge for semantic annotation due to their referential underspecification and context-dependent meaning. Although frequently classified under categories like {artefact} or {object}, their actual referents often belong to abstract or cognitive domains, as in Der Placeboeffekt ist eines der faszinierendsten Dinge in der Welt der Medizin. Drawing on valency grammar, this study shows that these nouns activate different argument structures depending on their syntagmatic environment, reflecting semantic flexibility and combinatorial variability. Lexical databases such as GalNet or GermaNet frequently assign multiple synsets to these nouns, illustrating their ontological ambiguity. This paper examines whether large language models (LLMs) can replicate this nuanced classification. Using a gold standard corpus annotated by linguists, we implement a two-step prompting strategy —supplying LLMs with predefined semantic tags and contextual windows— to test their performance. The results underscore the limitations of current LLMs in dealing with the lexical underspecification of generic nouns, even when provided with an extended context window. These findings contribute to ongoing discussions on the automation of semantic tagging and point to meaningful ways in which AI systems can complement human expertise in natural language processing tasks.
This article examines the influence of sociolinguistic factors on language development in Uzbek and Russian media texts. Media discourse is one of the most dynamic spheres of language functioning because it reflects social change, political processes, technological innovation, globalization, cultural values and communicative needs of society. The purpose of the article is to analyze how sociolinguistic factors such as bilingualism, language contact, globalization, digital communication, social stratification, audience orientation, language policy and media genre influence the lexical, semantic, stylistic and pragmatic development of Uzbek and Russian media texts. The study is based on descriptive, comparative, sociolinguistic and discourse-analytical methods. The results show that Uzbek and Russian media texts are actively influenced by neologisms, borrowings, anglicisms, code-switching, colloquialization, terminological renewal and changes in stylistic norms. Uzbek media discourse demonstrates a tendency toward national language development together with active borrowing from English and Russian, while Russian media discourse shows a high level of lexical innovation, hybrid forms and stylistic diversification under the influence of global and digital communication. The article concludes that sociolinguistic factors play a decisive role in the development of media language because they determine how linguistic units are selected, adapted, normalized and disseminated in public communication.
Since sound is a wave, it has all the properties pertaining to any wave, including frequency, amplitude, waveform, and duration, which, translated into musical terms, describe pitch, dynamics, tone, and duration. These four structural elements of sound each alter the listener’s experience and, by extension, the arousal, valence, and, in particular, emotions elicited by them. The significance of emotions and music within both modern society and ancient cultures suggests that any correlation or causation between the two must be inherently important, and the near universal ubiquity and experience of both music and emotion is further evidence of this. However, in order to determine if there was a relationship between the two factors, meta-analysis of both qualitative and quantitative data was conducted, revealing that the alteration of any single or combination of structural elements resulted in variation of the onset predictability, creating psychological engagement and reward. This positive correlation between sound and the degree of response on emotional, arousal, and valence levels within both the qualitative and quantitative data suggests that, depending on the pitch, dynamics, timbre, and tempo of the auditory stimuli, the affective state experienced by the listener varies. Additionally, the higher valency and arousal ratings expressed upon listening to expressive auditory stimuli, which involves the augmentation of tempo as well as an increase in pitch and dynamics and a decrease in duration, ultimately result in surprise, fear, or, most typically, happiness.
This research is situated at the intersection of digital humanities, the history of emotions, and computational linguistics. The article presents the results of the sentiment analysis of the epistolary heritage and diary entries of three key figures of the Russian monarchy: Catherine II, Alexander I, and Nicholas I. The total corpus of analyzed texts amounted to over 2 million word usages. Using models of deep learning BERT (XLMRoBERTaLarge and Conversational RuBERT) adapted for historical texts, the authors reconstruct the emotional dynamics of communication throughout a turbulent century—from Enlightened absolutism to the crisis of the Nicholas system. The study confirms the hypothesis of a stable correlation between the genre of the document (official letter vs. private diary) and the degree of emotional expressiveness, as well as identifies specific lexical markers of anxiety during periods of political instability (on the eve of the Decembrist revolt and during the Crimean War). Methodologically, the research is based on three conceptual foundations. Firstly, it is the theory of emotional communities by B. Rosenwein, according to which emotions are constructed within social groups with shared values and norms of expression. Secondly, it adopts the semiotic approach of Yu. M. Lotman in studying the everyday behavior of the Russian nobility. Thirdly, it employs methodologies of computational text analysis. The hypothesis of this study is as follows: the emotional tone of the personal correspondence and diaries of Russian monarchs is not so much a spontaneous expression of an individual psychological state, but rather a ritualized social action, subject to the cultural codes of the era and genre canon. The aim of the study is to conduct a comprehensive historical-linguistic analysis of the emotional tone of the epistolary and diary heritage of Russian statesmen of the 18th-19th centuries using digital text processing methods, to identify stable emotional patterns and their connection with historical-biographical context. It has been established that the sentimentalist tradition of the late 18th century paradoxically combined with hypertrophied emotional restraint in official communication, creating an effect of "emotional dissonance," which was resolved in the literature and journalism of the 19th century. This work contributes to the methodology of analyzing historical texts, demonstrating the possibilities and limitations of NLP tools when working with archaic vocabulary and bilingual corpora (Russian-French linguistic dualism).
With the rapid growth of computer technology, the application of the Tamil language has expanded globally. Overcoming the early limitations of 8-bit ASCII encoding systems, the introduction of 16-bit encoding standards such as Unicode and TACE16 has enabled the seamless use of Tamil characters across all computer and mobile platforms. This paper comprehensively examines the historical evolution of Tamil printing and typing systems to its modern spread across emails, blogs, online dictionaries, digital libraries, and social media. Furthermore, it evaluates the technical advancements achieved in modern Natural Language Processing (NLP) domains, including spell checkers, morphological analyzers, syntactic treebanks, Tamil Optical Character Recognition (OCR), speech technologies (TTS & Speech Recognition), and Artificial Intelligence-driven e-governance chatbots. This study highlights how the Tamil language maintains sustainability and international stature in the digital era through contributions from the International Forum for Information Technology in Tamil (INFITT) and various research organizations.
Campaign language in the United States has grown markedly polarized, with candidates increasingly speaking in party-distinctive ways. Yet the mechanisms driving this linguistic divergence remain insufficiently understood. This study proposes identity–policy fusion, a framing strategy in which candidates embed distinctive biographical vocabulary into policy statements, as one factor shaping this divergence. By framing policy as an authentic extension of lived experience, fusion may tie stances to in-group identity and biographical authority, raising the psychological and social costs of disagreement. Using computational text analysis of 41,842 policy statements from 3,343 U.S. House candidates in the 2018, 2020, and 2022 election cycles, the study operationalizes fusion as term frequency–inverse document frequency (TF–IDF)-weighted lexical overlap between biographies and policy texts, and polarization as a statement’s relative similarity to in-party versus out-party linguistic norms within policy domains. Ordinary least squares regression shows that fusion is significantly associated with higher polarization in campaign language, with the association approximately 26 percent stronger for Republicans than Democrats. This partisan asymmetry is consistent with fusion serving as an alternative source of legitimation under conditions of contested institutional authority, illuminating a potential mechanism through which elite messaging may harden partisan boundaries.
Considering the trends towards universalization and differentiation in procedural law, the paper hypothesises that similar patterns exist in the language of procedural law. It identifies the need to study the linguistic and textological distinctiveness of the texts of the Criminal Procedure Code, Civil Procedure Code, Arbitration Procedure Code, and Administrative Procedure Code. The study reveals certain features of the language of these laws in terms of semantics, grammar, vocabulary, syntax, and textology. Examples of homonymy, synonymy, oxymoron, violations of linguistic norms in terms of spelling and correct word order, the use of identical headings, combining heterogeneous elements into structural elements of a normative legal act, and lexical patterns are given. The normative meaning of headings is revealed; examples reflecting the problems of the structure of procedural laws are given. The author names the striking textual feature of the Criminal Procedure Code of the Russian Federation, namely Art. 5, which includes the definition of the main terms used in the law. The syntactic complication of the language of procedural laws has been noted in recent years. The lexical features of individual procedural laws are revealed, demonstrating their substantive originality. An assumption is made about the influence of the quality and features of the normative legal language on the language of judicial acts and the language of legal science.
Abstract Recognizing others’ emotions is central to social interaction. Traditional biological psychology infers emotional responding via laboratory measures, whereas contemporary computer vision algorithms claim to identify emotions unobtrusively from facial video. However, the validity of such algorithms for classifying spontaneous emotional responses occurring without explicit communicative intent remains debated. We compared established psychophysiological measures (EEG, facial EMG, EDA activity) with the open-source facial behavior toolkit OpenFace for classifying participants’ spontaneous responses during free viewing of happiness-inducing, disgust-inducing, and neutral pictures. Participants provided valence and arousal ratings and later selected the basic emotion that best matched their reaction which served as the classification criterion. Using within-participants single-trial support vector machine (SVM) classification, EEG achieved the highest accuracy (40%), followed by facial EMG (37%); OpenFace reached 36%. All methods except EDA exceeded chance performance (33.3%) and were lower compared to human raters (48%). Predictions declined slightly for across-participants SVMs, being at chance for OpenFace and EDA. The results indicate that in principle both, psychophysiological measures and video-derived facial action units, can capture diagnostically relevant aspects of emotional responding during picture viewing, but that their performance is limited when expressions are spontaneous and not produced for communicative purposes. Inter-individual variability in expressivity and physiological responding likely contributes to these limitations and should be considered when deploying automatic emotion recognition in research or applied settings.
BACKGROUND: Flat-detector computed tomography (FDCT) is increasingly used for peri-interventional cerebral imaging but is associated with a relatively high radiation exposure. Copper (Cu) filtration may reduce radiation dose. However, its impact on cerebral image quality and intracranial hemorrhage detection remains unclear. METHODS: In this retrospective single-center study, 31 patients undergoing neurointerventional procedures with intraindividual FDCT acquisitions with and without Cu filtration were analyzed. Quantitative image quality was assessed using contrast-to-noise ratio (CNR). Qualitative image analysis and intracranial hemorrhage detection were independently evaluated by two readers blinded to Cu filtration status using five-point scales. RESULTS: Cu filtration resulted in a significant radiation dose reduction of 25.9% for both entrance skin dose (145.19±13.18 mGy vs 195.89±18.05 mGy) and dose-area product (41.52±3.77 Gy·cm² vs 56.01±5.16 Gy·cm²), respectively (P<0.001). No differences in CNR were observed for unfiltered vs Cu-filtered FDCT (basal ganglia: 4.73±2.04 vs 4.37±1.99, P=0.419). Qualitative image ratings were similar between techniques (supratentorial cortex: 2.27±0.66 vs 2.08±0.75, P=0.089), with very good inter-reader agreement (κ=0.86; 95% CI: 0.80 to 0.91). All intracranial hemorrhages were correctly identified by both techniques. Correct exclusion of intracranial hemorrhage was 15/16 with Cu filtration and 14/16 without Cu filtration, without statistically significant difference. Differences were limited to hemorrhage mimics (n=2) and minor variations in diagnostic confidence without affecting binary classification. CONCLUSION: Cu filtration in cerebral FDCT enables substantial radiation dose reduction while preserving image quality and intracranial hemorrhage detection, supporting its clinical implementation as a practical dose optimization strategy for peri-interventional imaging.
Anonymised dataset for the article "The diachronic evolution of null subjects in French and Venetian: A treebank, statistical approach"
Respect plays a crucial role in successful interpersonal and intercultural communication. However, differences in linguistic norms and pragmatic conventions often lead to misunderstandings between speakers of different languages. This article examines pragmatic failures in expressing respect in English-Russian cross-cultural communication. The study aims to identify linguistic and cultural factors that cause misinterpretations of respect and to analyze how respect is pragmatically encoded in both languages. The findings suggest that pragmatic failures frequently arise from divergent politeness strategies, speech act realizations, and sociocultural expectations embedded in English and Russian communicative practices. The study emphasizes the importance of pragmatic awareness in developing intercultural communicative competence.
The current study examines how the presentation of separable components of rapport building in forensic interviews with children affect lay perceptions of both the child victim and defendant guilt. Mock jurors will read a forensic interview in which a child either alleges sexual abuse by an adult male perpetrator or does not. Interviews will either contain a ground rules phase or only a brief introduction. Thus, the study adheres to a 2 (ground rules: present or absent) x 2 (disclosure: present or absent) between-subjects design). Effects of ground rules on perceptions of defendant guilt are not expected. However, it is expected that disclosure will affect ratings of defendant guilt. It is also expected that both ground rules and disclosure will affect perceptions of the child. It is expected that the child will be perceived more positively than when both are present.
The article examines taarof as one of the key phenomena of Persian communicative culture, reflecting the peculiarities of speech etiquette and the system of interpersonal interaction in Iranian society. The subject of the study is taarof as one of the key phenomena of Persian communicative culture, reflecting the specifics of speech etiquette and the system of interpersonal interaction in Iranian society. This phenomenon represents a stable system of politeness formulas that regulate communicative behavior and ensure adherence to culturally conditioned norms of communication in various situations of verbal interaction. Special attention in the study is given to the lexical-semantic features of taarof formulas in modern Persian, as well as their functioning in real communication. The author thoroughly examines aspects of the topic such as the typology of taarof formulas, their pragmatic characteristics, and their role in organizing indirect speech acts. Particular focus is given to how these formulas are realized in various types of discourse, including conversational speech, artistic texts, and media sources. The methodological basis of the research consists of contextual, semantic, and pragmatic analyses aimed at identifying the characteristics of taarof formulas' functioning in various communicative situations. A corpus-based approach is employed, including the analysis of dialogical fragments of conversational speech, artistic texts, and quotes from films. The scientific novelty of the research lies in the systematic description of taarof formulas in terms of their lexical-semantic typology and pragmatic functioning in modern Persian. The work proposes a comprehensive approach to analyzing these units as elements of the system of indirect speech acts, taking into account not only their linguistic form but also the culturally conditioned mechanisms of interpretation. The classification of taarof formulas is refined by distinguishing hybrid types that combine several communicative functions. The study establishes that taarof formulas play a key role in organizing interpersonal interaction, ensuring the mitigation of speech acts, expressing implicit refusals, and regulating the social distance between communication participants. It is shown that their meaning is formed in the process of dialogue and depends on context, social roles, and communicative expectations. The conclusions drawn confirm that taarof is an important mechanism for maintaining the ritualized nature of communication and reflects the specificity of the Persian linguistic worldview.
Language is a living organism that evolves alongside technological and social advancements. This paper examines the phenomenon of neologisms—newly coined words or expressions—and their pervasive role in contemporary English mass media. The study categorizes recent neologisms based on their morphological formation processes, such as blending, compounding, and functional shift. Furthermore, it analyzes how mass media acts as a primary catalyst for the popularization of these terms. By investigating digital journals, social media platforms, and news broadcasts, the research highlights the pragmatic functions of neologisms in creating concise, engaging, and culturally relevant communication. The findings provide insights into the current trends of English lexicology and the impact of the digital age on linguistic norms.
This article analyzes lexical norms and their violations in media language through a comparison of Azerbaijani and English. Globalization and social media have increased lexical deviations, including misuse of foreign words, semantic inconsistencies, and slang. These issues reduce clarity and accuracy in communication. Strengthening editorial control and journalists’ language competence is essential. Keywords: lexical norm, media, digital, style, text, conceptual, global
Inside a modern language model sits a single internal direction that tracks how positive or negative a sentence feels. We show how to find this valence axis (V-axis) from just 9 emotion category names plus 50 short narrative paragraphs per emotion -- about 1,500 fewer labels than the usual supervised approach -- and that the same direction appears in vision, audio, and human-brain encoders never jointly trained. The recipe: embed nine emotion-anchored story sets in a frozen encoder, take the top principal direction of the nine averaged embeddings. Projecting new inputs onto it captures 93% of supervised performance on SST-2 (Llama-3-8B-Instruct, AUC 0.772 vs. 0.828), correlates with human valence ratings on 11,811 EmoSet images at r=0.636, reaches AUC 0.906 on ESC-50 audio (p<2.2e-15), and AUC 0.720+/-0.055 on EEG from 123 subjects (p<3.65e-8). The direction is mechanistically active: ablating it collapses sentiment accuracy by 5.5-37.2 pp across three LLMs vs. at most 0.88 pp for matched random directions (z>12). A 2-parameter classifier trained on text labels transfers to images (AUC 0.961), audio (0.764), and brain recordings (0.828) without target-modality labels; a generic 16-D subspace stays at chance (0.525). The recipe is bounded to continuous attributes -- seven tests on categorical concepts return near-chance -- and steering is family-specific (Llama/Mistral yes, Qwen/Gemma no).
The study examines the microstructural Russian narrative competence of native Chinese speakers studying Russian outside the Russian-speaking environment. The relevance of the study stems from the need for an in-depth analysis of the mechanisms underlying the development of written speech production in a foreign language and the identification of specific difficulties of Chinese learners producing written narratives. The research is aimed at identifying the key microstructural characteristics of Chinese students’ written narratives, classify their typical deviations from the Russian linguistic norm, and establish systematic links between microstructural features and different types of errors. The material includes 60 narratives produced by Chinese learners during an experiment with MAIN. The methodology combined quantitative and qualitative approaches: microstructural features markup, error classification, statistical procedures, and correlation analysis. The results show that text length does not correlate with the number of errors, whereas low lexical diversity is associated with higher error frequency. The research demonstrates that the most problematic area is morphology; the most frequent errors include “frozen” initial and oblique forms, and errors in aspect, case, and gender. Graphic errors are also frequent and are caused by indistinction of consonant pairs and an underdeveloped orthographic word image. Lexical difficulties primarily concern verbs of motion. Correlation analysis reveals stable error clusters that reflect different sublevels of acquisition of the target language system. The written speech of Chinese learners outside the language environment is characterized by certain difficulties which are not typical for other learners of Russian as a foreign language. The findings may enhance teaching, particularly in the aspect of paradigms, verb aspects, and verbs of motion.
Embodied theories of language emphasize the role of sensorimotor experience in linguistic knowledge. Central to testing these theories is the creation of large datasets of linguistic norms, which contain judgments about a word's sensorimotor associations and can be used to predict human behavioral or brain data - sometimes in contrast to competing variables, such as those derived from distributional language models. Yet many of these datasets contain judgments about words in isolation, despite the fact that most words are ambiguous, making it difficult to determine which meaning of a word is characterized by its rating (e.g., "wooden table" vs. "data table"). In the current work, we introduce a new lexical resource (directly inspired by the Lancaster sensorimotor for 112 English words, each rated in four different contexts (448 sentences total). We demonstrate: first, that these ratings encode overlapping but distinct information from the Lancaster sensorimotor norms; second, that decontextualized ratings likely reflect the more dominant meaning of ambiguous words; third, that homonyms have more distinct sensorimotor profiles than polysemes; fourth, that the contextualized sensorimotor distance between two uses of an ambiguous word predicts human judgments about semantic relatedness; and fifth, that ratings derived from GPT-4 align reasonably well with human judgments. We conclude by suggesting that contextualized ratings like these can be used both to inform competing theories of semantic representations and also to evaluate or "probe" the ability of LMs to recover sensorimotor information.
Quantitative phylogenetics in historical linguistics has relied almost entirely on lexical cognate data. This study asks a different question: how much genealogical signal can be recovered from structural features extracted from annotated corpora, and whether it survives at deep time depths. We compute 29 structural features-including Shannon entropies of dependency direction and of dependency-relation distributions, relation-specific directionality ratios, dependency-distance measures, and constructional ratios-across 25 Transeurasian languages from the five proposed groups (Turkic, Mongolic, Tungusic, Japonic, and Koreanic) and three outgroups (Chinese, Vietnamese, and Hindi), 28 languages in all. Most of the Tungusic and Mongolic languages have no running-text corpus, so we built new Universal Dependencies treebanks for them by glossing example sentences from reference grammars; thirteen are used here. Each feature was tested for phylogenetic signal (Pagel's λ and Blomberg's K, with FDR correction) under four competing reference topologies, and the features that passed were used for tree inference (Bayesian inference in MrBayes, with Neighbor-Joining as a check). The same pipeline was first run on Indo-European in a companion study, where it recovers only individual subgroups and does not resolve a stable tree. At the depth proposed for the Transeurasian family (a Proto-Transeurasian root of about 9000 years before present), the structural signal was not enough to reconstruct the family's internal relationships. The signal tests favoured a flat three-way division of the major branches (7 strict/20 relaxed features) over any nested hypothesis (≤2 strict features each), and the strongest signal lay in core word-order parameters (e.g., object direction, λ = 1.00, K = 6.06). But both Bayesian and distance-based inference returned near-complete polytomies: although the chains converged (ASDSF < 0.01), no branch reached a posterior probability above 0.75, and none of the three multi-language branches (Turkic, Mongolic, or Tungusic) was recovered. The outgroup test made the reason clear: Hindi, which is Indo-European but SOV, grouped with the head-final Transeurasian languages rather than with the other two (head-initial) outgroups, so the features are tracking typological similarity, not shared descent, at this depth. The study contributes 13 new treebanks for poorly documented languages, a reproducible framework for testing how much genealogical signal structural features carry, and direct evidence that, at Transeurasian time depths, this signal reflects typology rather than genealogy.
This paper develops a Weberian ideal‑type model explaining how linguistic norms diffuse through institutional environments via administrative incentives rather than decentralized cultural drift or explicit state-led language planning. It identifies a five‑stage sequence—Institutional Access, Resource Mobilization, Norm Diffusion, Coalition Reinforcement, and Norm Enforcement—through which upstream gatekeeping bodies, professional standards committees, and compliance-oriented organizations convert optional vocabulary into de facto mandatory administrative norms. The model distinguishes coordinated campaigns from emergent isomorphic steering, integrates sociolinguistics, institutional sociology, social psychology, political theory, and administrative law, and grounds each stage in directly observable primary-source evidence including organizational style guides, discourse‑analytic studies, automated language‑governance frameworks, and cross-domain case material from environmental governance, public health, and corporate HR. It specifies explicit boundary conditions and four failure modes—public ridicule, institutional pluralism, preference‑falsification collapse, and statutory friction—and provides empirical operationalization protocols using corpus linguistics, administrative documentation, and tribunal data. Rather than describing any single historical episode, the paper offers a falsifiable analytical framework for investigating how linguistic norms emerge, stabilize, or reverse within formal institutional systems. The next version will include datasets from the Islamic state of Iran and various other groups.
The subject of this study is the concept of «crime» from the perspective of cognitive linguistics. The object of the study is language as a cognitive mechanism involved in the representation and transformation of information. The objective of the research is to study the linguistic means used to represent the concept of «crime» in Russian and English, as well as to reveal common and specific features of the concept in question related to national specifics. The paper analyzes the terms denoting illegal and criminal acts. The authors come to the conclusion that the word «offense» is a generic term meaning any act that violates any legal norms in the Russian language, however in English, the broadest term is the lexeme «crime». Great attention is paid to the analysis of words verbalizing elements of the concept of «crime» in the Russian and the English languages. The paper considers lexical means reflecting various aspects of criminal activity such as types of crimes, their signs and stages, persons committing crimes, etc. These words form the public perception of crime and illustrate its significance by representatives of a particular nation. The study uses dictionary definitions to research the meaning of terms related to illegal acts; conceptual analysis to assess the structure and content of the concept of «crime»;a comparative analysis and an associative experiment to determine which components of the concept of «crime» are specific to Russian and English linguistic consciousness and which are universal. The scientific novelty of this study lies in the comparative analysis of the phenomenon of «crime» in the Russian and English linguistic worldviews by studying the linguistic means representing the given concept and conducting an associative experiment aimed at identification of universal and national-specific features of the concept of «crime». The main outcome of the work is the obtained data that the national and cultural specifics of the concept of «crime» is manifested at the level of the elements of the concept, in different ideas about the characteristic features of the criminal and the significance of the types of crimes in particular. For the first time, the relevance of conducting a linguistic analysis of the concept «crime» for law enforcement officers has been substantiated. Police officers' knowledge of the national characteristics of the perception of the phenomenon of «crime» and its components by representatives of certain cultures expands their possibilities for combating crime and allows for an accurate assessment of the public danger in cases where citizens of a certain nationality attempt to violate the law.
Dependency distance is a key measure of syntactic complexity and processing constraints, but as a scalar cannot capture how information is distributed between a head and dependent. We analyze dependency spans—the endpoints and intervening words—as position-aligned units. Interveners lie outside the binary dependency but constitute its sequential processing context. Using Universal Dependencies treebanks and XGLM-2.9B, we estimated word-level surprisal for distances 4–10 across 22 languages, yielding 154 mean curves. Dynamic time warping and clustering identified an approximately monotonic decline and a nonmonotonic contour with an initial decline, stable middle, and final rise. Membership was stable across distances in 20 languages; Russian had three Type 1 and four Type 2 curves, whereas German had six Type 1 and one Type 2 curve. Segmented models favored three stages for both types, differing mainly in the final stage. Principal component analyses revealed a more dispersed latent structure in the middle stage and concentration around fewer variables at the edges. The final-stage contrast was associated with dependency direction and verb roles. Dependency-span analysis thus complements dependency syntax with a position-sensitive account of linear realization, revealing cross-linguistic information patterns that dependency length alone cannot recover.
ABSTRACT Communication and language exist as complex and intertwined systems influenced by several overlapping dimensions that cannot be comprehended separately. In this research, the multi-dimensional nature of communication and language is explored using an interdisciplinary approach which considers linguistic aspects, sociocultural background, technological dimension, and cognition. The study employs a sequential mixed methods approach which includes the qualitative analysis of 48 cases of authentic communication (24 instances of face-to-face interaction and 24 examples of digital interaction) and the quantitative survey of 312 participants from three linguistic communities. Findings indicate that technological platforms have the potential to radically alter traditional linguistic norms by emphasizing sociocultural and cognitive components in such a way that hybrid linguistic practices develop. Multimodality, adaptive politeness, code-meshing, and platform-related linguistic features were used by participants in order to accomplish linguistic goals in different contexts of cognitive and sociocultural demands. Results from the factor analysis showed four correlated but independent dimensions that explained 68.4% of the variance, whereas results from multiple regression models showed significant predictors of technological factors on linguistic adaptation (β =.41, p <.001) and sociocultural adaptation (β =.37, p <.001). There was a significant negative correlation between cognitive load and communicative satisfaction in high-stakes intercultural digital communication. This paper points out some research gaps that still exist in terms of examining the combination of these aspects and presents a multidimensional approach. The results can be useful from a theoretical point of view and have several applications. Some limitations regarding the sample size and measuring the cognitive process are discussed.
The proliferation of artificial intelligence (AI) technologies has transformed academic research practices across higher education institutions around the world. While AI-powered tools such as large language models, neural machine translation systems, and automated writing assistants offer significant opportunities for enhanced research productivity, but also pose complex ethical and linguistic challenges. These issues are particularly acute in multilingual and postcolonial academic contexts, where researches are often required to navigate between multiple languages, traditions, and knowledge systems. This study explores the ethical implications of AI-assisted research in the Algerian higher education and presents a context-sensitive framework for responsible AI use in multilingual hypo thesis-based research. Using a convergent mixed-methods approach, this study examines the experience of 45 master students at the English department of Batna 2 University enrolled in a research methodology course. The results indicate that structured AI ethics education significantly improves students’ ethical decision-making skills, awareness of algorithmic bias, and their ability to engage critically with AI generated outputs. The study also demonstrates that mainstream AI systems tend to favor Anglophone linguistic norms, which could marginalize local epistemologies and multilingual scholarly practices. In response, the article proposes the Algerian AI Ethics Triade: a framework based on the principles of Transparency, Cultural Relevance, and Human Primacy. The framework offers practical recommendations for researchers, educators, and policymakers who want to incorporate AI into academic research scholarly without compromising integrity, linguistic diversity, and intellectual autonomy. The findings contribute to emerging discussions on responsible AI governance and offer implications for multilingual higher education systems across the Middle East and North Africa (MENA) region
The current research has discussed the challenges associated with teaching poetry in secondary school in Nepal among English teachers. The study utilized a qualitative research design and employed a narrative inquiry as the method of data collection, involving interviews, recordings, and diary notes. Ten community secondary schools in Kathmandu Metropolitan City were selected through purposive sampling, and one English teacher of each school was interviewed. Data were interpreted using thematic analysis. The results showed that the teachers have faced major linguistic problems, such as problems with archaic words, lexical and semantic complexities, figurative language, and odd syntactic structures. Cultural issues also came up as they reflected barriers to do with societal norms, traditions and values, and the cognitive issues were associated with understanding and interpretation of poetic forms. The respondents emphasized that it is especially challenging to teach poetry in the classroom, and the most common approach to teaching poetry in community schools is translation. The paper highlights the importance of contextually sensitive pedagogical approaches, teacher development, and curriculum-scaffolding to mitigate such factors and improve student engagement and comprehension of poetry.
Paper 6 (Silva 2026) introduced BPE Mean Vocabulary Morpheme Length (VMML) as a writing system classifier and showed that the Voynich Manuscript occupies a discriminant zone (VMML = 5.918, 95% CI 5.77-6.05) above all 15 tested alphabetic natural languages. This paper (v2.5) expands to 71 corpora across 40+ languages and reports six extended analyses: (1) Alphabetic ceiling confirmed at 5.76; (2) Tagalog (VMML=5.914) is the sole natural-language entry into the Voynich CI, but BC=0.202 distinguishes it from Voynich (BC=0.361); (3) Romanization inflates VMML by 2.4-5.3 units (methodological confound). Extended analyses: (4) Currier A vs B: delta VMML=+1.27, delta CBMI=+0.16 bits -- two quantifiably distinct writing registers; (5) BC coherent across all 7 manuscript sections (CV=6.7%) -- single writing system confirmed; (6) 3D discriminant (VMML x BC x CBMI): Voynich isolated, nearest natural-language neighbor Irish at distance 0.17; (7) Six named hoax mechanisms (monoalphabetic, Vigenere/barbavara, Vigenere/Italian-Knowles 2026, null insertion, syllabic compression, vocabulary shuffle) each fail all three criteria simultaneously; (8) BC orthogonal to all classical textual metrics (|r| < 0.23 vs entropy, TTR, hapax, Zipf) -- genuinely new structural dimension. All code and six extension scripts publicly available in companion repository. v2.3 (2026-06-08): Section 5.9 added - per-folio Currier A/B reanalysis using the Gaskell and Bowern (2022) canonical corpus (36,361 tokens, min_freq=5 BPE). Cross-boundary mutual information (CBMI) identified as primary discriminant: CBMI_A = 1.97 bits vs CBMI_B = 1.51 bits, Cohen d = -1.01, permutation p less than 0.001 (n = 10,000 shuffles, Bonferroni-corrected). CBMI survives within-quire control (pooled nA=46, nB=33; permutation p = 0.0008; Fisher combined within-quire p = 0.001), ruling out manuscript section as a confound. All three metrics (BC, BPE-ratio, CBMI) show A greater than B direction. Fisher combined full-corpus: chi-squared(6) = 40.66, p less than 0.000002. Section 5.1 corrected: direction is A greater than B on BC and CBMI. Finding is orthogonal to Parisel (2026) vowel-selection model. Conclusion 12 added. v2.4 (2026-06-09): §5.10 added — Currier-preserving null model (n = 200 iterations, size-matched) quantifying each metric's section-discrimination sensitivity independently of dialect. Key result: CBMI is the weakest section discriminant (mean |z| = 1.20 across six sections), confirming that the large CBMI A/B gap (§5.9) is not a section-composition artifact. STTR@100 is the strongest section discriminant (mean |z| = 4.75). Herbal section shows anomalously low vocabulary diversity (STTR z = -13.9); Stars shows anomalously high unique vocabulary (Hapax@500 z = +4.4). Demonstrates two independent organizational layers: CBMI tracks dialect, STTR tracks content domain. Conclusion #13 added. v2.5 (2026-06-10): Corpus expanded from 55 to 71 corpora across 40+ languages. §5.11 adds five medieval European corpora in native script via Universal Dependencies treebanks (Gothic transliteration, Old Church Slavonic, Old East Slavic, Ancient Greek PROIEL and Perseus; VMML 3.54-5.18 — all below alphabetic ceiling of 5.748). §5.12 adds 11 Australian Aboriginal language corpora via BibleNLP/eBible (Pama-Nyungan Western Desert, Ngumpin-Yapa, Arandic; Yolngu; Gunwinyguan; Daly; VMML 6.09-8.00 — predominantly above the Voynich zone). Warlpiri (VMML 5.851) is the sole near-entry on VMML but fails BC (0.233) and CBMI (0.244); 3D normalized distance from Voynich = 0.746 (vs. Irish = 0.200, the nearest neighbor from §5.4). Voynich zone is now charted on both sides: fusional alphabetic below (VMML 3.5-5.75), agglutinative-to-polysynthetic above (VMML 6.0-8.0). Voynich occupies a structural configuration not replicated by any of the 71 corpora tested. To our knowledge, this is the first systematic BPE profiling of Pama-Nyungan languages in the computational linguistics literature. Conclusions #14 and #15 added. v2.6 (2026-06-12): Section 5.10.1 adds a prose-only robustness check for the Section 5.10 Currier-preserving null model. Excluding all label, circular and radial loci (8.7% of tokens), every headline deviation survives essentially unchanged: Herbal STTR z = -13.5, Balneological z = -10.7, Stars Hapax z = +4.3; the sensitivity ranking is unchanged with CBMI last in both conditions. A mean-vs-median distributional note (both summaries rank lexical-diversity metrics first, boundary metrics last) and a coverage note (Astro/Zodiac folios carry no Currier tags and are outside any Currier-preserving design) are added. Erratum: Section 5.10 folio count corrected to 226 parsed / 196 Currier-labeled.
The Korean adnominal ending \texttt{ETM} occurs in diverse noun-modifying constructions, including relative-clause-like modifiers, adjectival and copular forms, bound-noun constructions, and lexicalized expressions. This paper argues that \texttt{ETM} is not a direct marker of relative-clause structure, but a morphological exponent shared by several adnominal constructions. We propose a corpus-based typology that distinguishes these constructions using predicate type, auxiliary structure, argument-structural compatibility, head-noun restriction, and lexicalized patterns. We operationalize the typology as a construction-sensitive annotation layer for the KLUE dependency treebank, implemented through an ordered rule-based procedure and evaluated by manual validation. Productive relative-clause-like uses account for 39.4\% of the analyzed instances; the remainder consists mainly of adjectival, copular, bound-nominal, modal, temporal, and collocational constructions. The findings show that Korean relative-clause-like modification cannot be identified from adnominal morphology alone.
Abstract Transcendental Logic (TL) offers a mathematisation of the unsayable background-reality ( noumenal domain, N-domain) that underlies the logical models ( phenomenal domain, P-domain). According to TL, the N-domain is unsayable not because it is unstructured or inaccessible, but because its structure is governed by the laws of orthomodular logic. In contrast, the P-domain represents reality as a structure of clearly distinct units of meaning; thus, its structure is governed by the laws of classical bivalent logic. This immediately raises two questions: on what basis does TL claim that the N-domain is governed by orthomodular logic, and into what kind of structure the N-domain is organised? The answer lies in TL’s philosophical roots – namely, in Béla Weissmahr’s transcendental metaphysics. His central logical method is dialectic, which governs both the structure of the N-domain and the N-P relation. The structure of the N-domain is analogia entis (the analogy of Being): a dynamic network of analogical-dialectical relations among beings. TL argues that Weissmahr’s dialectic can be expressed in orthomodularity (a dialectical pair is a pair of non-commuting elements), and thus formalises analogia entis as the structure of the N-domain. In Weissmahr’s spirit, TL argues that the dynamic-analogical relations among beings store the information later articulated in the P-domain as structured meaning. The act of this articulation is what TL calls univocalisation. Importantly, TL does not conceive of univocalisation as an act of cognition performed by a particular cogniser. Rather, TL understands univocalisation as the articulation capability of the relational structure itself – independent of any subject or agent. This article does not present the whole of TL, but only its philosophical and algebraic underpinnings. One of its central achievements is the distinction between two interpretations of analogia entis – a structural one and a respective one – and the exposition of their internal relation. In this context, respectivity refers to the algebraic formalisation of the relations among beings as determined by a particular respect. It provides the basis for both relational information storage and univocalisation. The corresponding theorems and proofs are presented in the Appendix. TL thus occupies a new position on the map of logical inquiry. It does not investigate the inferential or linguistic norms of valid reasoning, but rather how meaning is univocalised from the underlying N-domain – understood as analogia entis.
In the contemporary context of the development of the Ukrainian language as a key factor of statehood, national identity, and professional communication, the processes of adapting foreign terminology in the professional language of infrastructure specialties have gained particular relevance. The road construction sector, especially under wartime conditions, remains a strategically important field for Ukraine’s economy and infrastructure, as it ensures mobility, logistics, and national resilience. At the current stage, the terminological lexicon of road construction has been significantly enriched by foreign borrowings. French, German, English, and Slavic terms that entered the Ukrainian language from the 18th to the 21st centuries have become an integral part of professional discourse. These lexical units are incorporated into the national terminological system through phonetic, morphological, and semantic adaptation, which ensures their functional compatibility with Ukrainian linguistic norms. This study provides a detailed description of the road construction terminological system, considering the origin of terms, word-formation patterns, structural features, and methods of integrating borrowings into scientific and technical vocabulary. Particular attention is given to the interaction between native and borrowed elements and the productivity of word-formation models. The analysis demonstrates that foreign terms constitute both historically established and contemporary layers of Ukrainian road construction terminology, contributing to the systematization, internal consistency, productivity, and stability of the terminological system. The results of the research may serve as a theoretical basis for further terminological studies and provide practical guidance for developing linguistic recommendations regarding the standardization, codification, and normalization of terminology in the field of road construction, as well as for other infrastructure-related professional domains.
Static concreteness ratings are widely used in NLP, yet a word's concreteness can shift with context, especially in figurative language such as metaphor, where common concrete nouns can take abstract interpretations. While such shifts are evident from context, it remains unclear how LLMs understand concreteness internally. We conduct a layer-wise and geometric analysis of LLM hidden representations across four model families, examining how models distinguish literal vs figurative uses of the same noun and how concreteness is organized in representation space. We find that LLMs separate literal and figurative usage in early layers, and that mid-to-late layers compress concreteness into a one-dimensional direction that is consistent across models. Finally, we show that this geometric structure is practically useful: a single concreteness direction supports efficient figurative-language classification and enables training-free steering of generation toward more literal or more figurative rewrites.
Cement and concrete account for around 8% of global CO2 emissions and remain among the most hard-to-abate sectors, alongside steel and chemicals. Demand for these materials is projected to grow substantially over the coming decades, particularly across the Global South, driven by rapid urbanization and infrastructure development. While numerous decarbonization technologies and strategies have emerged and are being implemented, the absence of quantitative, context-specific definitions of low-carbon concrete make it difficult to quantify the extent of decarbonization, especially in developing economies. Using India as an illustrative case, where cement production is projected to grow roughly five-fold by 2070, this perspective examines why low-carbon concrete ratings and definitions applied in the developed world cannot be directly implemented in developing-country contexts. It also proposes a phased strategy towards quantitative definitions for the organized and unorganized concrete sectors. This offers a practical pathway to closing the definitional gap that currently limits the efficacy and accounting of cement and concrete sector decarbonization efforts in the global south.
Humans are inherently social beings, and social cues such as faces and voices guide attention and behavior. Auditory perception, especially binaural hearing, is essential for social cognition, enabling sound localization and speech comprehension in noisy environments. Deficits in auditory processing can impair social functioning, and conditions such as social anxiety are linked to reduced social functioning. Since social functioning is closely linked to overall well-being, improving social behavior represents a key objective in psychological research. Virtual reality (VR) is increasingly used to study social behavior due to its flexibility and ecological validity. However, users often report limited social presence, reducing the effectiveness of VR-based interventions especially for social anxiety. One reason may be the dominance of visual over auditory realism: audio is often presented in mono or stereo, reducing naturalness and presence. Binaural auralizations, which provide realistic, externalized spatial audio, may enhance presence and support virtual social interactions. This thesis pursues four main research objectives: identifying suitable behavioral and subjective measures for evaluating binaural realism; assessing immersion, realism, and audio quality across auralization techniques; comparing synthetic and natural speech in a socially stressful VR scenario; and examining effects of binaural audio on affect, presence, and attention under varying social stress levels. Study 1 examined how the virtual visual scene and measurement method affect localization and distance perception of physical sound sources. Across two experiments (N=60), audiovisual incongruence reduced localization accuracy but did not affect presence or realism. Distance estimation was influences by the interaction of task and scene: overestimation increased when using a placement task in a reduced-visibility scene. Study 2 compared localization accuracy for loudspeakers and four virtual audio renderings using a placement task and a gaze-based paradigm (N=49). Binaural renderings produced slightly lower localization accuracy but similar ratings of social presence and realism. A simple generic rendering performed as well as more complex ones. Only the anchor condition lacked externalization and was inferior across measures. Social presence and subjective realism were strongly correlated. Study 3 compared AI-generated text-to-speech with natural human speech in the Trier Social Stress Test (N=40). Both conditions elicited substantial stress responses and produced similar presence and affect ratings, demonstrating the practicality of synthetic speech in virtual social interactions. Study 4 investigated audiovisual realism in a virtual social stress scenario (N=78). A high-stress group showed stronger physiological and subjective stress responses than a low-stress group. Binaural audio increased perceived realism and externalization but did not affect social presence, stress responses, or gaze behavior. High arousal across all groups may have masked audio effects. Across all 4 studies, social anxiety did not consistently affect auditory perception or presence but influenced affective states and subjective evaluations of the interaction. Overall, the findings highlight the importance of VR-specific auditory perception and the role of acoustic immersion. Auditory realism enhances social and physical presence, though its impact varies by context. It appears most effective in low- to moderate-arousal scenarios and may be less critical in highly affective VR applications such as anxiety treatments. Practical advancesn such as TTS integration and simplified binaural rendering methods can support the broader use of realistic audiovisual VR environments in psychological research.
исследование посвящено анализу лингвистических механизмов и прагматических эффектов языковой игры в анимационном дискурсе (на примере мультсериала «Маша и Медведь») в рамках интегративной модели, сочетающей когнитивно-дискурсивный и семиотический подходы. Актуальность работы обусловлена возрастающей ролью аудиовизуальных медиа в формировании языковой личности ребенка и недостаточной разработанностью теоретических оснований описания воздействия ненормативных языковых явлений на когнитивное развитие. Выявлена корреляция между типологическими разновидностями языковой игры (интертекстуальные аллюзии, окказиональное словообразование, трансформация фразеологизмов, нарушение лексической сочетаемости и др.) и активацией метаязыкового анализа, семантической дешифровки и лингвокреативной деятельности реципиента. Доказано, что языковая игра создает лингвокогнитивное пространство, в котором нарушение языковых норм служит механизмом интенсификации рефлексии над языком и его узуальными паттернами. Полученные данные обосновывают рассмотрение языковой игры в детском дискурсе как инструмента формирования языковой компетенции и расширяют теоретические основания изучения когнитивных аспектов речевого воздействия в медиадискурсе и могут быть использованы при разработке методических материалов по развитию речи и лингвистического мышления у детей. the study is devoted to the analysis of linguistic mechanisms and pragmatic effects of language game in animated discourse (using the example of the animated series «Masha and the Bear») within the framework of an integrative model combining cognitive-discursive and semiotic approaches. The relevance of the work is due to the increasing role of audiovisual media in the formation in child`s linguistic personality and the insufficient development of theoretical foundations for describing the impact of non-normative linguistic phenomena on cognitive development. A correlation has been revealed between the typological varieties of language play (intertextual allusions, occasional word formation, transformation of phraseological units, violation of lexical compatibility etc.) and activation of metalanguage analysis, semantic decoding and linguistic creative activity of the recipient. It is proved that language game creates a linguocognitive space, where violation of linguistic norms serves as a mechanism for intensifying reflection on language and its usual patterns. The obtained data substantiate the consideration of language game in children's discourse as a tool for the forming language competence, expand the theoretical foundations for studying the cognitive aspects of speech impact in media discourse, and can be used in the development of methodological materials for the development of speech and linguistic thinking of children.
While language enables meaning, constituting knowledge in courts, schools, or parliaments, who gets to decide what can be known? Is meaning only use or a result of power too? Pitting Wittgenstein's forms of life against Foucault's regimes of discourse makes linguistic norms appear as instruments of exclusion. Marginalised speakers – subaltern, indigenous, and non-normative are often rendered unintelligible. Epistemic justice demands more than inclusion; it demands considering how rules are set, who enforces them, and how meaning is being contextually built. A discourse-sensitive, epistemic theory of justice is proposed, based on Kripke's rule-following paradox and Dijk's discourse analysis, to show that language is not neutral but a battleground of struggle over meaning, recognition, and epistemic authority.
Old English (OE) has traditionally been regarded as an extraordinarily complex language, for its intricate syntactic and morphological relations remain a conundrum for many historical linguists.This explains the substantial number of manuals which primarily focus on the basics of grammar (Mitchell and Robinson 1992; Hog 2012;Fulk 2014).Nevertheless, for those enthusiasts who would like to deepen their insight of Old English structures, Ojanguren López book constitutes a stimulating work which explicitly delves into one of the most elaborate aspects of syntax and semantics: the competition on the complementation of OE aspectual and manipulative verbs, namely, the ones that correspond to contemporary English aspectual end, try and fail, along with manipulative forbid, hinder and refrain.This research also addresses the question of "the rise of serial verb constructions in English" (Ojanguren López 2024, 17) and provides some clarification on the semantics of the verbs linked to complementation patterns, which represents a valuable addition to prior contributions (Callaway 1913;Mitchell 1985;Molencki 1991;Denison 1993;Martín Arista 2022).The aforementioned monograph is divided in nine chapters, which facilitates the comprehension of the matter of study, as it is presented sequentially and includes a well-defined methodology which is supported by the theoretical underpinnings of the research.Thus, chapter one overviews the major focus of the undertaking as well as its analytical structure, and contextualises the complementation patterns of verbs of aspect and manipulation in Old English within its most recent framework.Once the fundamentals have been established, chapter two underlines the relevance of the work that has been carried out considering several dimensions (synchrony, diachrony, typology), which contribute to a comprehensive approach of the objects of research.Furthermore, the lexicographical and textual sources specified in said chapter include dictionaries, lexical databases, grammars, translations, and lemmatisers.In order to describe the competition existing on the complementation of Old English verbs, the Role and Reference Grammar (RRG) and its Interclausal Relations Hierarchy (Van Valin and LaPolla 1997;Van Valin 2005) are proposed, for its "model of the interaction between morphology and syntax, including constituents and operators" is paramount for a language such as Old English (Ojanguren López 2024, 26).Additionally, several works on the clausal complementation of verbs are revisited to discuss different types of competition (Callaway 1913;Denison 1993;Molencki 1991;Los 2005) while introducing nominal complementation and verb serialisation as the innovative elements of the study.As a means to expand the foundations sustaining the analysis on OE verbal complementation, Ojanguren López reviews in chapter three the most relevant parts of the theory of RRG, as the former is central for the consideration of semantic and pragmatic categories in syntax, and for the concern of cross-linguistic validity.The book also devotes special attention to lexical representation, semantic roles and macroroles, which play a crucial role in the association of semantics and syntax, as well as the notion of Aktionsart (Brugmann 1922), for it constitutes "the semantic representation of the sentence" (Ojanguren López 2024, 37).Additionally, grammatical relations embodied in the function of Privileged Syntactic Argument (PSA) and clause structures represented by the Layered Structure of the Clause (LSC), which distinguishes multiple forms of complex referential clauses, complete RRG theory and indubitably help the reader understand the basis for the subsequent study, reinforced by the concepts of juncture and nexus acting as the main aspects of the theory of complex sentences.Chapter four, in turn, intends to become a further contribution towards "a more semantically oriented syntax of Old English" (Ojanguren López 2024, 57), hence it revolves around preceding research in the structural-functional analysis of Old English (Martín Arista 2000a, 2000b, 2020) with the aim of applying these earlier advances to the linking of semantics and syntax proposed by RRG.To successfully accomplish this, terminological remarks such as Verb Classes and Alternations (Levin 1993), linking and its
Social interactions are dynamic and complex, relying on tracking variation in your own and your partner’s affective and mental states. Being “in-sync” neurally and building a shared consensus can mark successful social interactions. In addition, social anxiety can impact social experiences. We used a novel, naturalistic paradigm to investigate relations between inter-brain neural similarity within mentalizing regions, affective similarity and social anxiety symptoms. Undergraduate student friend pairs (N = 34, 85% White, 65% Women) engaged in a social interaction while videos previously captured from their individual perspectives were recorded. Participants watched clips of the social interaction from both their own perspective (their personal view of the social interaction) and their friend’s perspective (their friend’s personal view of the social interaction) while fMRI data were collected. They rated their affect after each clip. Inter-brain neural similarity was computed across three conditions: 1) Same-Stimuli: both participants viewed identical visual stimuli as in a traditional neural similarity paradigm, 2) Self-Perspective: both participants viewed the clip from their own perspective like the originally experienced social interaction and 3) Friend-Perspective: both participants viewed the clip from their friend’s perspective, a novel perspective. Participants self-reported their social anxiety symptoms and affective similarity captured affect rating concordance within dyads. In contrast to the Same-Stimuli condition, when participants viewed the clips like they originally experienced social interaction (Self-Perspective), greater affective similarity was associated with greater inter-brain neural similarity. When participants viewed the clips from a novel perspective (Friend-Perspective), participants lower in social anxiety symptoms exhibited greater inter-brain neural similarity with greater affective similarity; whereas, participants higher in social anxiety symptoms exhibited greater inter-brain neural similarity with less affective similarity. The results suggest stimuli from socially relevant perspectives, rather than identical stimuli, may reveal more nuanced brain-behavior dynamics, allowing for a better understanding of individual differences in socioemotional experience.
Visual perception of built environments contributes to the affective impressions that people form in everyday life. However, how these impressions are represented within vision foundation models remains largely unexplored. To support the systematic investigation of this subject, we introduce the Emotional Impression of Spaces (EMOIS) dataset, comprising 1,544 real-world built-environment images. Each image is annotated with image-evoked valence and arousal ratings collected from Japanese adults by conducting a large-scale web-based survey, with approximately 120 ratings per image. Using Contrastive Language--Image Pre-training (CLIP) representations, we perform predictive and geometric analyses to systematically investigate how valence and arousal are encoded and organized within the representation space. These analyses reveal that valence exhibited stronger and more coherent organization than arousal. Cross-dataset analyses with the Open Affective Standardized Image Set (OASIS), a benchmark dataset of general affective photographs, reveal differences in affective organization between the two datasets. Regression analyses demonstrate high predictive performance for valence and arousal within EMOIS, with mean coefficients of determination of 0.865 and 0.807, respectively, across repeated internal hold-out evaluations. Finally, we present an example-based interface illustrating how learned representations can support qualitative interpretation of predicted affective values. These findings can help elucidate affective representations of built environments and establish EMOIS as a densely annotated resource for future affective computing research in this domain.
The Core Idea Classical combinatorics counts the number of ways to partition a set of n elements into non-empty blocks. That count is the Bell number B(n), and it depends only on how many elements there are, never on how they relate to each other. This manuscript introduces a strict refinement. Given a connected graph G on n labeled vertices, define a connected set-partition as a partition of V(G) in which every block induces a connected subgraph. The geometric partition function is p₍𝔾₎(G) = |{ {V₁, …, Vₖ}: each G[Vᵢ] is connected }| This count depends on G, not just n. Two graphs on the same number of vertices can have wildly different geometric partition counts. The entire theory follows from taking that dependence seriously. Main Theorems Theorem 1 (Subgraph Lattice Theorem). If G and H share the same vertex set and E(G) ⊆ E(H), then p₍𝔾₎(G) ≤ p₍𝔾₎(H). Adding edges can only increase the connected partition count, never decrease it. Theorem 2 (Tree Evaluation). For any tree T on n vertices, p₍𝔾₎(T) = 2ⁿ⁻¹. Every tree on n vertices, regardless of shape, has exactly the same geometric partition count. This is the universal floor for connected graphs. Theorem 3 (Complete-Graph Evaluation). For Kₙ, p₍𝔾₎(Kₙ) = B(n), the n-th Bell number. When every possible edge is present, every set-partition is automatically connected, so the geometric count recovers the classical count. This is the universal ceiling. Theorem 4 (Cycle Closed Form). For the cycle graph Cₙ with n ≥ 3, p₍𝔾₎(Cₙ) = 2ⁿ − n. This matches OEIS sequence A000325. Corollary (Sandwich Hierarchy). For every connected graph G on n vertices with n ≥ 4: 2ⁿ⁻¹ = p₍𝔾₎(Tₙ) < p₍𝔾₎(Cₙ) = 2ⁿ − n < p₍𝔾₎(Kₙ) = B(n) Both inequalities are strict. Trees sit at the floor, complete graphs at the ceiling, and the cycle sits strictly between them. What the Numbers Look Like For cycles versus trees versus complete graphs: n = 5: Tree = 16, Cycle = 27, Complete = 52 n = 6: Tree = 32, Cycle = 58, Complete = 203 n = 7: Tree = 64, Cycle = 121, Complete = 877 n = 8: Tree = 128, Cycle = 248, Complete = 4,140 n = 10: Tree = 512, Cycle = 1,014, Complete = 115,975 The gap between floor and ceiling grows super-exponentially. At n = 10, the complete graph has 227× more connected partitions than a tree on the same vertices. At fixed n, different topologies produce sharply different counts. At n = 6 for example: P₆ = 32 < C₆ = 58 < Ladder = 74 < Fan = 89 < K₂,₄ = 96 < Prism = 114 < W₆ = 118 < K₆ = 203 The ordering tracks graph density with Spearman ρ = +0.955 (p < 10⁻⁶) across all tested sizes n = 4 through n = 10. The Connection Curvature Define the partition connection on cycles as Γₙ = p₍𝔾₎(Cₙ₊₁) / p₍𝔾₎(Cₙ). The curvature κₙ = Γₙ − 2 measures deviation from the flat tree-growth rate: κₙ = (n − 1) / (2ⁿ − n) This is strictly positive and monotonically decaying to zero: n = 3: κ = 0.4000 n = 5: κ = 0.1481 n = 7: κ = 0.0496 n = 10: κ = 0.0089 The cycle approaches tree-like growth exponentially fast. The successive ratio κₙ / κₙ₋₁ converges toward ½ from above. Empirical Validation Program The deposit includes a preregistered validation suite with predeclared falsifiers, frozen metrics, and external ground truth. All outcomes, including failures, are retained as first-class results. Test 1 (OEIS Path Recursion). Brute-force enumeration confirms p₍𝔾₎(Pₙ) = 2ⁿ⁻¹ for n = 1 through 20. The sequence matches OEIS A000079 (powers of 2) exactly. Algebraic DP and direct bitmask enumeration agree at every tested value. Test 2b (Same-n Topology Comparison). At each n from 4 to 10, connected graphs on the same vertex count were compared. Spearman rank correlation between vertex connectivity and p₍𝔾₎ is ρ ≥ 0.83 at every tested n. Pooled edge-density correlation: ρ = 0.955 with p < 10⁻⁶. Negative control (max degree): ρ ≈ 0.35, not significant at any n. The lattice theorem's monotonicity prediction holds empirically, and counterexamples to vertex-connectivity ordering exist only between spanning-subgraph-incomparable pairs, exactly as the theory permits. Test 4 (Internet Autonomous Systems, 2010–2024). Using 180 monthly CAIDA BGP snapshots: AS count grew from 33,778 to 77,495 while edges grew from 93,998 to 501,473. The edge-succession to vertex-succession ratio Sₑ/Sᵥ increased with a linear slope of 3.68 per year, R² = 0.88, p = 2.23 × 10⁻⁷, Spearman ρ = 0.968. The pre-registered threshold was a slope above 0.05. The observed slope is 73× that threshold. Edge growth dominates vertex growth in a real infrastructure graph, consistent with the lattice theorem's prediction that denser graphs have exponentially more connected partitions. Test 5 (OpenAlex Citation Subgraphs). Among 10,000 co-citation subgraphs from 2020–2024: only 3 out of 10,000 were path graphs (0.03%). Among connected subgraphs, paths were 0.47%. For n ≥ 7, zero path graphs appeared among 527 connected subgraphs. Real citation networks avoid the tree floor almost entirely. Test 6/6b (Universal Dependencies Treebanks). The primary test fired its falsifier: English had 19.63% path-topology sentences, exceeding the 10% threshold. However, the pre-registered follow-up (Test 6b) conditioning on n ≥ 5 tokens found all six languages below 10%, with a maximum of 1.57% (Russian). At n ≥ 7 tokens, all languages drop below 0.15%. Short sentences are path-like; longer sentences are not. Test 7 (C. elegans Connectome). 279 neurons, 1,961 synapses. The biological network sits near the tree floor: 81% of 4-vertex connected induced subgraphs achieve p₍𝔾₎ = 2ⁿ⁻¹ exactly. KS tests showed no significant difference from Erdős–Rényi at matched density. Pre-registered hypothesis not supported. The sandwich bound is tight in sparse biological networks. Test 10 (Crystallographic Space Groups). Commutation graphs of all 230 space groups tested against a Ramsey proxy bound. The bound holds for 230/230 groups (100%). Saturation target (≥ 3 groups at n = 5 with zero triangles and zero independent triples) was not achieved. Verdict: partial. Tests 2, 8 (Falsifier Fired). Cycloalkane ring strain shows no monotonic observable matching p₍𝔾₎ (Pearson r = −0.11, p = 0.79). KEGG metabolic pathways achieved 14/18 hits for the Sₑ ≥ 2 × Sᵥ criterion versus a target of 15. Both failures are preserved in the record. Validation Ledger Supported: Tests 2b, 4, 5, 6b (4 of 10 completed tests) Falsifier fired: Tests 1 (by design), 2, 6, 8 (4 of 10) Not supported: Test 7 (1 of 10) Partial: Test 10 (1 of 10) Blocked/Deferred: Tests 9, 11 Negative and null results are first-class outcomes, not omissions. The program-level verdict is assessed across the full ledger, not cherry-picked from successes. Limitations Some tests depend on external APIs (OpenAlex, CAIDA, KEGG) whose upstream data evolves. Test 9 is blocked by endpoint observability constraints. This archive captures a dated snapshot; future reruns under changed data conditions should be treated as independent replications, not as invalidations. Citation Cite the Zenodo DOI for this archived package and the manuscript title. If reusing specific test scripts or results, cite the relevant script and JSON artifact in your methods section.
Part of speech and syntactically annotated dataset for modern Mongolian. The dataset is a fully annotated corpus of modern Mongolian texts written in Mongolian Cyrillic. Version 2.
The Prague Dependency Treebank framework is unique in its attempt to systematically include and link different layers of language, including a meaning representation with several types of inter-sentential phenomena, especially coreference and discourse relations. We present its second consolidated version (PDT-C 2.0), which concludes almost 30-years long project of sustained development of the resource to a uniformly and coherently annotated, genre-diversified, almost 4 million token language resource of Czech language, with accompanying fully compatible lexicons. In addition to continuous linguistic research, the richly linguistically annotated corpus is also widely used in international comparisons of the development of traditional and novel NLP tools as well as in conversions into other formalisms. The corpus and the trained parsers are available under the CC BY-NC-SA licence.
Studies of emotion often rely on standardized stimulus sets to elicit affective responses. Although established databases provide images with normative valence and arousal ratings, selecting suitable stimuli can be difficult when experiments require specific thematic or content constraints. This challenge is especially pronounced for negative stimuli, which are central to research on maladaptive emotions and behaviors in clinical contexts but are often scarce in necessary quantity or specificity. The present study evaluated the feasibility of using generative AI, specifically text-to-image generators, to create tailored negative and neutral affective stimuli. To assess whether these images can serve as alternatives to traditional stimuli, we compared their affective properties to those reported in standardized image databases. Across two studies, participants rated the valence and arousal of 160 and 200 AI-generated images. Our findings revealed that AI-generated negative and neutral images reproduced the characteristic inverse association between valence and arousal observed in standardized databases, with moderate to strong correlations between these dimensions. These results highlight the potential of generative AI as a practical methodological tool for creating customized affective stimuli aligned with specific research objectives and experimental designs.
Abstract Color-evasive language practices and ideologies are central to the creation of clinical spaces as white public spaces, which are inherently dangerous for non-white people. This chapter explores how race, language, and disability are shaped by the color-evasive logic taught to medical students through linguistic practices within autism diagnostic processes. In these color-evasive practices, culture can index race in a socially acceptable way, as it allows individuals to talk about race without talking about race. In this clinical vignette, a Black child’s use of Black linguistic norms is described by the white clinician as being “inappropriate, disruptive, and echolalic.” This is reinforced when the clinician tells his students that scoring an autism diagnostic test is “not about culture, it’s about the response.” The chapter explores how the scoring of the diagnostic test is grounded in the subjective assessment of the clinician administering the test and how they understand the patient’s behavior, which is shaped by underlying anti-Black logics about race, language, and disability. The clinician’s response shows us that despite his production of a color-evasive framework, he perceives and hears the patient as a racialized other. The chapter argues that color-evasive ideologies create white public space where clinicians police non-white bodies and conflate a Black child’s use of Black language with disability.
This dataset contains the coded lexical and contextual data used in a corpus-informed analysis of emotion-related vocabulary in Italian as a foreign language (IFL) textbooks at the beginner level (CEFR A1) used in Polish lower secondary education. The dataset is based on four textbooks from two series: Progetto Italiano Junior (Marin, 2017; 2018) and Va bene! (Kaliska & Kostecka-Szewc, 2021). All materials were analysed in their printed form. The dataset includes all lexical items identified as emotion-related based on their presence in the ANEW-IT database (Montefinese et al., 2014), which provides normative ratings of valence and arousal for Italian words. Each entry in the dataset corresponds to a single token occurrence of an emotion-related lexical item. The dataset includes both surface forms as they appear in the textbooks and their corresponding base forms (lemmas) as listed in ANEW-IT. For each token, the dataset provides contextual, linguistic, and affective information. The variables included are: • textbook and series identification • unit/chapter, page number, and exercise reference • material type (e.g., dialogue, reading text, exercise, review) • pedagogical focus (e.g., grammar, vocabulary, comprehension, mixed) • word form (surface form) and lemma • part of speech (POS) • contextual sentence or description of occurrence • frequency measures (FreqColfis, Ln_Colfis) • affective ratings (valence and arousal, scale 1–9) The dataset enables replication of the quantitative analyses reported in the study, including token counts, type–token ratios, valence and arousal distributions, and comparisons across textbook series and pedagogical contexts.
Abstract This paper examines the social contexts in Thomas Mann’s novel Buddenbrooks where the North German dialect Plattdeutsch is spoken. Beyond the technical challenges of translating these passages, the analysis focuses on the literary representation of code-switching that functions primarily as socially and symbolically charged act. Drawing on Pierre Bourdieu’s sociological theory, the study interprets deviations from linguistic norms and dialect use as instances of double negation – a strategy that appears to challenge social conventions but, in reality, affirms the most valuable social capital in classical bourgeois society: the certainty that one’s high status remains unthreatened. The difficulty of translating such passages stems from the specific cultural parameters embedded in the novel. Ultimately, the paper argues that culture is not an immediate given but requires analytical frameworks from the social sciences for proper understanding and interpretation.
Implicit discourse relation recognition (IDRR) addresses the classification of discourse relations between text segments without explicit connectives. Existing prompt-based methods for IDRR often rely heavily on predicting surface connectives as an indicator for the discourse relation, which is inherently limited by the capacity of pre-trained language models. Meanwhile, standard attention mechanisms in these models are easily distracted by task-irrelevant tokens. This paper proposes a Role-Focused Prompt Framework that addresses these limitations by introducing a role-centric perspective to IDRR. Our approach is built on two core innovations: (1) the incorporation of linguistically grounded semantic roles (e.g., Cause/Effect for Contingency relation) into IDRR, which directly captures the underlying argument structure that determines discourse relations, reducing reliance on connectives; (2) a focused prompt structure that condenses the input to its core semantic concepts (argument summaries, connective, and semantic roles), creating a high signal-to-noise environment for attention-based reasoning. Extensive experiments on Penn Discourse TreeBank 2.0 (PDTB 2.0) demonstrate that our framework achieves competitive results, providing complementary direction for IDRR research. Ablation studies validate that both innovations are essential to the framework. Our work demonstrates that incorporating linguistically grounded semantic roles and focusing on task-relevant concepts can effectively specialize pre-trained models for IDRR.