Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
18265 papers
Abstract While much work on language variation and change has stressed the role of social context, relatively little attention has been paid to the ways in which particular socio-cultural practices may guide processes of language change in locally or regionally variable ways. This paper explores the role of specialized discourse forms – such as shamanic incantation, ceremonial dialog, and other forms of verbal art – as loci for the emergence and propagation of linguistic innovations in the Amazon basin, particularly those associated with language contact. Specialized discourse in this region arguably enables a potent recipe for language change, via an emphasis on extensive circulation across speakers, communities, and languages; a particular value ascribed to dispersed and linguistically distinct forms; discursive norms that favor close replication while also licensing creative manipulation; and the social position of specialists themselves, who tend to bring together both diffuse social networks and relative status. Various examples of lexical and grammatical change in Amazonian languages are identified that have plausible ties to specialized discourse.
The following study explores the extent to which American English has influenced the phonology of Singapore English, with a specific look at the TRAP-BATH [æ]/[ɑː] vowels. While Singapore has traditionally regarded British English as the norm due to past historical colonial ties, increasing American media exposure raises the possibility of emerging and new phonological change. This study explores this possibility in Singaporean university students, examining their vowel productions with lexical items from the BATH lexical set. The author also hypothesises that American media influence would in fact cause the adoption of the [æ] vowel for lexical items in the BATH set. Six Singaporean university students participated in a reading task and elicitation task to capture their pronunciation of certain BATH lexical items. Results show an overwhelming preference for the British /ɑ/ vowel, with limited productions of the American /æ/ vowel. Productions of the /æ/ vowel appear to be word specific and did not reflect a larger systemic shift of Singapore English adopting American English features consistently. No significant differences were observed across gender and ethnic groups. Overall, this paper finds that American English influence may be present, but remains limited and is not reflective of a broader change in the vowel production of Singapore English. This study contributes to ongoing discussions on Singapore English and how it continues to navigate between established linguistic norms.
This paper revisits and extends Greenberg's Universal 45 on gender distinctions using Universal Dependencies 2.17, comprising 339 treebanks across 186 languages.A systematic analysis of morphosyntactic patterns confirms the implicational hierarchy (singular > plural gender marking), with 98.6 % conformity in pronominal categories.Only two potential exceptions are detected, both with minimal occurrences and likely attributable to annotation errors.Extending the analysis beyond pronouns to 13 UPOS categories shows that core categories maintain near-perfect compliance, while peripheral categories exhibit higher violation rates, primarily driven by annotation inconsistencies rather than genuine linguistic exceptions.A total of 90 treebanks display gender-number features in traditionally invariable categories (e.g., adpositions, conjunctions, adverbs), indicating annotation issues such as prepositional contraction handling, homophone merging, and erroneous feature assignment.The study establishes a replicable computational methodology for large-scale typological validation, highlighting both the potential of corpus-based approaches and key limitations, including genealogical sampling biases, annotation heterogeneity despite universal schemas, and the false sense of comparability across treebanks.
This describes how to calculate the reliability of word ratings (or other performance data, such as RTs). It gives R code and an example of how to use it.
Previous studies have shown that emotional stimuli, such as pictures and sounds, can affect the perception of time. The present study investigated the effects of arousal and valence of standardized emotional pictures on two different subjective experiences of time: duration and the passage of time (POT). Emotional pictures were presented for 2, 4, or 6 s, and after each presentation participants provided either a duration or a POT judgment (in different experimental halves) as well as valence and arousal ratings on visual analog scales. In contrast to previous findings, the results showed no effects of arousal and valence on duration judgments; however, there were significant effects of arousal and valence on POT judgments, with both higher arousal and valence leading to an accelerated POT. The absence of arousal and valence effects on duration judgments in the present study could be related to the fact that pictures with extreme content were not included and thus the entire spectrum of possible arousal and valence values was not tested. Importantly, however, the different patterns of result for duration and POT judgments align with a growing body of research indicating that the perception of duration and the experience of the POT are not necessarily directly linked. Furthermore, the present study highlights the importance of collecting and analyzing current affective ratings even when using standardized emotional pictures, as the current ratings differed significantly from the norm values of the standardized picture sets.
Digital communication has profoundly transformed the nature of textuality in both English and Azerbaijani contexts, affecting language structure, stylistic conventions, and the interactional norms of written discourse. The evolution of online platforms—including social media, messaging applications, and digital forums—has encouraged brevity, immediacy, and multimodality in textual expression. English, as a globally dominant language in digital spaces, exhibits widespread use of abbreviations, emojis, hashtags, and memes, reflecting both linguistic economy and the need to convey affective meaning rapidly. In contrast, Azerbaijani, while increasingly influenced by these global digital trends, maintains certain structural and grammatical particularities, including agglutinative morphology and context-dependent sentence constructions, which interact uniquely with digital shorthand and visual markers of communication. Comparatively, English digital texts often prioritize directness and global comprehensibility, leveraging standardized syntax and widespread conventions that facilitate cross-cultural interaction. Azerbaijani digital communication, while adapting similar strategies, demonstrates a more hybrid textuality: users negotiate between traditional linguistic norms and the pressures of digital brevity, producing hybrid forms that mix Latin-based abbreviations, code-switching, and local syntactic patterns. These patterns reveal not only the flexibility of the Azerbaijani language in accommodating technological mediation but also highlight the interplay between digital literacy, cultural identity, and language preservation.
This paper focuses on annotating the Greek Learner Corpus in Universal Dependencies (UD).It presents the annotation process, development of guidelines and evaluation of the attempted annotation of two annotators.This work is part of a larger annotation project which aims to compile a sizeable learner treebank that can be used to promote research on second language acquisition and its automatic processing.
Materials The selected pictures (248 in total: 120 neutral and 128 negative) were drawn from five different standardized affective picture sets (see Table 1 for details on which images were selected from each set). Participants were recruited from Utrecht University via campus advertisements. The Dutch sample consisted of 103 participants (90 females), all native Dutch speakers, with a mean age of 20.93 years (SD = 2.42). Their responses provide the arousal and valence ratings with these 248 images. Task Design This is a emotion rating test conducted online via the Gorilla platform. Each trial began with a 1500-ms fixation, followed by a 2000-ms presentation of the target image. After a 1000-ms delay, the SAM scales appeared and remained on screen until the participant responded, with arousal rated first and valence second (Figure 1). DATA · The folder titled “Open_Data_Dutch_Norm_Data” contains an Excel files named “V1_Dutch_norm_data_arousal_pp103” and “V1_Dutch_norm_data_valence_pp103”, which includes raw data on arousal and valence, along with their response times during the rating task, respectively. · The same columns in both files are: Subject, Gender, Image, Scale, Picture_Valence, Response, Response_Duration. · Within-subject factor: Picture_Valence (negative, neutral)
This study examines the participation experiences of multilingual students speaking English as an additional language (EAL) and their decisions to invest or disinvest in language practices during classroom discussions. Using an embedded multiple case study design, data were collected through classroom observations, weekly reflections, and individual interviews, and analyzed using thematic analysis. The findings show that course characteristics, such as structure and topic, influenced power dynamics and identity negotiation, reinforcing sociocultural and linguistic norms aligned with Western participation practices. The complex interplay of course dynamics, norms, and broader ideologies contributed to multilingual EAL students’ novice identity, creating barriers that made it challenging for them to disrupt participatory norms and invest in the language practices of the course community. This article underscores the importance of reframing participation as a collaborative and critical process and highlights the need to create inclusive classroom environments to support the equitable participation of multilingual EAL students.
Emotion recognition plays a crucial role in human–computer interaction, health monitoring, and affective computing by analysing physiological signals. Despite recent advancements, current research still faces challenges, including the lack of effective fusion strategies for diverse physiological modalities, difficulties in handling high-dimensional feature representations, and limited use of efficient temporal modelling techniques to capture complex emotional patterns. This study proposes a deep learning-based approach that fuses multiple physiological modalities, including Electroencephalography (EEG), Electrooculography (EOG), Electromyography (EMG), Galvanic Skin Response (GSR), Respiratory Rate (RR), Skin Temperature (SKT), and Photoplethysmography (PPG), to improve emotion recognition. Arousal and valence ratings were binarized into two classes (low/high) using a threshold of 4.5, formulating a binary classification problem. In addition to utilising Bidirectional Long Short-Term Memory (Bi-LSTM), the study employs Temporal Convolutional Networks (TCN), a widely used approach for time-series analysis, to efficiently capture temporal dependencies. The proposed model optimises feature selection through channel-wise strategies, incorporates advanced learning rate scheduling, and reduces computational overhead. Furthermore, window-wise, block-wise, and trial-wise evaluation protocols were investigated to assess the impact of temporal information leakage on emotion recognition performance. Using the DEAP dataset for validation, the proposed TCN-based approach achieved classification accuracies of 88.42% for valence and 86.35% for arousal under an overlapping block-wise evaluation protocol, demonstrating improved performance in binary emotion recognition and highlighting the importance of leakage-aware model assessment.
Cet article examine la représentation du mariage mixte et les enjeux de la nomination des enfants dans le roman À la frontière de nos deux mondes: La Déchirure de Leila Miloud Ropp, inscrit dans le contexte de l’Algérie coloniale, de l’après-Seconde Guerre mondiale aux prémices de la guerre d’indépendance. À travers l’histoire d’Alice, Française venue des Vosges, et de Mécheri, Algérien musulman de Mostaganem, l’œuvre construit un espace romanesque où l’intime se trouve continuellement traversé par des rapports de domination, des logiques d’appartenance et des fractures historiques. Notre analyse montre que le mariage mixte fonctionne comme un dispositif narratif de confrontation entre deux systèmes de normes; familiales, religieuses, politiques et que l’appellation des enfants issus de l’union constitue un acte social chargé d’idéologie: nommer, c’est assigner une place, tracer une frontière ou tenter une conciliation. En mobilisant une approche croisée analyse littéraire, stylistique et socio-linguistique, nous mettons en évidence les mécanismes discursifs par lesquels le roman fait émerger la « déchirure » identitaire, cristallisée dans les choix lexicaux, les voix rapportées et les scènes de sociabilité. Abstract This article examines the representation of mixed marriage and the issues surrounding the naming of children in Leila Miloud Ropp's novel À la frontière de nos deux mondes: La Déchirure (At the Border of Our Two Worlds: The Tear), set in colonial Algeria from the aftermath of World War II to the beginning of the war of independence. Through the story of Alice, a French woman from the Vosges, and Mécheri, an Algerian Muslim from Mostaganem, the work constructs a fictional space where intimacy is continually disrupted by relationships of domination, logics of belonging, and historical divisions. Our analysis shows that mixed marriage functions as a narrative device for the confrontation between two systems of norms—familial, religious, and political—and that the naming of children born of the union is a social act laden with ideology: to name is to assign a place, to draw a boundary, or to attempt reconciliation. Using a cross-disciplinary approach combining literary, stylistic, and sociolinguistic analysis, we highlight the discursive mechanisms through which the novel brings out the “tear” in identity, crystallized in lexical choices, reported voices, and scenes of sociability
This article offers a close study in using TEI to represent the distinctive forms of multilingual glossing found in premodern texts, including where existing norms for encoding practice may not always leave room for the specificities of medieval practice. Drawing on the data set from a new edition of Walter de Bibbesworth’s Tretiz (a thirteenth-century rhymed vocabulary of French, extensively glossed into Middle English across its many surviving manuscripts), it outlines how <term> and <gloss> elements might usefully be deployed to illustrate the relations between lexical items both on the manuscript page and across a broader manuscript tradition.
The curation of Armenian medical vocabulary from Late Antiquity to the early modern period reflects an intricate interplay between lexical borrowing and native word formation. Medical terminology entered Armenian mainly through Greek, Arabic and Persian. Each influenced scribal decisions based on factors such as semantic precision, syntactical compatibility and cultural relevance. Direct borrowings were often transliterated in accordance with phonological norms. Marginal glosses and commentary added pedagogical clarity. Medical concepts such as pharmacological compounds, anatomical parts and botany illustrate the linguistic continuity in the Armenian vocabulary. The resulting lexicon was neither passive adoption nor bulk translation, but instead a carefully synthesised glossary shaped by ideological and philological pressures. The Armenian medical language thus reveals the exigencies of scribes in mediating scientific knowledge, serving as a valuable case study in cross-cultural medical transmission and historical medical linguistics.
Digital mental health is undergoing rapid transformation through the deployment of artificial intelligence (AI) systems including conversational agents, risk-prediction tools, and clinical decision-support platforms. Despite growing enthusiasm, the field has yet to converge on an operational standard for what constitutes honest, trustworthy AI in psychiatric and psychological contexts-one that explicitly accounts for the limits, blind spots, and populations for whom deployed systems have not been validated. This Perspective draws on Richard Feynman's foundational principle of scientific integrity-articulated in his 1,974 Caltech commencement address, "Cargo Cult Science"-and Carl Sagan's dimensional metaphors to develop a practical, disclosure-oriented ethical framework for AI in digital mental health. The neurodiversity paradigm provides a critical stress test: growing empirical evidence indicates that AI tools trained on narrow behavioural and linguistic norms carry substantial risk of misrepresenting neurodivergent users, with the potential to silently re-encode stigma, misclassification, and inequity into clinical systems. Drawing on predictive processing theory and recent empirical evidence of AI bias against neurodivergent populations, this work proposes five criteria for the Feynman Honesty Standard: scope clarity, population-limit disclosure, uncertainty communication, role integrity, and participatory co-development. The analysis concludes with a forward-looking research agenda for clinicians, developers, and policymakers, arguing that the field will be judged not only by algorithmic performance but by its candour about what these systems know, miss, and cannot do.
OBJECTIVE: To investigate how resonance manipulation in Chinese vocal performance causally affects the perceived emotional valence, using an integrated two-phase approach that combines acoustic characterization with perceptual evaluation. METHODS: In the Stimulus Selection Phase, 120 native Chinese-speaking participants rated the emotional valence of 300 Chinese music excerpts (children's songs, pop, folk, operas, religious music, and modern folk-pop) on a 7-point Likert scale. Excerpts were temporally degraded to isolate melodic content. Valence scores were binned into 7 levels, and 10 excerpts per level (70 total) were re-recorded by two professional female vocalists in "full" and "thin" resonance conditions. In the perceptual evaluation phase, 80 additional participants rated the emotional valence of these recorded excerpts (280 trials) using a Latin square design. Acoustic parameters, i.e., H1-H2, mean formant bandwidth (MFB), singing power ratio (SPR), and harmonics-to-noise ratio (HNR), were analyzed, and linear regression models assessed the effects of valence level, resonance condition and their interaction on emotional ratings. RESULTS: Acoustic analyses confirmed significant differences between resonance conditions, with "full" resonance showing higher H1-H2, HNR, MFB and SPR, indicating distinct timbral qualities. Perceptual ratings revealed that valence level and resonance condition both significantly predicted emotional ratings (p < 0.001). "Full" resonance enhanced positive valence ratings and "thin" resonance reduced positive valence ratings across levels, particularly for emotionally ambiguous stimuli. CONCLUSION: Vocal resonance manipulation is a potent causal factor in shaping emotional perception in singing. The findings bridge the gap between pedagogical concepts and measurable acoustics, showing how the "full-thin" resonance continuum serves as a nuanced tool for emotional communication within the Chinese vocal tradition. These results indicate the importance of considering both acoustic implementation and perceptual outcome in understanding timbre-emotion relationships.
Abstract Music enjoyment and familiarity are closely related but often confounded in studies of neural music processing. Here, we investigated their distinct contributions to cortical oscillatory activity and neural tracking of music using electroencephalography (EEG). Thirty-two participants listened to self-selected all-time favourite songs, recent favourite songs, and tempo-matched songs from disliked genres. This novel paradigm dissociated familiarity from enjoyment by including highly enjoyed songs that differed in familiarity. Spectral power and cortical tracking (using Mutual Information) were analysed using linear mixed-effects models with enjoyment and familiarity ratings. Familiarity was associated with increased left-frontal alpha power, whereas enjoyment predicted increased theta and beta power, demonstrating distinct oscillatory signatures for these dimensions. An interaction revealed that the positive relationship between enjoyment and theta power was strongest for highly familiar music. Cortical tracking analyses showed that greater enjoyment was associated with reduced delta-band tracking, with a significant interaction indicating that this negative relationship was present for highly familiar songs but not for less familiar songs. These findings indicate that enjoyment and familiarity differentially shape neural responses to music and highlight the importance of modelling both factors to disentangle their distinct effects on neural activity during music listening.
of the norm,
Language is a crucial instrument of society, since communication ensures the effective functioning and development of social structures.Therefore, the relationship between language and society has become one of the key areas of sociolinguistic research.Another important aspect of social interaction is gender, which significantly influences communication patterns, linguistic behavior, and the formation of cultural norms.The article examines gender features in the context of modern media discourse and analyzes the relationship between language, gender, and mass media.Particular attention is paid to the representation of gender stereotypes in English-language media discourse and to the ways they are reproduced through journalism, advertising, television, cinema, and digital media.The study highlights the role of gender as a sociocultural factor that shapes linguistic practices, communication strategies, and public perceptions of masculinity and femininity.The paper explores the development of gender-inclusive language and analyzes the transformation of linguistic norms under the influence of feminist and sociolinguistic approaches.It is emphasized that media not only reflect social tendencies but also actively participate in constructing gender roles and behavioral models.The study demonstrates that media discourse often reproduces stereotypical portrayals of women and men, where men are associated with authority and activity, while women are more frequently connected with emotional or domestic roles.The article also investigates gender differences in communication styles, including variations in politeness, formality, lexical choice, and sentence structure.Special attention is devoted to advertising discourse and linguistic strategies aimed at different gender groups.The research confirms that gender stereotypes in media discourse influence public consciousness, cultural values, and social behavior.
This paper explores the systemic structure of the Nepali language through the lens of Bloomfield’s linguistic norms, grounded in the general characteristics of structural linguistics. It aims to explore how far Bloomfieldian norms are applicable to the way the Nepali language functions. This study is based on secondary sources, including books and articles available in the library and in open-access databases. The major findings of the research reveal that these structural principles are not confined to the English language alone; rather, every language possesses its own systematic sentence structure. Accordingly, Nepali also exhibits constituent relationships that conform to its own structural organization. This suggests that Bloomfield's concepts of endocentric and exocentric constructions are also applicable to Nepali sentence structures. In both English and Nepali, meaningful relationships exist between preceding and following constituents. These constituent relationships illustrate the hierarchical organization of sentence structure and the functional interdependence of its elements.
The emergence of a distinct Gen-Z sociolect, often termed "Genzie" or "Internet Slang," represents one of the most rapid and transformative linguistic developments of the digital age. This language is not a random collection of slang but a complex, rule-governed system born from the intersection of technology, social change, and identity formation. A comprehensive, data-driven analysis of its historical evolution, structural properties, and socio-pragmatic functions is critical to understanding contemporary communication, as it reflects fundamental shifts in how a generation conceptualizes interaction, community, and self-expression. This study aims to deconstruct the Gen-Z sociolect by tracing its historical development over a key 36-month period (2021-2023), analyzing its core structural components (lexical, semantic, syntactic, multimodal), and explaining its social functions within digital communities. The research seeks to move beyond anecdotal description to provide a rigorous, empirical account of this dynamic linguistic phenomenon, thereby establishing a benchmark for the academic study of internet-native dialects. We position this sociolect not as a degradation of Standard English, but as a legitimate linguistic innovation worthy of serious scholarly attention, with its own internal logic and systemic coherence. A mixed-methods, diachronic approach was employed, integrating the scale of computational linguistics with the nuance of qualitative discourse analysis. A large-scale corpus of approximately 6000 posts was compiled from three core platforms—Twitter/X, Instagram, and TikTok—across the 12-month timeframe, ensuring a representative sample of public-facing Gen-Z communication. Computational linguistics methods were used for quantitative analysis, including time-series modeling for lexical diffusion, diachronic word embeddings for semantic shift, and supervised machine learning for stylometric identification. This was complemented by qualitative discourse and pragmatic analysis of a stratified sample of posts to understand language-in-use, focusing on the interplay between text, image, and platform-specific conventions. The analysis reveals a clear, platform-influenced historical trajectory for Gen-Z language, with terms originating on niche, visually-driven forums like TikTok and Twitch before achieving mass diffusion on the text-centric environment of Twitter and finally being normalized on the broader social canvas of Instagram. We identified and modeled three primary mechanisms of lexical creation: neologism (e.g., "skibidi," "gyatt"), semantic reappropriation (e.g., "cap," "based," "fire"), and phono-semantic matching from online cultures (e.g., "ratio," "L + RIP bozo"). Gen-Z language is a legitimate and sophisticated dialect of the digital era, a natural linguistic adaptation to a hyper-connected, attention-economy-driven world. Its evolution is not chaotic but follows predictable patterns of cultural transmission that are dramatically amplified and accelerated by social media algorithms. Its structure efficiently manages cognitive load in fast-paced digital environments while its primary functions are the performance of a specific digital identity, the creation and policing of digital community boundaries, and a form of resistance to traditional linguistic and social norms.
This study explores the pragmatic typology of the functional-semantic field (FSF) of degree in English and Uzbek, focusing on how gradability, intensity, and comparison are expressed and interpreted across two typologically different languages. The concept of degree is treated as a universal semantic category realized through a range of linguistic means, including morphological forms, lexical items, and syntactic constructions. The research aims to identify both common patterns and language-specific features in the expression of degree, as well as to analyze the role of pragmatic factors in shaping its meaning. The findings demonstrate that English primarily relies on analytic and morphological devices, such as comparative and superlative forms and intensifiers, while Uzbek employs agglutinative mechanisms, lexical markers, and expressive forms such as reduplication. Despite these structural differences, both languages share a common semantic core based on scalarity and gradation. However, the interpretation of degree is highly context-dependent and influenced by speaker intention, discourse context, and cultural norms. The study also shows that degree expressions serve not only as markers of quantitative or qualitative comparison but also as pragmatic tools for expressing evaluation, emphasis, politeness, and implicature. The functional-semantic field of degree is organized into core and peripheral zones, where core elements provide basic gradation and peripheral elements introduce stylistic and contextual variation. Uzbek demonstrates a stronger tendency toward expressive and emphatic usage, while English often relies on more implicit and context-driven strategies. In conclusion, the research highlights the importance of integrating semantic and pragmatic approaches in the study of degree and contributes to a deeper understanding of cross-linguistic variation in functional-semantic categories. The results have practical implications for language teaching, translation, and intercultural communication.
ABSTRACT Emotional episodes involve partially coordinated changes in experience and autonomic activity, and subjective–physiological coherence appears to vary with aging and physiological stress. Physiological accounts of emotional aging propose that age‐related changes in autonomic and interoceptive function reduce the informativeness of bodily signals, weakening the coupling between feelings and bodily responses. We test whether subjective–physiological coherence is lower in older vs. younger normoxic adults, lower in hypoxic vs. normoxic younger adults, and similar in hypoxic younger adults and older normoxic adults; in secondary mechanism‐oriented analyses, we examine whether heart‐rate variability and cardiac interoceptive accuracy are statistically related to group differences in coherence. We will test these questions in three groups ( N = 120): younger adults in normoxia, younger adults under acute hypoxia (FiO 2 ≈ 12%), and older adults in normoxia. Participants will view emotional images and provide trial‐wise ratings of affect and approach motivation while heart rate and skin conductance are recorded. Emotional coherence will be indexed primarily by within‐person correlations between subjective arousal ratings and autonomic reactivity, with approach motivation and valence examined as complementary subjective dimensions.
Low-resource languages face a critical challenge in AI development: creating specialized conversational systems without access to massive training corpora. We present a systematic methodology for transforming structured linguistic resources into specialized AI systems, demonstrating that expert-curated lexical databases can serve as effective foundations for conversational AI development. Our approach converts Hindi WordNet into 1.25 million diverse instruction-response pairs, fine-tunes a 12B-parameter language model using resource-efficient LoRA with 4-bit quantization. Evaluation through a Hindi language learning chatbot demonstrates that structured-knowledge-based systems achieve superior pedagogical effectiveness (91.0 vs. 79.4-83.6 for general-purpose models) while maintaining competitive semantic performance and exceptional consistency. The complete pipeline demonstrates a proof-of-concept methodology using Hindi for developing specialized AI systems for any languages with WordNet resources. This work addresses the critical gap in AI accessibility for low-resource languages, offering a practical alternative to corpus-intensive approaches and potentially enabling specialized AI development for the hundreds of languages with existing WordNet resources.
We present MorfFlex, a morphological dictionary architecture suitable for languages with extensive regularity in both inflection and derivation. As the primary example of MorfFlex in use we introduce MorfFlex CZ, a morphological dictionary of Czech. It is distributed as a simple, unstructured list of <wordform, lemma, tag> triplets, however, its manually maintained, unpublished source files and conversion scripts encode a sophisticated system of inflectional and derivational patterns. These patterns dramatically reduce the otherwise enormous size of the dictionary, which currently contains over 100 million wordforms and more than 1 million lemmas. The MorfFlex CZ dictionary serves as an essential resource for ensuring the consistency of manual morphological annotation in the Prague Dependency Treebanks and underpins state-of-the-art automatic tools such as MorphoDiTa. In this paper, we focus on: (i) presenting an effective method for managing the rich morphological system within the dictionary, and (ii) demonstrating the utility of such a language resource for maintaining annotation consistency in corpora and supporting the development of advanced NLP applications.
Tooltips are among the most information-dense yet spatially constrained elements in software user interfaces. Their effective translation is far from a trivial problem. A tooltip must convey precise technical meaning, respect the lexical norms of the target language, fit within a character budget imposed by layout constraints, and remain coherent with surrounding interface text, all simultaneously. Conventional translation pipelines that treat each string in isolation consistently fail to satisfy these requirements together. This paper proposes a context-aware tooltip translation model that addresses the deficiencies of isolated string translation. The approach integrates sentence-level and document-level context into the translation decision, enforces terminology consistency through a dynamic glossary module, and applies a length-aware decoding strategy that respects display constraints without sacrificing semantic fidelity. A structured survey of methods published between 2018 and 2025 is presented, covering neural machine translation, context-sensitive models, terminology-constrained decoding, multimodal localization, and low-resource adaptation techniques. Reported BLEU score improvements over baseline isolated-string systems range from 4 to 18 points across evaluated language pairs, with the largest gains observed in morphologically rich and low-resource languages. The proposed framework is evaluated across six language pairs and three software domains, demonstrating consistent improvement over isolated-string baselines on both automatic metrics and human preference judgments. Significant challenges remain, particularly around handling domain-specific abbreviations, maintaining consistency at scale across large software products, and adapting to low-resource target languages with limited parallel UI data. This survey maps the current state of the field, identifies where existing methods fall short, and charts the directions most likely to yield practical progress.
This article presents a comparative analysis of the linguistic, stylistic, and structural features of official document texts in the Uzbek and Turkish languages. The study examines state documents, official correspondence, directives, orders, contracts, applications, and decisions as illustrative materials. The terminological systems, syntactic structures, lexical units, and stylistic norms of document texts in both languages are identified, and their similarities and differences are analyzed on a scientific basis. The results reveal the common Turkic roots of Uzbek and Turkish document language, contemporary trends in the development of official style, and factors related to national language policy.
This article examines the concept of occasional words (occasionalisms) and their linguistic features within modern discourse. It explores their definitions, structural characteristics, and functional roles from morphological, semantic, pragmatic, and stylistic perspectives. The study highlights the context-dependent nature of occasionalisms, emphasizing their expressive and evaluative functions as products of individual linguistic creativity. Special attention is given to their formation through analogy, their role in literary and journalistic texts, and their position between language norm and speech innovation. The article also differentiates occasionalisms from neologisms and nonce formations, focusing on their degree of lexicalization and communicative purpose. The findings demonstrate that occasional words are not random deviations but systematic, rule-governed innovations that contribute to linguistic development and reflect the dynamic interaction between language structure and usage.
In this study, it is examined patterns of code-switching, adoption of slang, lexical innovations, and multimodal elements in captions, comments, and hashtags among young Uzbek users on Instagram. Quantitative surveys of multilingual young adults in Kokand, Uzbekistan, are combined with qualitative analysis of public posts from influencers in the lifestyle, education, and marketing domains as part of a mixed-methods design. The results show that Uzbek, English, and Russian are frequently combined with platform-driven semantic changes (e.g., redefined terms for followers, content, and private messages) and informal syntactic patterns driven by visual-text interaction. Mixed tags, emojis, and acronyms are often used by users to adapt global trends, resulting in hybrid registers that represent local identities and social connections. Greater linguistic diversity is correlated with heavy daily engagement, particularly among younger demographics that lead in the use of slang and expressive forms. Instagram becomes a dynamic space for identity negotiation, speeding up multilingual flexibility in response to interactive and algorithmic limitations. While comments allow for smooth multilingual conversations catered to the needs of the audience, captions condense stories with artistic flair. These trends show how the platform supports linguistic vitality in a variety of non-Western digital contexts by encouraging creative evolution without undermining fundamental communication norms.
Translation is the process of substituting text in the target language (TL) for text in the source language (SL). Catford 1965 said determining translation equivalency between two distinct linguistic systems is the main challenge intranslating. This research between Arabic-English translation that is actually has a big differents, aim to the distribution and percentage of pronouns containing in Surah Al-Kahfi, and investigates the translation ideology in the Arabic-English translation of Surah Al-Kahfi. The study using a descriptive qualitative method supported by quantitative analysis or Mixed Methode. The data consist of 532 pronouns identified from the Arabic text and its English translation.The pronouns are classified into personal pronouns, possessive pronouns, attached pronouns, and implicit pronouns. The pronouns are classified into shift type, tehniques type, ideology type and equivalence in translation. Shift type divided into unit shift, structural shift and split spit. Tehniques type divided into literal, transposition, amplification, reduction. Translation Ideology divided into foreignization and domestication, and Equivalence divided into formal and dynamic. The Findings release that Attached Pronoun and Implicit Pronoun are the majority found at the type of pronoun because AlQuran especially Surah Al Kahfi has the uniqe character. There are also Unit Shift has and Structural Shift are the most dominant categoriesat translation shift according to Catford 1965. Furthermore, Domestication appears slightlymore Dominant than foreignization due to Lawrence Venuti 1995 about translation ideology. At the end equivalenece shows that dynamic is more dominant than formal according to Eugene Nida1964. It is shown that reflecting the translator’s tendency to adapt Arabic grammatical structures to English linguistic norms while maintaining selected Qur’anic character.
Structured Abstract Objective Ambient artificial intelligence (AI) tools are increasingly adopted in clinical practices. This study investigated whether and how clinicians edit AI-generated drafts and the linguistic differences between AI drafts and clinician-finalized notes. Materials and Methods This retrospective study analyzed real-world data from ambulatory clinics at a large academic health system spanning two vendor deployments. We quantified clinicians’ editing behavior using the Myers diff algorithm to compare AI drafts and final documentation. We then applied statistical and linguistic analysis to study factors associated with the frequency/intensity of editing across note sections, turnaround time, clinician characteristics, and encounter types. Results Across 23,760 notes that included one or more ambient AI sections, 84.4% were edited by clinicians before signing off. While rates of unedited notes differed across note sections and care settings, the dominant source of variation was individual clinician practice style rather than specialty-level norms. Notes signed after 24 hours had lower overall edit intensity. The final versions showed small but statistically significant linguistic changes and exhibited slightly higher lexical diversity and modest changes in readability. Editing is most intensive in the assessment and plan section, and varies across specialties. Conclusion and Discussion A majority of AI-drafted clinical notes were edited by clinicians, although the editing rate varies across note sections, medical specialties, and individual clinicians. Future research is needed to further analyze this editing behavior to inform improvement in AI-assisted clinical documentation to achieve better documentation quality, efficiency, and clinician satisfaction.
ABSTRACT This study examines how Arabic language teachers in Indonesian madrasahs negotiate transnational Islamic authority and epistemic dependency within curriculum practices. Drawing on qualitative interviews, the research explores how global centres of Islamic knowledge—particularly Middle Eastern institutions, textbooks and linguistic norms—shape local understandings of legitimacy, authenticity and pedagogical authority. Rather than indicating structural domination, the findings show that epistemic dependency operates primarily as symbolic and referential alignment with perceived centres of authority. Teachers retain curricular autonomy and actively mediate these influences through contextual adaptation, aligning instruction with Indonesian cultural values, national educational frameworks and local religious traditions. These practices demonstrate negotiated forms of authority in which transnational standards function as reference points rather than mechanisms of control. The study contributes to discussions of language, curriculum and epistemic hierarchy in Global South contexts by highlighting how authority is symbolically structured yet locally mediated in practice.
This paper introduces an updated, publicly accessible version of the Film, Music and Emotion Dataset (FME-24), designed to examine how perceived emotion in film music evolves over time. It provides a comprehensive introduction to the dataset and explores its potential applications across music information retrieval (MIR), psychology, and AI-training contexts. The FME-24 dataset utilises film's immersive qualities to study emotional perception in a naturalistic yet controlled setting. It contains data from 275 film scores spanning the past two decades, including experimental and mainstream works. The dataset integrates high-quality film compositions with time-stamped valence-arousal (V-A) annotations, emotion sentences, familiarity ratings, and detailed metadata. 98 Participants contributed to annotating these temporal emotion features. For each time-stamped point, a two-second audio segment was analysed, and 78 features were extracted, including low-level timbral descriptors (MFCC statistics, spectral centroid), rhythmic descriptors (onset density, tempo), and higher-level psychoacoustic and tonal features (inharmonicity, roughness, chord transitions, tonal entropy). Although full audio files are unavailable due to licensing, reproducibility is ensured via ISRC codes, precise segment timings, and open access to all metadata and feature files in CSV format. The paper details the dataset's structure, annotation, and feature-extraction procedures, highlighting applications in computational and perceptual research and laying a foundation for future studies on emotion, perception, and narrative in film music.
Situated at the intersection of corpus stylistics, translation studies, and multifactorial statistics, this paper focuses on the identification of the predictors of preservation or avoidance of repetition in English-to-Polish translation of repeatedly used reporting verbs signalling direct speech. Using a sample of 20 literary texts, we fit multiple negative binomial regression with mixed effects models to assess the effect that seven predictor variables (i.e. frequency of a source-text verb, number of its translation equivalents in lexical databases, its number of senses, its semantic type, its length in characters, date of translation of a novel, and individual translators) have on the response variable: the number of Polish target-text reporting verbs (types) an English source-text reporting verb is translated into. The overall model fit per the lowest AIC (Akaike Information Criterion) and BIC (Bayesian Information Criterion) values obtained through backward elimination reveals that the frequency of a source-text reporting verb, its semantic type as well as the individual translators have the largest incremental contribution to the model’s fit. More precisely, the proportion of variance in the outcome variable explained by both fixed and random effects (76%) is higher than the proportion of variance explained by fixed effects alone (72%). Without the translators treated as a random intercept, some variation would have been unaccounted for. The findings attempt to explain the translator’s decisions, whether to avoid or preserve patterns of repetition in source texts, with respect to rendering reporting verbs, which play an important stylistic effect in literary prose.
ToS-100 contains 100 Terms of Service documents from online platforms, splitinto 20,417 clauses. Each clause is annotated for five categories of potentialunfairness: arbitration (A), unilateral change (CH), content removal (CR),limitation of liability (LTD) and unilateral termination (TER). A clause maycarry none of them (18,843 clauses, 92.3%) or several at once, so the task ismulti-label rather than multi-class. This deposit is the corpus of Ruggeri et al. (2022), itself built on the ToScorpus of Lippi et al. (2019), redistributed as three CSV files for Assignment 1of the Natural Language Processing course (Prof. Paolo Torroni, University ofBologna, a.y. 2026-2027). Files: train.csv 80 documents, 15,837 clauses validation.csv 10 documents, 2,548 clauses test.csv 10 documents, 2,032 clauses Columns: document_ID, document, text, A, CH, CR, LTD, TER. The split is by document, not by clause: all clauses of one contract stay inthe same split. Clauses of a single contract repeat each other almost verbatim,so a clause-level split would leak the test set into training. Positive clauses per category (train / validation / test): A 75 / 20 / 11 CH 268 / 43 / 33 CR 165 / 32 / 19 LTD 504 / 75 / 47 TER 318 / 59 / 43 Text is lowercased and tokenized in Penn Treebank style: brackets appear as-lrb- / -rrb-, quotes as `` and '', and clitics are detached (mozilla 's,do n't). The knowledge-base columns of the original release (*_targets), which point tofree-text legal rationales rather than to spans over the clause, are notincluded. Please cite the original works: Lippi et al., 2019. CLAUDETTE: an Automated Detector of Potentially Unfair Clauses in Online Terms of Service. Artificial Intelligence and Law. Ruggeri et al., 2022. Detecting and Explaining Unfairness in Consumer Contracts through Memory Networks. Artificial Intelligence and Law.
This thesis aims to analyze the linguistic and pragmatic deployment of perlocutionary acts in diplomatic speech by exploring Uzbek and English discourse texts. It endeavors to discover how and in what linguistic and pragmatic ways perlocutionary effects are encoded, spread, and construed within institutional communication practices in general. From a comparative perspective, the study scrutinizes the effect on language/cultural norms of the perlocutionary influence. It demonstrates that these practices of perlocution are established by lexical, grammatical, and discourse practices of politeness norms, indirectness, and communication intentions. This research advances the theoretical study of linguopragmatic language and also presents tangible implications for intercultural and diplomatic communication.
This study examines the phenomena of language interference that naturally arise when lower secondary school pupils use Indonesian as a second language in writing activities. The aim of this study is to identify various forms of language interference occurring at the phonological and morphological levels in descriptive texts written by pupils, whilst also analysing the factors that cause such interference to occur. The method used in this study is a qualitative descriptive method employing a comparative analysis technique between Indonesian language structures that conform to linguistic norms and the language system commonly used by students in their daily lives as their first language. The results of the study indicate that the most dominant interference is found in the phonological aspect, for example, the change of the sound /p/ to /b/, /u/ to /o/, and /a/ to /e/. Furthermore, interference is also evident in the morphological aspect, such as the use of the suffix -an, the use of reduplication forms, and the combination of words that do not conform to the rules of word formation in Indonesian. The emergence of this phenomenon is influenced by the widespread use of Javanese in everyday communication, oral language habits, and students’ still limited mastery of the written rules of Indonesian.
Large language models encode rich semantic information, but how concreteness is represented across layers remains unclear. We examine layer-wise linear separability of concreteness by training linear probes on hidden representations from two opensource model families at multiple scales: Qwen3 and Gemma3-<br/>Instruct. Using human concreteness ratings, we build balanced prompt datasets with four difficulty levels: an extreme abstract–concrete contrast and three finer boundary comparisons at the abstract end, mid-range, and concrete end. Probes achieve high accuracy on the extreme contrast in shallow layers, showing that endpoint differences are strongly linearly separable. For finer distinctions, performance follows a stable hierarchy: midrange concreteness is easiest to separate, abstract-end distinctions are hardest, and concrete-end distinctions are intermediate. Across models, accuracy rises rapidly in early layers, peaks in<br/>middle layers, and declines in later layers. Together, these findings clarify how the linear accessibility of concreteness varies across LLM layers.
The expression of an association between a conditioned stimulus (CS) and an aversive unconditioned stimulus (US) can be weakened by presenting the CS by itself (extinction [Ext]), pairing it with an appetitive US (counterconditioning [CC]), or pairing it with a neutral stimulus (novelty-facilitated extinction [NFE]). The present research tested whether NFE is less susceptible to ABC renewal than Ext and CC. In two experiments, participants viewed streams of rapid trials. After each stream, participants rated how likely it was that the target CS would be followed by the target US (i.e., predictive learning) as well as the valence of the target CS (i.e., evaluative conditioning). A stream was composed of two phases: Phase 1 established an association between the target CS and target US while Phase 2 aimed at disrupting the expression of this association through Ext, CC, or NFE. Phase 1 occurred in Context A while Phase 2 occurred in Context B. Prediction and valence ratings occurred in either Context A, B, or C. Neither Experiment 1 nor Experiment 2 found differences across interference conditions with predictive testing, regardless of test context. In Experiment 2, better controlled for context effect, CC and NFE altered the CS valence (CC more than NFE) when testing occurred in B, but the difference disappeared when testing occurred in either A or C. The present data do not support the hypothesis that NFE is less susceptible to ABC renewal than either Ext or CC. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
This article states that in the formation of the norms of the Karakalpak literary language, especially the grammatical norm, not only units belonging to its own lexical layer, but also words imported from abroad, are of certain importance. At the same time, it is noted that it is very difficult to determine the linguistic source of a word and that a certain lexical unit has several connections with different languages, that is, the unit can be borrowed from the Karakalpak language through an intermediary language.
ObjectiveAchilles tendon protective shoes are medical orthoses that must be worn after Achilles tendon rupture surgery. Currently, these shoes predominantly feature black and gray color schemes and lack conspicuous warning markings, making them visually obscure and prone to accidental impacts, thereby increasing the risk of re-rupture. By conducting perceptual research on the application of warning colors and textures in the shoe upper, this study aims to explore design combinations that can effectively enhance visual warning, thereby improving the safety of the product during use and reducing the risk of secondary injury.MethodsBased on the theory of warning colors and combined with the visual warning arousal mechanism and eye-tracking physiological indicators, three typical biological warning textures (dots, cracks, stripes) and three warning color combinations (orange-black, yellow-black, purple-yellow) were selected to create nine experimental combinations. These were applied to a uniform model of Achilles tendon protective shoes as visual stimulus materials for the experiment. After collecting subjective valence ratings and eye-tracking data (including Time to First Fixation and pupil diameter), the data were analyzed using repeated-measures ANOVA, followed by post-hoc multiple comparisons using the LSD method. Pearson correlation analysis was conducted to examine the consistency between subjective and objective indicators, thereby comprehensively evaluating the perceived warning intensity of each combination.ResultsThe nine combinations of warning textures and colors showed no significant difference in Time to First Fixation, but exhibited significant differences in pupil diameter response and subjective warning valence ratings, with a significant interaction effect between texture and color. A moderate positive correlation was observed between subjective warning valence ratings and the increment of pupil diameter. Among these, subjects exhibited stronger alertness and a greater tendency to make avoidance decisions when viewing striped yellow-black and striped orange-black patterns. In contrast, alertness levels were significantly lower when observing striped purple-yellow or dotted pattern combinations. This indicates that while stripes are superior to dots in eliciting visual alertness, their actual effectiveness is significantly influenced by color interactions. The results of multiple comparisons revealed that the differences identified by subjective warning assessments were not fully captured by the objective physiological indicator of pupil diameter. The striped yellow-black pattern received significantly higher subjective ratings, influenced by cultural learning, symbolic semantics, and individual experience. In contrast, while the cracked-orange-black pattern elicited significant pupil dilation, it was not explicitly perceived as highly alerting. The striped orange-black pattern elicited a pupil diameter response ratio reaching the sympathetic activation threshold, accompanied by high alertness ratings, forming a complete alertness perception cycle that integrates both bottom-up and top-down processing. In contrast, the dotted yellow-black pattern produced a pupil diameter response ratio reaching the parasympathetic activation threshold, which can suppress sympathetic activity during the early stage of alertness stress. The findings suggest that the design of warning colors for Achilles tendon protective footwear should prioritize the selection of colors based on texture, rather than independently selecting texture or color.ConclusionThe stripe pattern in yellow-black serves as the optimal biological warning color paradigm for Achilles tendon protective shoes, significantly enhancing safety. The orange-black stripe pattern can be considered a secondary option, while the dot pattern in yellow-black should be avoided in warning signs.
The article examines euphemisms used to name disability in Russian and English and interprets them as culturally sensitive linguistic tools that mediate between social norms, institutional regulation and personal identity. Drawing on pragmatics, sociolinguistics, linguistic politeness theory and disability studies, the research analyses how euphemistic nominations emerge, stabilise and shift across media discourse, legal and bureaucratic communication, education, and everyday interaction. The study compares the dominant euphemistic strategies in both languages and shows that disability naming tends to move from direct stigmatized labels to person oriented and rights oriented forms, while simultaneously generating new cycles of avoidance, abstraction and bureaucratisation. The empirical basis consists of illustrative examples from contemporary Russian and English public texts, including official documents, journalistic materials and public awareness communication. The analysis demonstrates that euphemisms of disability form a layered system in which lexical choice reflects competing ideologies, including person first versus identity first language, medical versus social models of disability, and charity narratives versus inclusion narratives. It is argued that euphemisms do not merely soften meaning but actively participate in constructing social reality, shaping attitudes to difference, normality and participation. In the context of globalisation and digital media, disability euphemisms circulate across languages, producing convergence in polite formulas but also local divergences rooted in distinct institutional histories and cultural expectations.
This study proposes a dual-stream late fusion approach to dimensional emotion recognition by fusing facial landmarks and electrocardiogram data. Unlike category-based methods, the proposed method identifies emotions in the valence-arousal space. By individually optimizing the classifiers for both modalities and then fusing their outputs, the model encourages flexibility, interpretability, and robustness. The ASCERTAIN dataset is used, which consists of facial landmark trajectories, physiological signals like electroencephalograms, electrocardiograms, galvanic skin responses, and electromyograms, and self-reported arousal and valence ratings of fifty-eight subjects who viewed thirty-six videos. Arousal and valence were each modeled using an individual classifier. The random forest-based feature selection achieved eighty-six point ninety-one percent accuracy in valence prediction, while a soft voting ensemble of support vector machines, K-nearest neighbors, and random forest achieved sixty-one point eighteen percent for arousal. These results were merged in order to classify four emotional quadrants: High Arousal High Valence, High Arousal Low Valence, Low Arousal High Valence, and Low Arousal Low Valence, with an aggregate accuracy of sixty-nine point sixteen percent. The findings indicate that the model introduced accurately makes use of complementary modalities and a late fusion strategy, and is suitable for real-world emotion-aware systems.
Abstract Impaired musical emotion perception is common in hearing loss, yet how reduced spectral audibility shapes cortical processing during emotion judgments remains unclear. Because alpha-band activity can index top-down control when sensory evidence is degraded, we examined whether spectral degradation increases the cognitive demand required to form stable affective judgments in music. Forty-eight healthy participants were divided into three groups: high-frequency hearing loss simulation (HF sim ), low-frequency hearing loss simulation (LF sim ), and normal hearing (NH). Participants rated the arousal and valence of filtered musical stimuli (happy, sad, neutral) during EEG recording. HF sim showed dimension- and context-dependent alpha modulations. In the happy condition, arousal ratings and alpha power were comparable across groups, whereas valence judgments showed behavioral differences and late-stage alpha increases in HF sim, consistent with reduced certainty when high-frequency cues supporting positive valence are degraded. In the sad condition, behavioral ratings were preserved, yet HF sim showed sustained alpha increases during arousal judgments, suggesting compensatory inhibitory-gating processes that may support stable appraisal under degraded listening. Overall, spectral degradation appears to elicit compensatory cognitive processing: alpha power increases index higher demands for happy-valence with reduced HF cues, and compensatory gating that maintains sad low-arousal appraisal.
Evaluative processes facilitate motivated responses to a wide range of stimuli in the environment, allowing individuals to make adaptive decisions in response to potential threats and rewards. These evaluations are influenced by many perceptual and biological factors. Recent perspectives have highlighted a need for more research examining how these individual factors interact to shape variability in evaluative processes. The present work examined how two such factors, resting parasympathetic activity and depressive symptoms, relate to affective evaluations of images in a nonclinical sample. Measures of depressive symptoms and resting parasympathetic activity were collected from a sample of young adults. Participants also rated positive, negative, and emotional arousal ratings to pleasant, neutral, unpleasant, and disgust pictures. Results showed that higher depressive symptoms were associated with increased positivity ratings of pleasant and neutral pictures, increased negativity ratings of pleasant, neutral, and disgust pictures, and ambivalence (high ratings of both positivity and negativity) of pleasant, neutral, and unpleasant pictures. Lower parasympathetic activity at rest was related to increased positivity ratings of pleasant and neutral pictures, increased negativity ratings of neutral pictures, increased ambivalence of neutral pictures, and increased emotional arousal ratings of pleasant pictures. This work suggests depressive symptoms and resting parasympathetic activity have complex effects across dimensions of evaluative processes. It points to a need for more research examining how these factors influence both positive and negative evaluations in both clinical and nonclinical populations.
Frontiers in Psychology Corpus (2010–2021)A comprehensive text corpus of 21,084 papers published in Frontiers in Psychology between 2010 and 2021, processed for computational linguistics and semantic analysis.Dataset OverviewThis corpus contains full-text articles from Frontiers in Psychology, converted from XML to plain text and preprocessed for natural language processing tasks.Files and Structurefpsyg_filtered.zipContains the filtered text corpus with the following preprocessing applied:- XML to text conversion: Original XML documents converted to plain text format- Sentence segmentation: Text segmented into individual sentences- Boilerplate removal: Journal metadata, headers, footers, and other non-content elements removed using filter_text_corpus.pyfpsyg_tagged.zip.* (2 parts)Contains the linguistically annotated corpus. This archive is split into multiple parts due to size constraints:- fpsyg_tagged.zip.001- fpsyg_tagged.zip.002To extract: First combine the parts using 7-Zip:bash7z x fpsyg_tagged.zip.0017-Zip will automatically detect and combine all parts, then extract the contents.After extraction:- File: fpsyg_filtered_tagged.conllu- POS tags: Penn Treebank part-of-speech tags- Dependency parsing: Universal Dependency (UD) tags- Tagger: Processed using Stanza- Format: CoNLL-U formatfpsyg_index.zip.* (6 parts)Contains a compiled binary index for semantic mining. This archive is split into multiple parts:- fpsyg_index.zip.001 through fpsyg_index.zip.006To extract: First combine the parts using 7-Zip:bash7z x fpsyg_index.zip.0017-Zip will automatically detect and combine all parts, then extract the contents.After extraction:- Compatible with ConceptSketch semantic mining software- Usage: Point ConceptSketch to the extraction directoryProcessing PipelineOriginal XML → Text Conversion → Sentence Segmentation → Boilerplate Removal → POS/UD Tagging (Stanza) → Final CorpusLicenseThis dataset is distributed under the Creative Commons Attribution License (CC BY), derived from the original papers' licensing terms.CitationIf you use this dataset in your research, please cite:Dataset compiled by: Marcin MiłkowskiAffiliation: Cognitive Metascience Lab & Center for AI in Society, Institute of Philosophy and Sociology, Polish Academy of SciencesRequirementsStanza (for re-tagging or custom processing): https://github.com/stanfordnlp/stanzaConceptSketch (for semantic mining): https://github.com/cognitive-metascience/concept-sketchContactFor questions regarding this dataset, please contact the Cognitive Metascience Lab at the Institute of Philosophy and Sociology, Polish Academy of Sciences: https://cognitive-metascience.github.io