Papers reviewed and determined not to be word norm studies. Use the flag icon to report errors or suggest re-inclusion.
16504 papers
This article investigates the pragmatic functions of first-person pronouns (I/we) in political speeches and formal writing across multiple language systems. By analyzing the discourse of prominent political leaders such as Shavkat Mirziyoyev, Joe Biden, and Emmanuel Macron, the study explores how first-person pronouns function to construct authority, inclusiveness, and responsibility. The research highlights variation across languages in terms of politeness, formality, and rhetorical strategy. Methodologically, the study employs comparative discourse analysis and pragmatic interpretation of political and academic texts. The findings demonstrate that the usage of "I" versus "we" reflects not only linguistic norms but also culturally embedded leadership styles.
The incorporation of real-world contexts into mathematical problems has become increasingly significant on a global scale, which is also reflected in Vietnamese education reform. This emphasis has resulted in the development of different series of textbooks designed to enhance student engagement and foster active learning through real-world problem contexts. In particular, the geometry and measurement strand presents substantial opportunities for the integration of contexts. In this study, we investigate the theoretical foundations of contexts, context familiarity, and context authenticity in mathematics education, with a focus on the plane geometry content in ninth-grade Vietnamese textbooks. The analysis is supported by four illustrative examples drawn from three reformed textbook series, demonstrating how context familiarity and authenticity are represented. A 5-point Likert-type survey of 109 students (72 males and 37 females) revealed that the soccer context was rated as less familiar than the airplane context overall, with no significant difference in familiarity ratings between boys and girls. Both the bike-riding and tripod contexts were rated as moderately authentic but not highly so. The authors also acknowledge several limitations and offers certain recommendations for future research in the area of contextual mathematics education.
Salient sexual cues (erect penis, attractive individuals) are thought to capture initial attention and automatically trigger genital arousal in women. Conscious appraisal activates subjective sexual arousal and further visual attention. The present study tested whether the attractiveness category (attractive/unattractive) and/or sexual arousal condition (in underwear, naked with flaccid penis, naked with erect penis) of male stimuli predicted female sexual responding and attentional patterns. Genital arousal, visual attention, and subjective ratings (subjective sexual arousal, pleasantness) of 26 predominantly heterosexual women (Mage = 31.2, SDage = 6.8) were measured while exposed to male stimuli across the experimental conditions. Neither attractiveness nor sexual arousal condition of male models significantly predicted genital arousal, subjective sexual arousal ratings or visual attention patterns of women. Results however showed high concordance between genital and subjective sexual arousal measures, and the pleasantness rating of the stimuli positively predicted subjective sexual arousal, suggesting a positive feedback loop of female sexual response.
French is often celebrated for its clarity and precision – a legacy shaped by Cartesian rationalism and prescriptive language policies. However, the evolving forms of spoken French challenge this ideal of fixed linguistic norms. This study examines one such feature: the right-peripheral duplication of the subject pronoun je with its tonic counterpart moi, a recurrent but underexplored phenomenon in spoken French. The primary objective is to understand how this syntactic feature functions pragmatically and emotionally in real-life discourse. Using a corpus of movie dialogues, the analysis shows that duplication plays a role in managing conversational flow, expressing personal stance, and enabling self-repair. Through a multidisciplinary lens that draws from sociolinguistics, pragmatics, and applied linguistics, the study argues that such variation enriches the expressive potential of French and complicates the rigid divide between written norms and spoken practice. It also suggests that incorporating these features into language pedagogy can support a more inclusive, realistic understanding of French as a living language.
Large language models (LLMs) are widely deployed in settings where both reliability and efficiency matter. We present a calibrated, seed‑robust empirical comparison of an encoder fine‑tuned model (bidirectional encoder representations from transformers (BERT)‑base) and a decoder in‑context model (generative pre-trained transformer (GPT)‑2 small) across Stanford question answering dataset v2.0 (SQuAD v2.0) and general language understanding evaluation (GLUE)-multi-genre natural language inference (MNLI), Stanford sentiment treebank 2 (SST‑2). Beyond accuracy, we assess reliability (expected calibration error with reliability diagrams and confidence–coverage analysis) and efficiency (latency, memory, throughput) under matched conditions and three fixed seeds. BERT‑base yields higher accuracy and lower calibration error, while GPT‑2 narrows gaps under few‑shot prompting but remains more sensitive to prompt design and context length. Efficiency benchmarks show that decoder‑only prompting incurs near‑linear latency/memory growth with k‑shot exemplars, whereas fine‑tuned encoders maintain stable per‑example cost. These findings offer practical guidance on when to prefer fine‑tuning versus prompting and demonstrate that reliability must be evaluated alongside accuracy for risk‑aware deployment.
The current study aimed to examine the effects of topic familiarity and language proficiency on linguistic complexity, accuracy, and fluency of argumentative essays in an EFL context. This study involved 64 college freshmen, who were divided into two groups according to a TOEIC score of 700, which corresponds to the B1 level on the CEFR scale: a high group (n = 31) and an intermediate group (n = 33). Participants were asked to write two argumentative essays on different topics, controlling for order effects through counterbalancing: one familiar (driving) and the other unfamiliar (smoking). They also completed a questionnaire that included background information and a topic familiarity rating on a 10-point Likert scale. The participants’ writing samples were analyzed in terms of lexical complexity, syntactic complexity, accuracy, and fluency. The results indicated that, in terms of topic familiarity, EFL learners tended to produce texts with lower levels of lexical and syntactic complexity, as well as reduced accuracy, when writing about unfamiliar topics. With regard to language proficiency, advanced learners demonstrated a broader vocabulary range, employed longer and more complex sentence structures, and produced more accurate and extensive texts compared to their intermediate-level peers. In-depth analysis and pedagogical implications are discussed.
This study investigates the thematic structure of Persian imperative clauses within the framework of Systemic Functional Grammar (SFG), focusing on the concept of Theme. In this framework, the Theme is defined as the first element of a clause that is a Participant, a Process, or a Circumstantial adjunct. In Persian, the unmarked Theme in most constructions follows the pattern of declarative clauses; however, the thematic position in imperative clauses has been little studied. To address this, 2,099 imperative clauses were extracted from the Persian Syntactic Dependency Corpus, and half of them were analyzed for thematic structure. It was hypothesized that in some clauses—particularly those with motion verbs—the verb might serve as the unmarked Theme. The findings reveal that more than 90% of Persian imperatives are verb-final. In about 10% of the clauses, a verb with a clausal complement appears initially, with its complement at the end. Among these, only about 0.5% are simple clauses in which the verb itself occurs in the initial position. Overall, the results show that the thematic pattern of Persian imperative clauses conforms to the verb-final typology of the language.IntroductionSystemic Functional Grammar (SFG) conceptualizes grammar as a meaning-oriented system shaped by social use. Within this framework, each clause simultaneously realizes three metafunctions: experiential, interpersonal, and textual. The experiential metafunction is concerned with processes, participants, and circumstances, while the interpersonal metafunction organizes relations between speakers and addressees; and the textual metafunction organizes the clause as a message in discourse.Central to this metafunction is the notion of Theme, which identifies the element chosen as the point from which the clause unfolds. Thematic choices reflect how speakers manage information flow, relate clauses to their context, and guide addressees toward an intended interpretation. As such, the analysis of thematic structure provides insight into the interface between grammar, discourse, and communicative intent.Cross-linguistic research has shown that unmarked thematic patterns often vary according to clause type. In imperative clauses, for example, most languages display subject omission as a conventional strategy. In English, this pattern typically coincides with verb-initial imperatives, where the Process functions as the unmarked Theme. Persian, however, differs typologically as a verb-final language, raising questions about how imperative clauses are thematically organized within its grammatical system.Although word order and thematic structure in the Persian language have been widely discussed for declarative and interrogative clauses, imperative constructions have received comparatively little systematic attention. This gap is noteworthy given the frequency of imperatives and their central role in instructional, persuasive, and interpersonal discourse.Assuming subject omission in imperative clauses as a cross-linguistic common pattern, the present study focuses on how thematic structure is realized in Persian imperatives. It examines whether the unmarked thematic pattern of Persian declarative clauses extends to imperatives and whether the Process can function as the unmarked Theme, as observed in English. By addressing these questions, the study seeks to identify the unmarked thematic structure of Persian imperative clauses through a corpus-based analysis.Research QuestionsThis study addresses the following research questions:Does the unmarked thematic pattern of Persian declarative clauses extend to imperative clauses?Can the Process (verb) function as the unmarked Theme in Persian imperative clauses, as it does in English?Literature ReviewStudies of word order have long been shaped by typological research, notably the work of Greenberg (1963, 1974), Dryer (1992), and Bybee (2002), which seek to explain cross-linguistic ordering patterns in functional terms. Functional approaches treat word order as a flexible means of organizing discourse rather than a purely syntactic constraint, emphasizing its role in information management and contextual interpretation. Research on Persian has largely concentrated on declarative and interrogative clauses, showing that verb-final order is predominant but readily modified for discourse-related purposes (Dabir-Moghaddam & Sharifi, 2005; Rezaei & Tayyeb, 2006; Tabatabaei & Modarresi, 2017). Corpus-based studies of Persian Wh-questions further indicate that unmarked thematic patterns in Persian differ from those in English and vary according to clause type (Mirzaei, 2019). Despite these advances, the thematic structure of imperative clauses in Persian has received little systematic attention.MethodologyThe data for this study were extracted from the Persian Syntactic Dependency Treebank (Rasooli, et al., 2013). Using verbal mood annotations, 2,099 imperative clauses were identified. From this dataset, 1,276 clauses were selected sequentially for detailed analysis. Each clause was examined for word order and coded for the presence and position of the Process, Participants, Circumstantial adjuncts, and vocatives. This approach allowed for a quantitative and qualitative analysis of thematic patterns in Persian imperative clauses.ResultsThe analysis shows that the vast majority of Persian imperative clauses are verb-final. Verb-initial structures are rare and largely restricted to clauses containing clausal complements.Figure 1Position of the verb in Persian imperative clausesIn verb-final imperatives, Circumstantial adjuncts are the most frequent unmarked Themes, followed by Participants. Figure 2Ordering of circumstances in relation to direct and indirect objectsVocatives typically occur in the initial position but do not function as Themes. The modal adjunct lotfan, ‘please’ appears frequently at the beginning of clauses, but does not alter the underlying verb-final thematic structure. DiscussionThe findings show that Persian imperative clauses differ fundamentally from their English counterparts in thematic structure. The strong preference for verb-final structures reflects broader typological constraints and suggests that thematic choices in the Persian language are shaped primarily by language-specific syntactic patterns rather than by clause type alone. The prominence of Circumstantial adjuncts as Themes highlights the role of contextual framing in Persian directive discourse.ConclusionThis study demonstrates that the unmarked thematic structure of Persian imperative clauses conforms to the verb-final typology of the language. The Process does not function as the unmarked Theme, even in clauses with motion verbs. Instead, non-verbal experiential elements—particularly Circumstantial adjuncts—are favored in the thematic position. These findings contribute to functional and typological studies of Persian and open new avenues for research on the interaction between thematic structure, modality, and polarity in imperative constructions.
Metaphorical mappings play a central role in embodied cognition, with prior research suggesting that power is associated with higher vertical spatial locations. This study investigated whether spatial cues influence gender classification and dominance attributions for androgynous faces that lack inherent gender or power cues. Using a multi-laboratory approach across twelve countries (N = 546), participants categorized gender-ambiguous faces as male or female and rated their dominance after brief exposures (100 msec vs. 500 msec) at different spatial positions. We hypothesized that faces in higher locations would be more frequently classified as male and rated as more dominant, particularly with longer exposure durations. However, results revealed that while exposure duration affected gender classification (longer durations increased the likelihood of male categorizations), spatial positioning had no significant effect on response times or dominance ratings. Instead, dominance ratings were strongly linked to gender classification, with faces categorized as male being rated as more dominant, regardless of spatial positioning. These findings suggest that the power-space metaphor may not operate automatically in social perception without salient contextual cues, while highlighting the persistent influence of implicit gender biases in dominance attribution.
This article explores the role of five expressive punctuation marks – multiple exclamation marks (!!!), exclamation mark (!), full stop (.), ellipsis (...), and null punctuation (ø) – as cues to writer attitudes in CMC. Specifically, it investigates how the underlying meaning of expressive punctuation influences the perceived emotional valence of discourse referents within exclamative constructions. In a 1x5 between-subjects repeated measures design, valence ratings were collected for 120 discourse referents embedded in exclamative constructions manipulated by message finale punctuation mark (e. g., What a view!!!/!/./…/ø) on a 1 (negative) to 9 (positive) scale. For inherently positive discourse referents, a clear positivity hierarchy in the overall valence of embedded discourse referents emerges, indicating a differential influence of punctuation on perceived valence: multiple exclamation marks > exclamation marks > null punctuation > full stop > ellipsis. For inherently negative discourse referents, the differences between the conditions are less distinct. Notably, only multiple exclamation marks yield a significantly lower valence rating within that range of values. While the findings for inherently positive referents align with prior assumptions on the expressive meaning of different punctuation marks in CMC, the observed pattern for inherently negative referents cannot be readily explained by existing literature on expressive punctuation in CMC.
The article systematizes and comprehensively analyzes the phenomena of the phonetic level of the Internet vocabulary of the modern Kazakh language. Instagram, Facebook, social networks (Threads, Instagram, Facebook), and instant messengers (WhatsApp, Telegram), which have been actively used in recent years, were chosen as the object of the study. The research used methods of observation, generalization, comparative and descriptive analysis. As a result, it is revealed that new forms of linguistic usage are being formed in the Internet space, characterized by a mixture of elements of spoken and written speech. At the phonetic level, phenomena such as sound compression of words and, conversely, the repetition of graphemes to convey emotions in writing are widespread. The active use of Latin graphics and the development of foreign-language sounds indicate a new stage of phonetic adaptation in the Kazakh-speaking Internet space. The article provides specific examples of these phenomena, reveals their causes and impact on the modern linguistic norm and writing culture. According to the results of the study, it was found that the phonetic features of the Internet vocabulary reflect the natural development and adaptability of the Kazakh language.
This paper presents a new version of the Spoken Slovenian Treebank (SST), a balanced and representative collection of transcribed spontaneous speech with manually annotated lemmas, part-of-speech tags, morphological features, and syntactic dependencies, recently expanded with over 3,000 newly annotated utterances. After a brief overview of the data sampling, anno-tation, and consolidation processes—presented in detail in previous work—we evaluate the significance of this new language resource for both linguistic research and natural language pro-cessing by first highlighting its distinctive lexical and morphosyntactic features in comparison to writing, and then assessing their impact on the performance of tools for automatic grammatical annotation. Finally, we reflect on the methodological insights gained during treebank creation, discuss the potential of SST for advancing spoken language research, and argue for the necessity of such resources in supporting linguistic diversity in language technology.
In ”Bartleby, the Scrivener,” Herman Melville presents a character whose passive refusal, encapsulated in the repeated phrase “I would prefer not to,” challenges power, agency, and social norms. This essay examines how Bartleby’s refrain acts as both an assertion of autonomy and a critique of the violence inherent in language. By rejecting his employer’s commands, Bartleby disrupts the rational, efficiency-driven logic of the workplace, exposing the violence embedded in linguistic norms. Slavoj Žižek’s concept of language as inherently violent—through its imposition of norms and standards— illuminates how Bartleby’s refusal goes beyond protest, creating a space of resistance that defies interpretation and subverts power dynamics. Bartleby’s language, neither a clear denial nor an expression of desire, becomes a radical negation that questions the very nature of meaning. Ultimately, Bartleby’s refusal does not propose a new order but disrupts the structures of meaning and authority, forcing us to confront the limits of language itself.
<div> We address the challenge of syntactic parsing for Urdu, a morphologically rich language, and present state-of-the-art results for both constituency and dependency parsing. This paper offers four major contributions: 1) the conversion of the CLE-UTB phrase structure treebank into a dependency treebank by developing language-specific head-word and phrase-to-dependency label mapping rules; 2) a novel sequence labeling scheme that transforms the parsing task into a unified representation; 3) the training of contextualized word representations on a large 220 million tokens Urdu corpus collected from the web; and 4) development of parsing framework using two learning paradigms, single-task and multi-task learning. Several post-processing rules are applied to improve the quality of the automatically converted dependency structure treebank. The proposed sequence labeling scheme enables the use of a shared architecture that learns the syntactic structures from both grammatical structures simultaneously and hence improves generalization. Experiments show that the multi-task learning setup significantly enhances parsing performance, achieving an F1 score of 91.39 for constituency parsing (an improvement of 3.29 points) and a labeled attachment score of 85.69 for dependency parsing (an improvement of 1.49 points). These results demonstrate that learning cross-task representations provides measurable benefits and advances the state of syntactic parsing for Urdu. </div>
BACKGROUND AND OBJECTIVES: The Seeking Proxies for Internal States (SPIS) model of OCD posits that reduced access to internal states plays a key role in the development and maintenance of the disorder. The current work sought to provide further support for the model's central claim that obsessive-compulsive tendencies are associated with reduced access to internal states. METHOD: Participants (N = 170) listened to 60 sound stimuli, rated how each one made them feel, and completed a measure of obsessive-compulsive tendencies. Following past procedure, we compared participants' ratings to each sound's normative valence rating, such that higher deviations between the ratings reflect a noisier perception of affective internal states. RESULTS: As hypothesized, higher obsessive-compulsive tendencies predicted greater deviations for both normatively-positive and normatively-negative sounds. CONCLUSIONS: The current work provides additional, novel support for the SPIS model, showing that with increasing obsessive-compulsive tendencies, people exhibited reduced attunement to how auditory stimuli made them feel.
Math anxiety poses significant challenges for university psychology students, affecting their career choices and overall well-being. This study employs a framework based on behavioural forma mentis networks (i.e. cognitive models that map how individuals structure their associative knowledge and emotional perceptions of concepts) to explore individual and group differences in the perception and association of concepts related to math and anxiety. We conducted 4 experiments involving psychology undergraduates from 2 samples (n1 = 70, n2 = 57) compared against GPT-simulated students (GPT-3.5: n2 = 300; GPT-4o: n4 = 300). Experiments 1, 2, and 3 employ individual-level network features to predict psychometric scores for math anxiety and its facets (observational, social and evaluational) from the Math Anxiety Scale. Experiment 4 focuses on group-level perceptions extracted from human students, GPT-3.5 and GPT-4o's networks. Results indicate that, in students, positive valence ratings and higher network degree for "anxiety", together with negative ratings for "math", can predict higher total and evaluative math anxiety. In contrast, these models do not work on GPT-based data because of differences in simulated networks and psychometric scores compared to humans. These results were also reconciled with differences found in the ways that high/low subgroups of simulated and real students framed semantically and emotionally STEM concepts. High math-anxiety students collectively framed "anxiety" in an emotionally polarising way, absent in the negative perception of low math-anxiety students. "Science" was rated positively, but contrasted against the negative perception of "math". These findings underscore the importance of understanding concept perception and associations in managing students' math anxiety.
This paper discusses the concept of child agency in home language maintenance in the context of transnational Polish-Australian families. Smith-Christmas’ (2022) model has been applied as a vantage point to present how the four intersecting dimensions of compliance regimes, linguistic competence, linguistic norms and generational positioning illustrate the dynamics of interactional practices in a family. To exemplify how complex and multi-layered child agency is, three conversational excerpts have been selected. While addressing the aforementioned dimensions, it occurred that children exerted agency through certain acts to shape language practices in the family. The four interrelated dimensions either unfolded individually to prove the children’s agentive use of language and/or behavior being a testimony to the fact that children are already fully-fledged actors in the family. On other occasions, the said dimensions transpired in the form of an exhaustive model, with all of them evolving in parallel, where compliance regimes, linguistic norms, linguistic competence and generational positioning showcased the multidimensional character of child agency.
The study of linguistic variation within the administrative structures of small-town America reveals a complex intersection between language, social identity, and institutional behavior. When approaching the linguistic environment of these communities from a purely academic perspective, without relying on personal immersion narratives or experiential accounts, one must begin with the foundational premise that English in the United States is profoundly regionalized. This regionalization is not a superficial matter of accent or vocabulary; it is a system of deeply embedded linguistic norms that shape how communication occurs, how authority is interpreted, and how institutional legitimacy is constructed. My interest as a researcher lies not in documenting local flavor or collecting curiosities from rural life but in understanding the mechanisms by which language operates as a structural force within governance. This requires an examination of sociolinguistic corpora, regional dialect research, institutional discourse studies, and the extensive literature on American dialect geography that has accumulated since the mid-twentieth century.
This research aims to analyze the Politeness and Speech Acts of the Community (Ojol Community). This research uses a qualitative approach with descriptive method. Data were obtained through direct observation and recording of conversations between online ojek drivers and customers in real situations. Recording is done naturally without intervention to reflect authentic speech acts. The audio data is then transcribed and analyzed using Searle's speech act theory. The analysis is done descriptively qualitative by classifying and interpreting the form and function of utterances in the context of the conversation. The results show that nonstandard language is more dominantly used in informal communication, such as conversations between online ojek drivers and passengers, because it is considered more familiar, relaxed, and efficient. However, mastery of standardized language is still important, especially in official contexts, to maintain clarity and politeness. People are expected to be able to adjust the use of language according to the context so that communication remains effective and in accordance with linguistic norms. Keywords:,,,,.
The article addresses the concept of linguistic purism in the context of globalization, when language barriers weaken and borrowings become commonplace. Purism, as an ideology, focuses on preserving the purity of the language, its stability and protection from external influences. The basic principles of purism, such as protecting the national language from foreign borrowings, preserving traditional norms, and countering linguistic changes are considered in the article. Two types of purism can be singled out, such as gustatory, based on subjective criteria, and scientific, which requires further clarification. The typology of linguistic purism is considered in terms of orientation, goals, and the nature of relations to linguistic facts. The experience of puristic activity in different languages and, accordingly, different societies, cultures and historical contexts are analysed. In general, linguistic purism is becoming the subject of topical discussions in the context of standardization and codification of the Russian literary language. Key words: language, vocabulary, linguistic norm, codification of linguistic norm, socio-cultural functions of language, sociolinguistics, psycholinguistics, linguistic (linguistic) purism.
Kyrgyz, a Turkic language with over 4.4 million speakers concentrated primarily in Kyrgyzstan and adjacent regions of Central Asia, faces a significant disparity in computational linguistic resources compared to languages with similar or even smaller speaker populations. Despite its status as a government language and cultural cornerstone, Kyrgyz remains underrepresented in the digital linguistic landscape. This investigation examines the application of the Universal Dependencies (UD) framework – an annotation system engineered to facilitate cross-linguistic syntactic comparability – to the structural complexities of Kyrgyz. We endeavor to identify optimal annotation strategies that faithfully represent Kyrgyz-specific syntactic phenomena while adhering to the principled constraints of the UD paradigm. The establishment of standardized syntactic resources for Kyrgyz carries dual significance: it advances linguistic typology by incorporating data from an underrepresented language family, while simultaneously laying groundwork for practical natural language processing applications crucial for Kyrgyz speakers’ participation in the digital sphere. Our methodological approach encompasses rigorous analysis of nascent Kyrgyz treebanks, comparative evaluation of annotation strategies employed for genetically related Turkic languages, and systematic examination of four fundamental annotation challenges: the representation of Kyrgyz’s defective copula system, the classification of multifunctional grammatical particles, the annotation of constructions with implicit heads, and the demarcation between inflectional and derivational morphology in this highly agglutinative language. Our analysis reveals that achieving the dual objectives of linguistic fidelity and cross-linguistic consistency necessitates judicious adaptation of UD guidelines to accommodate Kyrgyz-specific structures. We advance unified annotation solutions that preserve the integrity of Kyrgyz linguistic patterns while facilitating meaningful cross-linguistic comparison. This research not only contributes substantively to computational resources for Kyrgyz but also establishes annotation principles with broader applicability to typologically similar agglutinative languages. The practical implications extend to enhanced guidelines for Kyrgyz treebank development, which will consequently improve parser accuracy and catalyze the development of essential language technology tools for Kyrgyz speakers.
This article examines the methodological aspects of using the National Corpus of the Kazakh language in school-based Kazakh language lessons and analyzes its theoretical and practical significance. The study demonstrates the effectiveness of the national corpus in developing students’ communicative skills—namely, reading, speaking, listening, and writing—and in forming linguistic norms. It was identified that the corpus provides natural language data, enabling students to understand word meanings through text structure, observe the functions of grammatical forms, and distinguish stylistic and communicative features of speech. In addition, corpus materials help students master the structure of dialogue, use words accurately, and understand phraseological units. Regarding linguistic orientation, the corpus fosters learners’ research skills by enhancing their ability to independently identify language patterns, compare linguistic phenomena, and draw conclusions. The findings of the study show that the national corpus is an important tool for improving linguistic competence, developing speech culture, and strengthening research skills. The article focuses on two main directions—developing communicative skills and forming linguistic norms through the use of the corpus. The results confirm that systematic use of the national corpus significantly enhances students’ linguistic competence, speech culture, and research abilities.
Chinese word segmentation is a foundational task in natural language processing (NLP), with far-reaching effects on syntactic analysis. Unlike alphabetic languages like English, Chinese lacks explicit word boundaries, making segmentation both necessary and inherently ambiguous. This study highlights the intricate relationship between word segmentation and syntactic parsing, providing a clearer understanding of how different segmentation strategies shape dependency structures in Chinese. Focusing on the Chinese GSD treebank, we analyze multiple word boundary schemes, each reflecting distinct linguistic and computational assumptions, and examine how they influence the resulting syntactic structures. To support detailed comparison, we introduce an interactive web-based visualization tool that displays parsing outcomes across segmentation methods.
Abstract: Emotions impact pain; appetitive (pleasant) emotions reduce pain, and aversive (unpleasant) emotions increase pain. Emotion regulation (ER) strategies can alter emotional experience, and we have shown that ER can alter emotional modulation of pain, but not emotional modulation of spinal nociception (as assessed by nociceptive flexion reflex, NFR). The current study examined whether ER influences the emotional modulation of cortical event-related potentials (ERPs) in response to nociceptive input. To investigate, 68 pain-free individuals viewed pleasant (erotic), neutral, and unpleasant (mutilation) pictures during which painful electric stimulations were delivered. Participants viewed one block of pictures without engaging in ER and were then randomly assigned to ER (suppress or enhance) employed during a second block of pictures. Picture-evoked valence and arousal ratings, skin conductance response, and corrugator electromyogram (EMG) suggested that ER successfully regulated emotional experience. Instructions to suppress led to a significant reduction of emotional modulation of self-reported pain (i.e., reducing pleasure-induced pain inhibition and displeasure-induced pain facilitation), but neither NFRs nor ERPs were affected. Paradoxically, enhance instructions had no effect on pain or NFR, but were associated with a general suppression of the P260 ERP, indicating non-specific emotional arousal up-regulation. Findings indicate ER can independently impact emotional modulation at perceptual and supraspinal levels, but does not impact spinal nociception. Given that emotional suppression led to decreases in displeasure-evoked pain facilitation without reducing displeasure-evoked facilitation of spinal or supraspinal responses to nociception, future research should determine whether this divergence is associated with positive or negative long-term consequences.
There is ample evidence of the influence of both self-reference and the emotional content of words in language processing and memory. This study examines the conjoint influence of both factors in a variant of the affective HisMine-Paradigm. Participants were presented with pairs of words, that comprised emotional (positive and negative) or neutral words, preceded either by the first-person possessive pronoun "my" (self-reference) or by the definite article "the" (no-reference). Emotional words were divided into emotion-label words (e.g., happiness) and emotion-laden words (e.g., party). Participants were asked to perform an affective evaluation task (i.e., to decide if the word pair conveyed a positive, negative or neutral meaning), followed by a valence rating task (i.e., to rate the word pair in terms of their valence) and an unexpected free recall task. The results for positive words, but not negative words, showed that self-reference facilitated the affective evaluation task and led to more extreme valence ratings. These modulatory effects were observed in emotion-laden words, but not in emotion-label words. These findings support a self-positivity bias, and the literature about the modulation of emotional word processing by self-reference, yet point out the relevance of the distinction between emotion-label words and emotion-laden words.
The article investigates the role of language and accent in the United Kingdom as instruments of social stratification and as carriers of ideological constructs. It traces the historical development of the linguistic landscape, from the Celtic languages and the formative period of English to the contemporary situation of minority (Celtic) languages and the languages of migrant communities. Particular attention is devoted to accents as powerful social markers: the standard variety, Received Pronunciation, has traditionally been associated with high social status and elitism, whereas regional and ethnic accents may be subject to prejudice and function as indicators of class and group affiliation. The study highlights how the education system and the media reinforce the hierarchy of accent prestige, thereby shaping opportunities for social mobility. It also examines the discourse surrounding migrant languages and the role of English as both a vehicle of integration and a tool of social control. The article concludes that language in British society operates not only as a medium of communication but also as a mechanism intrinsically linked to ideology: linguistic norms, accents, and social varieties both reflect and reproduce existing hierarchies.
This study investigates the translation of 145 Arabic idioms by two prominent machine translation systems: Google Translate and ChatGPT. The research employs both quantitative and qualitative methodologies to analyze translation approaches and assess accuracy in conveying idiomatic meanings from Arabic to English. Data collection involved Arabic idiomatic expressions from literary sources, cultural texts, and linguistic databases. The analysis framework builds upon Baker's (1992) taxonomy of translation strategies. Quantitative findings reveal that Google Translate employed literal translation in 74% of cases, while ChatGPT demonstrated more varied approaches with 48% literal translations. For sense-based translations using non-figurative language, ChatGPT led with 41%, compared to Google Translate's 15%. When examining figurative language translations, ChatGPT achieved 11% compared to Google Translate's 11%. The qualitative analysis highlights persistent challenges in both systems regarding cultural context preservation and semantic accuracy. The study concludes that while technological advances have improved machine translation capabilities, rendering Arabic idioms into English remains problematic due to cultural-linguistic gaps and contextual complexities inherent in idiomatic expressions.
Learner Handover (LH) involves sharing information about learners between faculty supervisors, aligning with a growth mindset. Previous studies, however, demonstrate LH can bias subsequent ratings. Most of these studies collect ratings after a single encounter but faculty often have multiple interactions with learners potentially mitigating LH-related bias. This study explored if LH influences faculty ratings, entrustment decisions and feedback after observing several encounters of the same learner. Internal medicine faculty (n = 57) from five medical schools were randomly assigned to one of three study groups. Each group received either positive, negative or no LH prior to watching five simulated resident-patient encounter videos of the same white male resident. Participants rated each video using an entrustment scale, the Mini-CEX and provided written feedback. Feedback was assigned a valence score (-3 to + 3). There were no statistically significant differences between the mean ratings across the LH conditions (positive, control, negative) for entrustment [3.42, 3.26, 3.62], Mini-CEX [6.00, 5.90, 6.28] or feedback valence ratings [-0.34, -0.99, -0.74]. In the post-study questionnaire, most raters reported the LH had minimal effect on their decisions. Only 29% of raters guessed the true purpose of the study. Unlike previous studies, LH had no effect on ratings, entrustment decisions, or feedback after one encounter, nor over subsequent encounters with the same resident. These findings suggest LH's influence may vary and highlight the need for replication under different conditions, including diverse genders and equity-deserving groups, to identify factors that contribute to or mitigate bias.
The increasing use of English in many professional and academic contexts has played a pivotal role not only in expanding the teaching of English at universities worldwide but also in determining the growth of teaching in English. This has heightened the need to teach thesis writing in many English-medium contexts. Focusing on the tension between individuality – expressing an individual point of view – and commonality – adopting the rhetorical and linguistic norms of the specific discourse community addressed, we discuss our experience of designing an introduction to thesis writing for an EMI (English as a Medium of Instruction) multidisciplinary MA programme in Italy. The chapter begins by exploring the interplay between commonality and individuality, then proceeds to describe the course’s context and its educational underpinnings. The general principles are illustrated through an analysis of activities meant to teach the students how to engage with the discourse community of their choice and to express their stance: from engaging with discourse communities to finding relevant sources, incorporating other voices into a literature review, and working collaboratively on the conclusion section. The chapter closes with a brief summing up of the issues raised.
With the increasing application of automated news writing and digital anchors in journalism, AIGC (Artificial Intelligence-Generated Content) has become highly realistic in form and increasingly conforms to the linguistic norms of traditional news. However, AIGC news often suffers from ambiguous sources and insufficient factual support, leading to a crisis of authenticity. Drawing on Baudrillards theory of simulacra, this paper systematically analyzes the manifestations of the authenticity crisis in AIGC news at the levels of content generation, expressive form, and audience perception, revealing its impact on the ontological logic of news content. The study argues that AIGC drives a shift in news authenticity from a logic of facts to a logic of perception. Its high imitation of traditional news styles in language, layout, and headline structure creates a hyperreal illusion, blurring the standards of authenticity and forcing audiences into a state of non-judgmental reading, thereby deepening the publics cognitive crisis. This paper aims to provide a critical perspective for the development and theoretical study of AIGC news.
Poor relationship quality common among individuals with borderline personality disorder (BPD) may result, in part, from biased interpersonal decision-making. We examined memory biases for hypothetical interpersonal partner choices varying in the degree of familiarity. In Part 1 of our study, participants (n = 192) were asked to choose between novel or familiar partners based on lists of traits across six vignettes, and in Part 2, they completed a trait recognition task 36-60 hours later. Lower perceived social support was associated with a memory bias toward novel (over familiar) partners. BPD features were negatively related to an overall interpersonal memory bias (i.e., remembering both partners more negatively). However, when accounting for idiographic valence ratings, BPD features were positively related to this bias among those also low in social support. Memory biases may be related to partner choices associated with BPD features; however, it is critical to assess the role of perceived social support.
This paper presents the creation of a Universal Dependency (UD) treebank for Amahuaca (Peru), marking the first UD treebank within the Headwaters subbranch of the Panoan family, spoken mostly in Peru and Brazil. While the UD guidelines provided a general framework for our annotations, language-specific decisions were necessary due to the rich morphology of the Amahuaca language. The paper also describes specific constructions to initiate a discussion on several general UD annotation guidelines, particularly those concerning clitics and morpheme-level dependencies.
UD_Nheengatu-CompLin is the highest-rated and second-largest Amerindian language treebank in Universal Dependencies. It is annotated with Yauti, an analyzer for Nheengatu that uses special tags to handle unknown words. This paper presents a major revision of the special tag mechanism, extending coverage to phenomena such as reduplication, typos, and stylistic variation. A multi-level validator was implemented and passed 416 test cases. All 231 treebank sentences with special tags parsed, yielding a macro F1 score of 0.92 and both weighted and micro F1 scores of 0.96 in feature assignment.
There are limited discussions on how translanguaging practices may be tailored according to needs and contexts by providing examples of the implementation of translanguaging in different countries. Thus, this chapter reports on implementing translanguaging premises and pedagogies in light of critical needs analysis to compare and offer practical recommendations. Türkiye (at a state university where medical English was delivered using CLIL), Brazil (at a state university where linguists and computer scientists develop annotated treebanks of diverse dialects to be used in Natural Language Processing applications, among them a language used by the Warao refugees in Brazil, emerging from the contact between the Warao language, Venezuelan Spanish, and Brazilian Portuguese). Here, the target situation is a site for possible transformation and a translanguaging approach allows. To the knowledge of this chapter's authors, this study in two different contexts is the first of its kind on translanguaging.
Classical Arabic (CA) presents unique and significant challenges for computational syntactic analysis, primarily due to its complex morphosyntax, pervasive ellipsis, flexible word order, and the critical scarcity of comprehensive annotated resources. This paper introduces Noor, a novel multi-stage pipeline framework specifically engineered for robust, high-performance syntactic parsing of CA. The Noor framework integrates several key innovations: (1) a dedicated pre-parsing module (TAQDIR) employing a hybrid rule-based and heuristic approach for explicit ellipsis resolution; (2) parallel dependency (ARC-IIRAB) and constituency (ARC-MAHAL) parsers utilizing Bidirectional Long Short-Term Memory (BiLSTM)-based representations and tailored transition systems; and (3) a synthesis mechanism generating enriched hybrid syntactic representations. A primary contribution of this work is the creation and public release of the Quranic Treebank—the first complete, validated, and fully machine-readable hybrid syntactic resource for the CA, reconstructed and augmented using the Noor framework. Empirical evaluation demonstrates that Noor achieves state-of-the-art (SOTA) performance on CA parsing, with the hybrid parser significantly outperforming previous benchmarks by over 4.6 F1 points on the Extended Labeled Attachment Score (ELAS). The framework is complemented by IIRAB Vis, a novel visualization tool designed to render complex hybrid structures in alignment with traditional Arabic grammatical principles (I’rāb). By providing both a validated SOTA methodology and a foundational public dataset, this research substantially advances the computational analysis capabilities for CA across Natural Language Processing (NLP), digital humanities, and linguistic studies.
Abstract The paper introduces Parallel Trees, a novel multilingual treebank collection that includes 20 treebanks for 10 languages. The distinguishing property of this resource is that the sentences of each language are annotated using two syntactic representation paradigms (SRPs), respectively based on the notions of dependency and constituency. By aligning the annotations of existing resources, Parallel Trees represents an example of exploiting pre-existing treebanks to adapt them to novel applications. To illustrate its potential, we present a case study where the resource is employed as a benchmark to investigate whether and how BERT, one of the first prominent neural language models (NLMs), is sensitive to the dependency- and constituency-based approaches for representing the syntactic structure of a sentence. The case study results indicate that the model’s sensitivity fluctuates across languages and experimental settings. The unique nature of the Parallel Trees resource creates the prerequisites for innovative studies comparing dependency and phrase-structure trees, allowing for more focused investigations without the interference of lexical variation.
A variety of pruning methods have been introduced for over-parameterized Recurrent Neural Networks to improve efficiency in terms of power consumption and storage utilization. These advances motivate a new paradigm, termed `hyperpruning', which seeks to identify the most suitable pruning strategy for a given network architecture and application. Unlike conventional hyperparameter search, where the optimal configuration's accuracy remains uncertain, in the context of network pruning, the accuracy of the dense model sets the target for the accuracy of the pruned one. The goal, therefore, is to discover pruned variants that match or even surpass this established accuracy. However, exhaustive search over pruning configurations is computationally expensive and lacks early performance guarantees. To address this challenge, we propose a novel Lyapunov Spectrum (LS)-based distance metric that enables early comparison between pruned and dense networks, allowing accurate prediction of post-training performance. By integrating this LS-based distance with standard hyperparameter optimization algorithms, we introduce an efficient hyperpruning framework, termed LS-based Hyperpruning (LSH). LSH reduces search time by an order of magnitude compared to conventional approaches relying on full training. Experiments on stacked LSTM and RHN architectures using the Penn Treebank dataset, and on AWD-LSTM-MoS using WikiText-2, demonstrate that under fixed training budgets and target pruning ratios, LSH consistently identifies superior pruned models. Remarkably, these pruned variants not only outperform those selected by loss-based baseline but also exceed the performance of their dense counterpart.
Abstract In many fields, such as language acquisition, neuropsychology of language, the study of aging, and historical linguistics, corpora are used for estimating the diversity of grammatical structures that are produced during a period by an individual, community, or type of speakers. In these cases, treebanks are taken as representative samples of the syntactic structures that might be encountered. Generalizing the potential syntactic diversity from the structures documented in a small corpus requires careful extrapolation whose accuracy is constrained by the limited size of representative sub-corpora. In this article, I demonstrate—both theoretically and empirically—that a grammar’s derivational entropy and the mean length of the utterances (MLU) it generates are fundamentally linked, giving rise to a new measure, the derivational entropy rate. The mean length of utterances becomes the most practical index of syntactic complexity; I demonstrate that MLU is not a mere proxy, but a fundamental measure of syntactic diversity. In combination with the new derivational entropy rate measure, it provides a theory-free assessment of grammatical complexity. The derivational entropy rate indexes the rate at which different grammatical annotation frameworks determine the grammatical complexity of treebanks. I evaluate the Smoothed Induced Treebank Entropy (SITE) as a tool for estimating these measures accurately, even from very small treebanks. I conclude by discussing important implications of these results for both NLP and human language processing.
This study investigates the sentiment polarity (positive, negative, neutral) and specific emotions (joy, sadness, anger, surprise, trust, anticipation, disgust, and fear) expressed by Generation Z in digital platform comments regarding seven female duets with famous male singer. A dataset of 500 digital comments (250 from YouTube, 125 from Twitter, 125 from Instagram) was collected. The sample was then refined to include comments from 100 individuals (50 men, 50 women) affiliated with a private university in Mexico City, ensuring gender balance. Sentiment polarity was classified using a Bidirectional Encoder Representations from Transformers (BERT) model, with its hyperparameters (learning rate, epochs, batch size) optimized via a Particle Swarm Optimization (PSO) metaheuristic, leading to a 4% accuracy improvement over default settings. Emotion detection was performed concurrently using the NRC Emotion Lexicon, a lexical database mapping terms to eight emotional categories. Results indicate a clear correlation between musical tone and expressed sentiment: melancholic duets elicited predominantly negative sentiments, whereas more energetic collaborations generated a higher proportion of positive comments and the emotion 'joy.' Furthermore, significant differences using $\chi^2$ ($p < 0.01$) were observed in the distribution of 'anger' and 'sadness' between intimate and collaborative duets. These findings offer valuable insights into how Generation Z, segmented by gender, emotionally interprets this singer's musical productions. This research has significant implications for developing targeted music marketing strategies and content production for digitally native audiences.
“Pictures are worth a thousand words," yet most platforms like Yelp, Google Maps, Instagram, Walmart, and Amazon require users to provide text, ratings, and images. Images often capture a user's intent, and the features within the images typically correlate with that intent. In this paper, we extract various features from images (such as edge distribution, color distribution, text within the image, focus, etc.) and compare simple vs. complex models to predict the ratings associated with these images. We find that features such as brightness and contrast significantly explain the rating at image-level, and models such as random forest and logistic regression provide a 0.84 F-1 score when predicting the rating. In the era of generative AI, we anticipate that sharing an image will allow platforms to auto-generate user intent and image ratings, thereby simplifying the dissemination of information.
Although aerobic exercise modulates self-experienced pain, its impact on empathy for pain remains unclear. Moreover, whether exergaming, which combines exercise with interactive gaming, influences empathy-related neural responses is unknown. The present study investigated the effects of exergaming on the neural mechanisms underlying empathy for pain, comparing them with those of conventional aerobic exercise (cycling) and a non-active control condition. A total of ninety-one participants were randomly assigned to one of three conditions: exergaming (Nintendo Fitness Ring Adventure), moderate intensity cycling, or rest. After a 30-min intervention, participants completed a pain judgement task while event-related potentials (ERP) were recorded. Behavioral outcomes (reaction time, accuracy, pain intensity, and emotional valence ratings) and ERP components (N1, P2, N2, P3, LPP) were analyzed. Results revealed that both exergaming and cycling enhanced emotional valence ratings for painful images relative to the control condition. ERP analyses demonstrate that exergaming significantly amplified late-stage components (P3 and LPP) in response to painful stimuli, indicating enhanced cognitive appraisal processes associated with empathy for pain, while early components (N1, P2, N2) remain unaffected across conditions. These findings suggest that exergaming, through its combination of multisensory and cognitive engagement, uniquely enhances cognitive empathy for pain.
Banua Language is one of the languages originating from the province of East Kalimantan. Banua Language is part of the Austronesian language family. This research on Banua Language uses a syntactic typology approach. The aim of this study is to explain the behavior of subjects in clause construction within Banua Language. As a regional language in East Kalimantan, Banua Language has a unique syntactic structure. Not much research has been done on the function and position of subjects in this language. Therefore, this study examines subjecthood through canonical position, relativization, and control, the three main tools according to Keenan and Comrie (1983). The distributional method was used to analyze data collected from native speakers through elicitation and uninvolved observation techniques. The study shows that in Banua Language, the subject always precedes the verb in both intransitive and transitive clauses. Additionally, the subject can be relativized exclusively using the form "anu". Furthermore, the subject can be controlled in serial verb constructions, but arguments without subjects do not exhibit the same characteristics. The findings indicate that Banua Language has subjecthood properties that align with universal syntactic principles and contribute to the typological mapping of Indonesian languages. The theoretical implication of this research is expected to strengthen the theory of subjecthood from a syntactic typology perspective and provide a global linguistic database. In addition, the practical implication expected from this research is to serve as a means for preserving the regional language and a foundation for developing the Banua Language grammar.
This paper introduces the S M Nazmuz Sakib Cohesion Conservation Principle for English–Bangla translation and formally defines the associated Sakib Cohesion Constant. The central hypothesis is that, for natural and high-quality translations between English and Bangla, a weighted scalar measure of textual cohesion is approximately conserved across the source and target, even though the distribution of cohesive devices (lexical, morphological, and discourse-particle-based) differs substantially between the two languages. Building on Halliday and Hasan’s theory of cohesion in English and Blum-Kulka’s work on shifts of cohesion in translation, the present work extends cohesion analysis by (i) distinguishing three channels of cohesive realization—lexical/grammatical connectives, morphologically encoded agreement and argument marking, and discourse particles—and (ii) proposing a quantitative conservation law specific to the English–Bangla language pair. To ground the proposal empirically, we use public datasets: large-scale parallel corpora such as the Samanantar English–Indic corpus, the SUPara and EMILLE English–Bangla corpora, a multi-source BanglaNMT training corpus, the DiMLexBangla lexicon of Bangla discourse connectives, and the Bangla RST Discourse Treebank. From these, we derive corpus-level indicators of cohesion (lexical diversity, connective inventory size, distribution of text types, and asymmetries in tokens per sentence) and estimate a Sakib Cohesion Constant as a cross-corpusstable ratio capturing how English and Bangla distribute cohesive load. Twenty data-based figures and several summary tables are provided to illustrate the empirical behaviour of the proposed constant and its relevance for translation studies, discourse analysis, and machine translation evaluation.
Rapid Eye Movement Sleep (REM) is thought to process emotions via memory reactivation. Such REM reactivation can be triggered by presenting a tone associated with the target memory. This reduces subjective arousal ratings for negative stimuli. Here, we measure arousal objectively in brain and autonomic system. Participants rated negative image-sound pairs, half of which were then re-presented during subsequent REM. All images were re-rated in a Magnetic Resonance Imaging (MRI) scanner with pulse oximetry 48 h after encoding. Reactivation in REM reduced responses in the brain's Salience Network (SN), including Anterior Insula and dorsal Anterior Cingulate Cortex (dACC), and associated emotion-processing regions: orbitofrontal cortex, subgenual cingulate, and left amygdala. Memory reactivation in REM reduced heart rate deceleration (HRD). Subjective arousal ratings were reduced for more upsetting images and increased for less upsetting images. Our findings have implications for the use of memory reactivation to treat depression and anxiety disorders.
We are using a Wikibase instance (https://lilamorph.wikibase.cloud) for publishing a Latin verb forms dataset, with the final goal of enriching Wikidata Latin lexemes, and for corpus annotation (matching tokens in morphologically annotated corpora to Wikibase forms). Building on the PrinParLat lexicon of Latin verb principal parts, we generate the complete set of inflected forms for over 8,000 verbs, encoded as RDF in a dedicated Wikibase instance. These data are linked to the Index Thomisticus Treebank (ITTB), whose morphologically annotated tokens are related to corresponding forms based on segmental identity, lemma alignment, and mapped morphological features. Our method achieves over 95% coverage of ITTB verbal tokens, demonstrating the robustness of our generation and linking pipeline even for Medieval Latin data. By aligning Paralex, Wikidata, and LiLa ontologies, we ensure semantic interoperability and facilitate future integration into Wikidata. Beyond Latin, this workflow provides a reproducible model for linking inflectional paradigms and corpus attestations in other languages. With the different forms lexica built on our Wikibase instance, we are now in the position to contribute to a discussion in the Wikidata community, comparing different options of representation of inflected forms. We would like to highlight corpus token linking as central use case for Wikibase forms, which entails to adopt the data model that caters best for that application, namely a separate listing of orthographically identical but morphologically ambiguous forms. Having chosen Wikibase as platform for the experiments presented here, all datasets remain now ready for intervention of human or algorithmic users, who would mark ambiguous links (from token to form, or from token to lila lemma), as “preferred” or “deprecated”, so that the ambiguity is resolved.
The intricate interplay between visual perception and emotion determines how waking experience influences mentation through a 'day residue' at once conspicuous yet hard to predict. Here we set out to map the neural sources associated with the visuo-affective processing of the 'day residue' during hypnagogic sleep. To this end, we assessed 28 healthy participants on a combined nap protocol with serial awakenings, pre-sleep stimulation with affective visual images, yoked measures of the semantic similarity between image and imagery reports, affect ratings, estimation of 64-channel EEG sources, and functional connectivity analysis. Overall, low-frequency EEG power was associated with weaker residues, and high-frequency EEG power was associated with stronger residues. The source networks most significantly correlated with imagetic and affective residues were markedly different across wake-sleep states, partially overlapping with the default mode network during N1 for up to 50% and 61%, respectively. The results allowed us to identify neural correlates of the visuo-affective processing of the day's residue, showing that the hypnagogic processing of the waking experience involves complex, dynamic and sequential bi-hemispheric interactions among multiple cortical, subcortical, and cerebellar structures with visual, limbic, optokinetic, and cognitive functions.
Internet overuse is a widespread phenomenon in today's digital society. Existing interventions, such as time limits or grayscaling, often rely on restrictive controls that provoke psychological reactance and are frequently circumvented. Building on prior work showing that emotional responses mediate the relationship between content consumption and online engagement, we investigate whether regulating the emotional impact of images can reduce online use in a non-coercive manner. We introduce and systematically analyze three regressor-guided image-editing approaches: (i) global optimization of emotion-related image attributes, (ii) optimization in a style latent space, and (iii) a diffusion-based method using classifier and classifier-free guidance. While the first two approaches modify low-level visual features (e.g., contrast, color), the diffusion-based method enables higher-level changes (e.g., adjusting clothing, facial features). Results from a controlled image-rating study and a social media experiment show that diffusion-based edits balance emotional responses and are associated with lower usage duration while preserving visual quality.
Previous research regarding verb production deficits in Alzheimer’s disease (AD) primarily concentrated on either the quantity of verbs (inflections) or verb-related semantic units, with little consideration given to verb production within syntactic contexts, i.e., verb collocations. This study explored verb collocations in the connected speech of Chinese AD patients within the framework of dependency syntax. The findings include: (1) The frequency distribution of verb collocation patterns in AD follows the Mixed-Poisson function similar to that in the healthy control elderly (HCE) and healthy control young (HCY) groups, but it differs in the use of low- and high-collocation patterns; (2) In the static aspect, the AD patients exhibit the lowest overall mean collocation pattern (MCP) among the three treebanks, followed by the HCE group. In the dynamic aspect, the MCP and sentence length in the three groups show a similar synergistic relation, but differences exist in the quadratic regression parameters; (3) Based on the probabilistic distribution of verb-governed dependencies, the AD patients exhibit the lowest syntactic proficiency, followed by the HCE group. The differences between the AD patients and the HCE group confirm the presence of verb production deficits and a decline in syntactic proficiency in AD. Although the HCE group also shows mild language deterioration compared to the HCY group, the extent of these changes is considerably smaller than that observed in the AD patients. These findings suggest that while aging may contribute to a partial decline in language abilities, AD markedly exacerbates and accelerates this deterioration process, following a pathological trajectory distinct from normal aging.
The development of lexicalized grammars, particularly Tree-Adjoining Grammar (TAG), has significantly advanced our understanding of syntax and semantics in natural language processing (NLP). While existing syntactic resources like the Penn Treebank and Universal Dependencies offer extensive annotations for phrase-structure and dependency parsing, there is a lack of large-scale corpora grounded in lexicalized grammar formalisms. To address this gap, we introduce TAGbank, a corpus of TAG derivations automatically extracted from existing syntactic treebanks. This paper outlines a methodology for mapping phrase-structure annotations to TAG derivations, leveraging the generative power of TAG to support parsing, grammar induction, and semantic analysis. Our approach builds on the work of CCGbank, extending it to incorporate the unique structural properties of TAG, including its transparent derivation trees and its ability to capture long-distance dependencies. We also discuss the challenges involved in the extraction process, including ensuring consistency across treebank schemes and dealing with language-specific syntactic idiosyncrasies. Finally, we propose the future extension of TAGbank to include multilingual corpora, focusing on the Penn Korean and Penn Chinese Treebanks, to explore the cross-linguistic application of TAG's formalism. By providing a robust, derivation-based resource, TAGbank aims to support a wide range of computational tasks and contribute to the theoretical understanding of TAG's generative capacity.