Naturalistic paradigms provide ecologically valid insights into affective and cognitive processes but often require costly and time-consuming human annotations. Large language models (LLMs) provide a scalable tool for generating human-like affective ratings that could complement traditional behavioral approaches. In this study, we compared affective ratings of narrative segments obtained from young adults, five OpenAI's GPT models, Meta's Llama 3.1, and lexical-level norms from the SCOPE metabase. LLM-derived ratings of hedonic valence showed strong correlations with human ratings and outperformed lexical norms. When applied to fMRI data, LLM-derived ratings identified affective brain networks that substantially overlapped with those revealed by human ratings. These findings demonstrate that LLMs can approximate group-level affective ratings from young adults in naturalistic contexts and serve as a useful complement to traditional behavioral data collection, while underscoring the need for careful evaluation of their generalizability and potential biases.
No citing papers are currently in WordNorms