Normed linguistic stimuli are fundamental in psycholinguistics because they capture lexical and semantic properties that influence comprehension. However, generating these norms at scale is challenging, often leading researchers to rely on ad hoc norms collected from small samples, which can introduce inconsistencies and limit cross-study comparisons. In the present study, we investigated how large language models (LLMs) can support psycholinguistic research by prompting eight current LLMs to norm 300 English two-word metaphor combinations, such as sharp mind. We selected the dimensions of familiarity, aptness, concreteness, metaphoricity, and constituency, as these tap distinct cognitive processes and may provide insight into which aspects LLMs capture accurately and which they do not. We varied stimulus presentation (in context vs. in isolation) and response format (categorical vs. numerical) to examine which manipulation yields norms most closely aligned with human ratings. We then assessed the reliability and validity of model responses and used them to replicate existing analyses of metaphor comprehension. Overall, LLM-generated norms aligned best with familiarity and metaphoricity, which rely on word co-occurrence. In contrast, aptness, concreteness, and constituency—which require reasoning about the relationship between the topic (e.g., mind ) and the vehicle (e.g., sharp )—proved more challenging for LLMs.
No citing papers are currently in WordNorms