Our study examines the integration of large language models (LLMs) on psycholinguistic research by prompting several LLMs to rate figurative expressions on familiarity, aptness, concreteness, metaphoricity, and constituency, with two manipulations: (1) stimuli presentation (in context versus in isolation) to examine whether LLMs benefit from the presence of context when rating familiarity—but not aptness—as observed in a recent study (Pissani and de Almeida, in press.) and (2) type of instructions (graded ratings versus categories) to examine whether LLMs perform better when assigning categories rather than numerical ratings, as observed in a recent study (Bavaresco et al., 2025). In addition, we will substitute LLM-produced ratings into existing studies of metaphor comprehension to determine whether they can predict the same effects observed with human ratings. Our study serves three purposes. First, we aim to examine whether LLM outputs replicate humans in rating tasks involving figurative language and whether they can serve for data augmentation when human ratings are unavailable. Second, we aim to understand whether LLMs process these ratings in the same way as humans (e.g., familiarity ratings may reflect both frequency of use and ease of understanding, while aptness ratings may reflect the quality of the metaphorical mapping, regardless of context). Third, we aim to provide best practices for obtaining reliable ratings from LLMs, whether using graded scales or categorical outcomes.