Automated item generation (AIG) using large language models (LLMs) is emerging as a promising solution for scalable creativity assessments, yet the potential of AIG to create psychometrically rigorous creativity items has thus far received little attention. In the present study, we evaluated the psychometric properties of LLM-generated items for the classic Consequences task of divergent thinking, comparing them to human-written items. A sample of 192 undergraduates completed six Consequences tasks—two human-written, two standard LLM-generated, and two genre-inspired LLM-generated, drawing on sci-fi and fantasy themes—along with measures of personality, cognitive ability, and creative behavior. We found that LLM-generated items elicited significantly higher fluency and flexibility. Genre-inspired items (fantasy and sci-fi) additionally elicited higher originality, enjoyment, and positive valence ratings; LLM-generated items were similar to human-written items. Genre-inspired items did not unfairly advantage participants with more exposure to genre content, demonstrating measurement invariance. Together, these findings suggest that LLMs can generate consequences items with psychometric properties comparable to those of human-written items. We discuss the implications for creativity assessment, item bank development, and best practices for AIG in creativity research, and we provide the prompts we used for item generation to enable researchers to generate new Consequences items at scale.