This work investigates the capabilities of a multimodal large language model, Microsoft Copilot, to make emotion-related judgments for visual stimuli. We examine whether its ratings for affect and emotion are comparable to human ratings.