Abstract Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts, yet the extent to which these systems faithfully convey intended emotions remains largely underexplored. In this study, we introduce AImoclips, a benchmark designed to evaluate emotion conveyance to human listeners in TTM systems using a dimensional valence-arousal framework. We constructed the dataset with 991 instrumental music clips from six TTM systems prompted with 12 emotion words spanning four valence-arousal quadrants. A total of 111 participants provided 6,162 valence and arousal ratings using a 9-point scale. The results revealed that all systems perform above chance in conveying quadrant-level emotion intent, yet overall accuracy remains limited. Notably, substantial model-dependent biases were present. Commercial systems tended to exhibit positive valence deviations, whereas open-source models more often produced outputs with a negative valence shift, with diverging arousal shifts. Furthermore, overall audio quality (Fréchet Audio Distance (FAD)) correlates with both valence and arousal, while text-audio alignment (Contrastive Language-Audio Pretraining (CLAP) score) primarily relates to valence. These findings highlight underlying challenges in conveying linguistic emotional semantics precisely through musically conveyed affect for future research on controllable and perceptually consistent music generation.