Abstract Unintended misalignment in LLMs is highlighting the need to examine not only LLMs’ explicit outputs but also the functional representations of models that may shape their behavior. Building on this perspective, the present study explored whether two multimodal LLMs (MLLMs), GPT and Gemini, may form representations of their functional identity that remain across contexts. To this end, we adapted the reverse-correlation (RC) method and generated personified classification images (personified-CIs) based on the human face images that ChatGPT and Gemini selected as better reflecting their own image. The findings were as follows. First, across two RC tasks conducted one week apart, the temporal stability of the personified-CIs of both MLLMs was partially supported. Second, both ChatGPT and Gemini rated their own personified-CIs as more self-resembling than randomly generated filler-CIs. Third, both models rated their personified-CIs as higher in positive than negative valence and assigned higher valence ratings to their own personified-CIs than to filler-CIs. These findings provide preliminary evidence for the possibility that GPT and Gemini may form representations of their functional identity, suggesting that such representations warrant closer monitoring as LLMs continue to advance toward AGI capabilities and expand their domains of application.