Understanding and expressing emotions is subjective, leading to varying levels of ambiguity in how emotions are perceived. Humans naturally handle this ambiguity, and it's crucial for machines to do the same for natural human-machine interaction. Efforts are being made to develop emotion recognition systems that can handle ambiguity by modelling emotions as distributions of arousal/valence labels. However, existing approaches often assume that an underlying distribution can be inferred from ground truth ratings. Yet, the underlying distribution is never observed and the inferred distribution from the ratings can accurately represent the true distribution only when enough raters are involved. Our study investigates how the number of raters impacts the uncertainty in inferring distributions from ground truth ratings in the context of modelling emotion ambiguity. We then explore how many raters may be sufficient to effectively approximate the ambiguity present in the emotion annotations. Using the Belief Mismatch Coefficient (BMC) for analysis, we compare distributions to different numbers of ground truth ratings leading to quantitative and interpretable inferences. Experimental analysis was conducted using both simulated ratings and the real emotion (both arousal and valence) ratings collected from MSP-Conversation dataset. Our results suggest that using 5 raters may be sufficient to approximate a stable distribution representing the underlying emotional ambiguity when time continuous emotion ratings are captured and annotated from speech. Our study provides insights and practical advice for collecting emotion labels. Additionally, the analyses methods presented in this paper are expected to be relevant in any field of study dealing with labelling of subjective quantities that can vary from person to person.