Recognising emotions as they naturally emerge in everyday contexts remains a key challenge in affective computing due to the inherent complexity, subtlety and variability of human emotional expression. While recent advances in multimodal deep learning have shown promising results, their effectiveness in real-world applications is often limited by the datasets used for training. Most existing emotion datasets are collected in controlled, acted scenarios, which tend to exaggerate expressions and fail to capture the subtle, low-intensity signals that reflect how emotions naturally unfold. To address this gap, we introduce EmU, a novel dataset grounded in spontaneous emotional expressions through autobiographical narratives. Expert consensus provides two levels of annotation, comprising seven discrete emotion categories alongside discretised valence ratings. We benchmark the dataset under both unimodal and multimodal settings using an attentive feature fusion framework, offering baseline performances to guide and facilitate future research.