This study proposes a dual-stream late fusion approach to dimensional emotion recognition by fusing facial landmarks and electrocardiogram data. Unlike category-based methods, the proposed method identifies emotions in the valence-arousal space. By individually optimizing the classifiers for both modalities and then fusing their outputs, the model encourages flexibility, interpretability, and robustness. The ASCERTAIN dataset is used, which consists of facial landmark trajectories, physiological signals like electroencephalograms, electrocardiograms, galvanic skin responses, and electromyograms, and self-reported arousal and valence ratings of fifty-eight subjects who viewed thirty-six videos. Arousal and valence were each modeled using an individual classifier. The random forest-based feature selection achieved eighty-six point ninety-one percent accuracy in valence prediction, while a soft voting ensemble of support vector machines, K-nearest neighbors, and random forest achieved sixty-one point eighteen percent for arousal. These results were merged in order to classify four emotional quadrants: High Arousal High Valence, High Arousal Low Valence, Low Arousal High Valence, and Low Arousal Low Valence, with an aggregate accuracy of sixty-nine point sixteen percent. The findings indicate that the model introduced accurately makes use of complementary modalities and a late fusion strategy, and is suitable for real-world emotion-aware systems.