Depression, a pervasive mental health issue, highlights the critical need for early detection and effective intervention. A database (InStant-EMDB) is developed to analyze variations in the spoken responses of individuals with depression compared to healthy counterparts. Spoken responses are ob-tained from English and Malayalam bilingual speakers who respond spontaneously to a set of 15 emotionally evocative words. These words are sourced from the Affective Norms for English Words (ANEW) dataset, which includes valence and arousal ratings for each word. The speech in both English and Malayalam is manually transcribed to accurately reflect the spoken content. The dataset also contains self-reported af-fective ratings and data from a mental health survey (PHQ-9) collected from the participants to determine their mental state, which we considered the self-reported depression labels. A preliminary analysis is conducted on the collected speech using current state-of-the-art deep learning models such as Con-volutional Neural Networks (CNN), Long Short-Term Mem-ory networks (LSTM) and Bi-directional Long Short-Term Memory networks (Bi-LSTM). Although distinct linguistic patterns exhibited by individuals struggling with depression are successfully identified by all three models in both Malay-alam and English spoken responses, the highest accuracy is achieved by the LSTM models. Our findings and dataset em-phasize the potential of linguistic patterns as valuable cues for the early identification and intervention of depression and could contribute to enhancing accessibility for diverse populations.