Recent years have witnessed significant improvement in ASR systems to\nrecognize spoken utterances. However, it is still a challenging task for noisy\nand out-of-domain data, where substitution and deletion errors are prevalent in\nthe transcribed text. These errors significantly degrade the performance of\ndownstream tasks. In this work, we propose a BERT-style language model,\nreferred to as PhonemeBERT, that learns a joint language model with phoneme\nsequence and ASR transcript to learn phonetic-aware representations that are\nrobust to ASR errors. We show that PhonemeBERT can be used on downstream tasks\nusing phoneme sequences as additional features, and also in low-resource setup\nwhere we only have ASR-transcripts for the downstream tasks with no phoneme\ninformation available. We evaluate our approach extensively by generating noisy\ndata for three benchmark datasets - Stanford Sentiment Treebank, TREC and ATIS\nfor sentiment, question and intent classification tasks respectively. The\nresults of the proposed approach beats the state-of-the-art baselines\ncomprehensively on each dataset.\n