While Out-of-distribution (OOD) detection has been well explored in computer\nvision, there have been relatively few prior attempts in OOD detection for NLP\nclassification. In this paper we argue that these prior attempts do not fully\naddress the OOD problem and may suffer from data leakage and poor calibration\nof the resulting models. We present PnPOOD, a data augmentation technique to\nperform OOD detection via out-of-domain sample generation using the recently\nproposed Plug and Play Language Model (Dathathri et al., 2020). Our method\ngenerates high quality discriminative samples close to the class boundaries,\nresulting in accurate OOD detection at test time. We demonstrate that our model\noutperforms prior models on OOD sample detection, and exhibits lower\ncalibration error on the 20 newsgroup text and Stanford Sentiment Treebank\ndataset (Lang, 1995; Socheret al., 2013). We further highlight an important\ndata leakage issue with datasets used in prior attempts at OOD detection, and\nshare results on a new dataset for OOD detection that does not suffer from the\nsame problem.\n