Over the past 15 years, there has been increasing use of linguistically annotated sentence collections such as the LDC Penn Tree Bank (PTB) for constructing statistically based parsers. While these parsers have generally been built for engineering purposes, more recently such approaches have been advanced as potential cognitive solutions, e.g., for the problem of human language acquisition. Here we examine this possibility critically: we assess how well these Treebank parsers actually approach human/child language competence. We find that such systems fail to replicate many, perhaps most, empirically attested grammaticality judgments; seem overly sensitive, rather than robust, to training data idiosyncrasies; and easily acquire unnatural syntactic constructions never attested in human languages. Overall, we conclude that existing statistically based treebank parsers fail to incorporate much knowledge of language in these three senses. We discuss the implications of these results for the improvement of Treebank parsers and their cognitive relevance.