In this paper, we combine methods used in penalized generalized empirical likelihood (GEL) frameworks with feature-extraction techniques that project textual data into numerical spaces. We develop recent penalization techniques used in GEL when the number of features gets very large compared to the sample size. We relate this approach to the Maximum Entropy (MaxEnt) principle used in several tasks of Natural Language Processing (NLP), in particular in Part Of Speech (POS) Tagging. Since the features belong to a large dimensional space, we propose a penalization method based on the dual representation of the original problem: this yields an explicit approximation of conditional probabilities of tags given the context. This method considerably reduces computational costs. As a byproduct, for different GEL methods, we obtain the corresponding POS-Tagging classifiers generalizing the MaxEnt method. We apply it successfully to the Penn-Treebank corpus with an error rate less than 5%.