Open-vocabulary image classification can recognize arbitrary textual categories but often fails to capture hierarchical relationships, for example, a Ragdoll cat should also be recognized as both cat and animal. Therefore, we introduce OpenHier, a comprehensive framework for open-vocabulary hierarchical classification that encompasses a hierarchical structure, a dataset, a benchmark, and a model. To capture hierarchical relationships, we construct a systematic real-world hierarchical structure based on the linguistic lexical database WordNet and Vision-Language Large Model. Building upon this structure, we develop a large-scale dataset with 4M annotated images, as well as a benchmark with 50K annotated images. Meanwhile, we design an evaluation pipeline to assess open-vocabulary hierarchical classification performance using several curated metrics. Furthermore, we propose a hierarchical consistency constraint method and multimodal alignment strategy to build a hierarchical classification model. Comprehensive experimental results demonstrate that our model achieves superior performance on both multi-label image classification and hierarchical image classification tasks in the open-vocabulary setting.