This resource contains a cross-linguistic lexical database comprising more than 2.6 million entries from over 8,000 languages and dialects worldwide. The dataset is enriched with Glottolog classification data and additional genealogical metadata, enabling large-scale comparative and typological analyses. A subset of the vocabulary has been grouped into higher-level semantic categories to facilitate macro-typological and semantic investigations. The present public extract includes only those components suitable for open dissemination. Additional internal structural layers — most notably the phonological slot structures (K1/K2/K3) — were designed for custom Python programs and must be extracted or reconstructed separately depending on the research focus. This work operates at the intersection of: genealogical and historical-comparative linguistics, etymological research, sound symbolism and sound–meaning correspondences, philosophy of language and epistemology. Special attention is given to the relationship between linguistic development and the development of consciousness, and to the question of how extensive empirical data may contribute to an epistemic approach to these foundational issues. A comprehensive monograph as well as several research papers are currently in preparation and will be linked here once available.