This article discusses the issues of forming a database of phraseological units based on the Uzbek language corpus. A linguistic database is a structured collection of information that stores words, phrases, grammatical forms, idioms and other linguistic elements, and is used in lexicography to create dictionaries and reference books, in natural language processing to support technologies related to automatic translation, speech recognition and text generation, in linguistic research to analyze the structure of the language, the change and use of language units, and in language education to create educational materials and tools for language learning.