Although generic (i.e. domain independent) and specialized (i.e. domain specific) lexical resources are usually developed with different aims, an integrated consultation seems to be necessary for many NLP based applications. We describe an integration procedure based on the definition of plug-in relations that are established to manage overlaps and inconsistencies between the two resources. The approach has been experimented connecting ItalWordNet, a generic lexical database for Italian, and Economic-WordNet, a specialized wordnet for the economic and financial domain. In this paper we address the issue of integrating the information included in a generic lexical database with the information included in a specialized (i.e. domain specific) lexical database. We restrict our investigation to wordnet-like lexical resources, that is lexical databases whose model is derived from WordNet (Miller, 1990) (Fellbaum, 1998). The aim is to define a set of procedures to allow an of the two resources, such that overlapping senses are merged and conflicting situations are properly managed. Our starting point are two existing wordnets for Italian: ItalWordnet (IWN)1, a generic wordnet deriving from EuroWordNet (Vossen, 1999), and Economic-wordnet (ECOWN), a wordnet for the economic and financial domain. The current application scenario is a recommender system (Magnini and Strapparava, 2001), which includes a 1 This research has been supported by SI-TAL (Integrated System for the Automatic Treatment of Language), a National Project devoted to the creation of large linguistic resources and software for Italian written and spoken language processing. document processing module based on a light form of word sense disambiguation. The experimental domain is that of financial news. As a first attempt to use the two resources in conjunction, a “specific-first” strategy was implemented, which, given a lemma in the document, first looks up in the specific wordnet and just in case of failure resorts to the generic wordnet. An evident drawback of this approach is that it does not cover cases of words belonging to both of the databases when the correct interpretation of the word is placed in the generic wordnet. For instance, in the phrase “building its share market” (taken from a financial news), the word “share” is used with the generic meaning of “portion”, while in the specialized wordnet we would have the economic sense of “share” as part of the capital stock. To overcome this limit we tried with a “union” strategy which, for a given lemma, considers the sum of the senses for that lemma in the two resources. The problem here is that senses for words belonging to both the databases tend to proliferate, making disambiguation harder. What seems necessary is a deeper integration of the respective word senses, such that overlapping senses are merged and conflicting situations are detected and solved. Our scenario allows some significant simplifications with respect to the general problem of merging two distinct ontologies (Hovy, 1998). On the one side we have a specialized database, whose content is supposed to be more accurate and precise as far as specialized information is concerned; on the other side we can assume that the generic resource guarantees a more uniform coverage as far as high level senses are concerned. These two assumptions provide us with a powerful precedence criterion to be used for managing inheritance in the integration procedure. The contribution of our work consists in a!# $ % &(' ) approach which allows the connecting of the generic and the specialized wordnets in a flexible and modular way. This is realized by means of a semi-automatic procedure with four main steps: (i) first, a minimum set of specialized “basic synsets” is identified; (ii) basic synsets are aligned to corresponding generic synsets and a particular plug-in configuration is selected; (iii) for each plug-in configuration a merging algorithm reconstructs the corresponding portion of the integrated wordnet; (iv) possible inconsistencies are solved. There are two main benefits of this approach. First, already existing specialized resources can be connected to a generic resource without any change in the resource being necessary, a part from the data conversion into a wordnet-like format. Second, the inheritance of linguistic oriented information makes the specialized resources usable in existing wordnet-based applications. The paper is structured as follows. Section 2 presents the lexical resources we have used. Section 3 introduces the basic notions of the plug-in approach with some technical details. Section 4 describes the plug-in procedure and reports the results of an application of the approach. Section 5 places our proposal in the context of related works. * +-,., / 0 132.5476 8,91 0 2;:, 4@?BA / 45., C D In this work we assume a taxonomic wordnetlike structure of the lexicon, where nodes in the hierarchy are synsets (i.e. synonym sets), and a rather large set of conceptual relations (e.g. Part-of, Cause, Hypernymy, Pertains-to, etc.) are available to build a semantic net among synsets. The EuroWordNet model has been adopted, which is rich enough to encompass most of the relations used in existing non wordnet-like terminological databases. We focus on the integration of already existing generic and specialized wordnets; both the data acquisition modalities and the evaluation of the quality of the resources do not affect our approach. As for generic lexical databases, typically they contain knowledge with no specific coverage and much attention is placed on coding linguistic-oriented information, such as subcategorization relations and fine-grained sense distinctions. There are several examples of existing generic lexical databases, including the English WordNet (Miller, 1990), monolingual wordnets for several European languages (e.g. Dutch, Spanish, Italian, Basque, etc.), the SENSUS (Hovy, 1998) and the Mikrokosmos (Mahesh, 1996) ontologies. Specialized databases focus on a certain domain, providing sub-hierarchies of highly specialized concepts with a limited use of lexical and linguistic relations. Synset variants tend to assume the shape of complex terms (i.e. multiwords) and the role of the domain expert is crucial for establishing correct relations. In addition high level knowlegde (i.e. the top ontology) tend to be simplified and domain oriented. Many specialized lexical databases have been developed, particularly for concrete applications, including, for example, a taxonomy for the medical domain (Gangemi et E F., 1999), the Art and Architecture Getty Thesaurus and the Getty Thesaurus of Geographical Names. The plug-in model we present has been applied within the SI-TAL project to connect a generic wordnet and a specialized wordnet that have been created independently. GIH J K LNM O P Q R H (Roventini et S T., 2000) created as part of the EuroWordNet project and further developed through the introduction of adjectives and adverbs, is the lexical database involved in the plug-in as a generic resource and consists of about 45,000 lemmas. U5V W X W Y@Z V []\^W _ ` acb d is a specialized wordnet for the economic domain and consists of about 5,000 lemmas distributed in about 4,700 synsets. Table 1 summarizes the quantitative data of the two resources considered.