1358 norm sets
Die vorliegende Dissertation beschreibt Aufbau und Funktionalit{\"{a}}t der bayerischen Dialektdatenbank BAYDAT. Die Datenbank fasst die Erhebungsdaten der Teilprojekte des Bayerischen Sprachatlas (BSA) zusammen, speichert sie zukunftssicher und macht sie zentral nutzbar. Die Arbeit zeigt die Vorgehensweise bei der Aufbereitung der Quelldateien, beschreibt die einzelnen Datenbanktabellen der BAYDAT-Datenbank und widmet sich der Realisierung und Funktionalit{\"{a}}t der Onlineoberfl{\"{a}}che, {\"{u}}ber die die BAYDAT-Datenbank einem weltweiten Nutzerkreis aus Dialektologen und interessierten Laien zur Verf{\"{u}}gung stehen soll. The PhD thesis at hand describes the setup and functionality of the Bavarian dialect database BAYDAT. The database integrates the data of the subprojects of the Bayerischer Sprachatlas (BSA, Atlas of the Bavarian language). The database ensures a future-proof storage of the data and makes the data available at one central point. The thesis describes the methods used in formatting the source data. It also describes the structure of the different database tables. It also contains a description of the development and the functionality of BAYDAT's graphical user interface that will allow linguists and laypersons to access the database online.
India is a multilingual country where machine translation and cross lingual search are highly relevant problems. These problems require large resources- like wordnets and lexicons- of high quality and coverage. Wordnets are lexical structures composed of synsets and semantic relations. Synsets are sets of synonyms. They are linked by semantic relations like hypernymy (is-a), meronymy (part-of), troponymy (manner-of) etc. IndoWordnet is a linked structure of wordnets of major Indian languages from Indo-Aryan, Dravidian and Sino-Tibetan families. These wordnets have been created by following the expansion approach from Hindi wordnet which was made available free for research in 2006. Since then a number of Indian languages have been creating their wordnets. In this paper we discuss the methodology, coverage, important considerations and multifarious benefits of IndoWordnet. Case studies are provided for Marathi, Sanskrit, Bodo and Telugu, to bring out the basic methodology of and challenges involved in the expansion approach. The guidelines the lexicographers follow for wordnet construction are enumerated. The difference between IndoWordnet and EuroWordnet also is discussed.
In this work we present SENTIWORDNET 3.0, a lexical resource explicitly devised for supporting sentiment classification and opinion mining applications. SENTIWORDNET 3.0 is an improved version of SENTIWORDNET 1.0, a lexical resource publicly available for research purposes, now currently licensed to more than 300 research groups and used in a variety of research projects worldwide. Both SENTIWORDNET 1.0 and 3.0 are the result of automatically annotating all WORDNET synsets according to their degrees of positivity, negativity, and neutrality. SENTIWORDNET 1.0 and 3.0 differ (a) in the versions of WORDNET which they annotate (WORDNET 2.0 and 3.0, respectively), (b) in the algorithm used for automatically annotating WORDNET, which now includes (additionally to the previous semi-supervised learning step) a random-walk step for refining the scores. We here discuss SENTIWORDNET 3.0, especially focussing on the improvements concerning aspect (b) that it embodies with respect to version 1.0. We also report the results of evaluating SENTIWORDNET 3.0 against a fragment of WORDNET 3.0 manually annotated for positivity, negativity, and neutrality; these results indicate accuracy improvements of about 20{\%} with respect to SENTIWORDNET 1.0.