1358 norm sets
Spoken corpora are important for speech research, but are expensive to create and do not necessarily reflect (read or spontaneous) speech ‘in the wild’. We report on our conversion of the preexisting and freely available Spoken Wikipedia into a speech resource. The Spoken Wikipedia project unites volunteer readers of Wikipedia articles. There are initiatives to create and sustain Spoken Wikipedia versions in many languages and hence the available data grows over time. Thousands of spoken articles are available to users who prefer a spoken over the written version. We turn these semi-structured collections into structured and time-aligned corpora, keeping the exact correspondence with the original hypertext as well as all available metadata. Thus, we make the Spoken Wikipedia accessible for sustainable research. We present our open-source software pipeline that downloads, extracts, normalizes and text–speech aligns the Spoken Wikipedia. Additional language versions can be exploited by adapting configuration files or extending the software if necessary for language peculiarities. We also present and analyze the resulting corpora for German, English, and Dutch, which presently total 1005 h and grow at an estimated 87 h per year. The corpora, together with our software, are available via http://islrn.org/resources/684-927-624-257-3/. As a prototype usage of the time-aligned corpus, we describe an experiment about the preferred modalities for interacting with information-rich read-out hypertext. We find alignments to help improve user experience and factual information access by enabling targeted interaction.
We present DEMoS (Database of Elicited Mood in Speech), a new, large database with Italian emotional speech: 68 speakers, some 9 k speech samples. As Italian is under-represented in speech emotion research, for a comparison with the state-of-the-art, we model the ‘big 6 emotions’ and guilt. Besides making available this database for research, our contribution is three-fold: First, we employ a variety of mood induction procedures, whose combinations are especially tailored for specific emotions. Second, we use combinations of selection procedures such as an alexithymia test and self- and external assessment, obtaining 1,5 k (proto-) typical samples; these were used in a perception test (86 native Italian subjects, categorical identification and dimensional rating). Third, machine learning techniques—based on standardised brute-forced openSMILE ComParE features and support vector machine classifiers—were applied to assess how emotional typicality and sample size might impact machine learning efficiency. Our results are three-fold as well: First, we show that appropriate induction techniques ensure the collection of valid samples, whereas the type of self-assessment employed turned out not to be a meaningful measurement. Second, emotional typicality—which shows up in an acoustic analysis of prosodic main features—in contrast to sample size is not an essential feature for successfully training machine learning models. Third, the perceptual findings demonstrate that the confusion patterns mostly relate to cultural rules and to ambiguous emotions.
Definitional knowledge has proved to be essential in various Natural Language Processing tasks and applications, especially when information at the level of word senses is exploited. However, the few sense-annotated corpora of textual definitions available to date are of limited size: this is mainly due to the expensive and time-consuming process of annotating a wide variety of word senses and entity mentions at a reasonably high scale. In this paper we present SenseDefs, a large-scale high-quality corpus of disambiguated definitions (or glosses) in multiple languages, comprising sense annotations of both concepts and named entities from a wide-coverage unified sense inventory. Our approach for the construction and disambiguation of this corpus builds upon the structure of a large multilingual semantic network and a state-of-the-art disambiguation system: first, we gather complementary information of equivalent definitions across different languages to provide context for disambiguation; then we refine the disambiguation output with a distributional approach based on semantic similarity. As a result, we obtain a multilingual corpus of textual definitions featuring over 38 million definitions in 263 languages, and we publicly release it to the research community. We assess the quality of SenseDefs’s sense annotations both intrinsically and extrinsically on Open Information Extraction and Sense Clustering tasks.
Algeria’s socio-linguistic situation is known as a complex phenomenon involving several historical, cultural and technological factors. However, there are three languages that are mainly spoken in Algeria (Arabic, Tamazight and French) and they can be mixed in the same sentence (code-switching). Moreover, there are several varieties of dialects that differ from one region to another and sometimes within the same region. This paper aims to provide a new multi-purpose parallel corpus (i.e., DZDC12 corpus), which will serve as a testbed for various natural language processing and information retrieval applications. In particular, it can be a useful tool to study Arabic–French code-switching phenomenon, Algerian Romanized Arabic (Arabizi), different Algerian sub-dialects, sentiment analysis, gender writing style, machine translation, abuse detection, etc. To the best of our knowledge, the proposed corpus is the first of its kind, where the texts are written in Latin script and crawled from Facebook. More specifically, this corpus is organised by gender, region and city, and is transliterated into Arabic script and translated into Modern Standard Arabic. In addition, it is annotated for emotion detection and abuse detection, and annotated at the word level. This article focuses in particular on Algeria’s socio-linguistic situation and the effect of social media networks. Furthermore, the general guidelines for the design of DZDC12 corpus are described as well as the dialects clustering over the map.
The aim of word sense disambiguation (WSD) is to correctly identify the meaning of a word in context. All natural languages exhibit word sense ambiguities and these are often hard to resolve automatically. Consequently WSD is considered an important problem in natural language processing (NLP). Standard evaluation resources are needed to develop, evaluate and compare WSD methods. A range of initiatives have lead to the development of benchmark WSD corpora for a wide range of languages from various language families. However, there is a lack of benchmark WSD corpora for South Asian languages including Urdu, despite there being over 300 million Urdu speakers and a large amounts of Urdu digital text available online. To address that gap, this study describes a novel benchmark corpus for the Urdu Lexical Sample WSD task. This corpus contains 50 target words (30 nouns, 11 adjectives, and 9 verbs). A standard, manually crafted dictionary called Urdu Lughat is used as a sense inventory. Four baseline WSD approaches were applied to the corpus. The results show that the best performance was obtained using a simple Bag of Words approach. To encourage NLP research on the Urdu language the corpus is freely available to the research community.
This paper introduces Emilia, a speech corpus created to build a female voice in Spanish spoken in Buenos Aires for the Aromo text-to-speech system. Aromo is a unit selection text-to-speech system, which employs diphones as units of synthesis. The key requirements and design criteria for Emilia were: to synthesize any text in Spanish into high-quality speech with a minimum corpus size. The text corpus was designed to guarantee the phonetic and prosodic coverage. A three-stage strategy was used: in the first stage, 741 sentences were designed with all of the syllables of Spanish spoken in Argentina, with and without stress, and in all positions within the word; in the second stage, 852 sentences were added to balance out the distribution of the diphones; and after a perceptual evaluation of the quality of synthesized speech, in the third and final stage, 625 sentences were added to achieve the specified unit coverage, and to introduce sentences with more complex syntactic and prosodic structures. Issues from all three corpus building stages are reported. The paper also presents the results from the quality perceptual evaluations of the synthesized voice. Emilia has a duration of three hours and 15 minutes; its speech quality synthesized with Aromo system is similar to the level obtained with commercial systems, with a real-time ratio less than one.
The latest reference corpus of written Slovene, the Gigafida corpus, was created as part of the ‘Communication in Slovene’ project. In the same project, a web concordancer was designed for the broadest possible use, and tailored to the needs and abilities of user groups such as translators, writers, proofreaders and teachers. Two years after the corpus was published within the new tool, its features were assessed by the users. With an average rate of 4.36 on a scale between 1 and 5 (1 = I strongly disagree, 5 = I strongly agree), the results indicate that most survey participants agreed or strongly agreed with positive statements about the new implementations (e.g. “The corpus results are displayed in a clear manner”). This is a considerable improvement in user experience from the previous reference corpus of Slovene, i.e. the FidaPLUS corpus within the ASP32 concordancer (rated with 3.67). In the user feedback, the simplicity of search options and the interface clarity are highlighted as the main advantages, while for the future development, advanced visualizations of corpus data and improved search of word-phrases are suggested. The evaluation also highlighted some relevant user habits, such as not taking the time to learn systematically about the tool before they start using it. The findings will be implemented in future editions of the Gigafida corpus, but are relevant to any project that aims at facilitating a wider use of reference corpora and corpus-based resources.
The paper describes the creation of the first open access multi-genre historical corpus of Emergent Modern Hebrew, made possible by implementation of digital humanities methods in the process of corpus curation, encoding, and dissemination. Corpus contents originate in the Ben-Yehuda Project, an open access repository of Hebrew literature online, and in digital images curated from the collections of the National Library of Israel, a selection of which have been transcribed through a dedicated crowdsourcing task that feeds back into the library’s online catalog. Texts in the corpus are encoded following best practices in the digital humanities, including markup of metadata that enables time-sensitive research, linguistic and other, of the corpus. Evaluation of morphological analysis based on Modern Hebrew language models is shown to distinguish between genres in the historical variety, highlighting the importance of ephemeral materials for linguistic research and for potential collaboration with libraries and cultural institutions in the process of corpus creation. We demonstrate the use of the corpus in diachronic linguistic research and suggest ways in which the association it provides between digital images and texts can be used to support automatic language processing and to enhance resources in the digital humanities.
Around the world, a growing interest has been seen in learner translator corpora, which are invaluable resources for teaching and research. This paper introduces a new resource to support researchers from different interdisciplinary areas such as computational linguistics, descriptive translation studies, computer-aided translation technology, Arabic machine translation applications, cognitive science, and translation pedagogy. Motivated by the lack of learner translator resources that provide data about learners of translation from and into Arabic, the undergraduate learner translator corpus (ULTC) is an ongoing, error-tagged sentence-aligned parallel corpus of English, Arabic, and French, with Arabic as its main language. The present corpus, consisting of parallel texts of female learners of translation from English or French into Arabic, is the first of its kind in terms of the languages represented, tasks covered, and number of students involved. It is also unique in terms of combining many complementary corpora of cross-lingual data, each of which has its own web-based query interface and corpus analysis tools. This paper describes the ULTC compilation process, preliminary findings, and planned future expansion and research.
This paper introduces MadSex, a spoken corpus of 54 sociolinguistic interviews in Spanish based on the topic of sexuality. It was collected in order to study the cognitive sociolinguistic variation of sexual concepts. The paper presents and justifies methodological decisions taken during design, collection and transcription stages. Informants were selected in Madrid, based on a pre-stratified sample divided by sex, age and level of education. The interview methodology relied on an opinion questionnaire designed for the indirect elicitation of sexual concepts, which overcame successfully the limitations imposed by the low frequencies of semantic variables in discourse and the impact of sexual taboo in interaction. Relevant aspects of fieldwork, empathy and ethical protocols are also detailed in the paper. Transcription and markup are explained. Finally, an overview of the corpus is given, as well as some research papers based on it. Examples of the questionnaire are also provided.
Arabic is a widely-spoken language with a long and rich history, but existing corpora and language technology focus mostly on modern Arabic and its varieties. Therefore, studying the history of the language has so far been mostly limited to manual analyses on a small scale. In this work, we present a large-scale historical corpus of the written Arabic language, spanning 1400 years. We describe our efforts to clean and process this corpus using Arabic NLP tools, including the identification of reused text. We study the history of the Arabic language using a novel automatic periodization algorithm, as well as other techniques. Our findings confirm the established division of written Arabic into Modern Standard and Classical Arabic, and confirm other established periodizations, while suggesting that written Arabic may be divisible into still further periods of development.
Corpora play an important role when training machine learning systems for sentiment analysis. However, Spanish is underrepresented in these corpora, as most primarily include English texts. This paper describes 20 Spanish-language text corpora—collected to support different tasks related to sentiment analysis, ranging from polarity to emotion categorization. We present a brand-new framework for the characterization of corpora. This includes a number of features to help analyze resources at both corpus level and document level. This survey—besides depicting the overall landscape of corpora in Spanish—supports sentiment analysis practitioners with the task of selecting the most suitable resources.
The article details the formational process of the FinnTransFrame corpus, a part of the FinnFrameNet project. In addition to a large annotated frame semantic corpus of natural language examples, the project created a separate corpus of examples translated from English to Finnish. The research question when creating the FinnTransFrame corpus was to see to what extent the various frames of the original Berkeley FrameNet transfer into Finnish in translated examples, i.e. what are the main problems and how can they be categorized? A variety of Berkeley FrameNet examples were chosen from different frames and then translated by professionals. The FinnFrameNet annotation team checked all the examples and their translations to see if the frames remained intact in translation. Problematic examples were tagged according to the type of the encountered problem, with the main focus on the type of fine-grained mismatches of meaning that caused frame changes even when the translation was the best possible one. The frame-loss amounted to 4.2% of the 88,209 relevant example sentences. Filtering out sentences with other types of problems, we found that 88.1% of all the frame instances still translated into Finnish with their frame intact. In addition, the article analyzes the error types in the problematic frames.
What are known as specialized or specialist dictionaries are much more than lists of words and their definitions with occasional comments on things such as synonymy and homonymy. That is to say, a particular specialist term may be associated with many other concepts, including quotations, different senses, etymological categories, semantic categories, superordinate and subordinate terms in the terminological hierarchy, spelling variants, and references to background sources discussing the exact meaning and application of the term. The various concepts, in turn, form networks of mutual links, which makes the structure of the background concepts demanding to model when designing a database structure for this type of dictionary. The Dictionary of medical vocabulary in English, 1375–1550 is a specialized historical dictionary that covers the vast medical lexicon of the centuries examined. It comprises over 12,000 terms, each of them associated with a host of background concepts. Compiling the dictionary took over 15 years. The process started with an analysis of hand-written manuscripts and early printed books from different sources and ended with the electronic dictionary described in the present paper. Over these years, the conceptual structure, database schema, and requirements for essential use cases were iteratively developed. In our paper, we introduce the conceptual structure and database schema modelled for implementing an electronic dictionary that involves different use cases such as term insertion and linking a term to related concepts. The achieved conceptual model, database structure, and use cases provide a general framework for reference-oriented specialized dictionaries, including ones with a historical orientation.
Vowels in Arabic are optional orthographic symbols written as diacritics above or below letters. In Arabic texts, typically more than 97 percent of written words do not explicitly show any of the vowels they contain; that is to say, depending on the author, genre and field, less than 3 percent of words include any explicit vowel. Although numerous studies have been published on the issue of restoring the omitted vowels in speech technologies, little attention has been given to this problem in papers dedicated to written Arabic technologies. In this research, we present Arabic-Unitex, an Arabic Language Resource, with emphasis on vowel representation and encoding. Specifically, we present two dozens of rules formalizing a detailed description of vowel omission in written text. They are typographical rules integrated into large-coverage resources for morphological annotation. For restoring vowels, our resources are capable of identifying words in which the vowels are not shown, as well as words in which the vowels are partially or fully included. By taking into account these rules, our resources are able to compute and restore for each word form a list of compatible fully vowelized candidates through omission-tolerant dictionary lookup. In our previous studies, we have proposed a straightforward encoding of taxonomy for verbs (Neme in Proceedings of the international workshop on lexical resources (WoLeR) at ESSLLI, 2011) and broken plurals (Neme and Laporte in Lang Sci, 2013, http://dx.doi.org/10.1016/j.langsci.2013.06.002). While traditional morphology is based on derivational rules, our description is based on inflectional ones. The breakthrough lies in the reversal of the traditional root-and-pattern Semitic model into pattern-and-root, giving precedence to patterns over roots. The lexicon is built and updated manually and contains 76,000 fully vowelized lemmas. It is then inflected by means of finite-state transducers (FSTs), generating 6 million forms. The coverage of these inflected forms is extended by formalized grammars, which accurately describe agglutinations around a core verb, noun, adjective or preposition. A laptop needs one minute to generate the 6 million inflected forms in a 340-MB flat file, which is compressed in 2 min into 11 MB for fast retrieval. Our program performs the analysis of 5000 words/second for running text (20 pages/second). Based on these comprehensive linguistic resources, we created a spell checker that detects any invalid/misplaced vowel in a fully or partially vowelized form. Finally, our resources provide a lexical coverage of more than 99 percent of the words used in popular newspapers, and restore vowels in words (out of context) simply and efficiently.
This article describes the procedures employed during the development of the first comprehensive machine-readable Turkish Sign Language (TiD) resource: a bilingual lexical database and a parallel corpus between Turkish and TiD. In addition to sign language specific annotations (such as non-manual markers, classifiers and buoys) following the recently introduced TiD knowledge representation (Eryiğit et al. 2016), the parallel corpus contains also annotations of dependency relations, which makes it the first parallel treebank between a sign language and an auditory-vocal language.
Corpus-based research has formed the backbone of linguistic research in recent decades. Large text corpora are used for solving various kinds of linguistic problems, including those of quantitative linguistics, cognitive linguistics, and psycholinguistics. This paper reports the creation of two corpora of contemporary Vietnamese. It also describes the construction of these two equally sized Vietnamese corpora (a corpus from Vietnamese film subtitles, subtlex-viet, and a general corpus of varieties of online newspapers and stories, genlex-viet). We document the general steps of the construction and extraction of linguistic information from the language corpora and provide a road map for others who would like to create similar corpora. The resultant corpora are available in three versions: plain text, tokenized, and POS tagged. In the second half of the paper, the construction of a lexical database derived from the corpora is described. The database includes measures such as frequency of occurrence, dispersion, Mutual Information, Inverse Document Frequency, as well as vector space measures based on Latent Semantic Analysis and Hyperspace Analogue to Language. We conclude by reporting a comparison of the lexical predictors and a validation using psycholinguistic data from visual lexical decision experiments.
Studies on morphological processing in French, as in other languages, have shown disparate results. We argue that a critical and long-overlooked factor that could underlie these diverging results is the methodological differences in the calculation of morphological variables across studies. To address the need for a common morphological database, we present MorphoLex-FR, a sizeable and freely available database with 12 variables for prefixes, roots, and suffixes for the 38,840 words of the French Lexicon Project. MorphoLex-FR constitutes a first step to render future studies addressing morphological processing in French comparable. The procedure we used for morphological segmentation and variable computation is effectively the same as that in MorphoLex, an English morphological database. This will allow for cross-linguistic comparisons of future studies in French and English that will contribute to our understanding of how morphologically complex words are processed. To validate these variables, we explored their influence on lexical decision latencies for morphologically complex nouns in a series of hierarchical regression models. The results indicated that only morphological variables related to the suffix explained lexical decision latencies. The frequency and family size of the suffix exerted facilitatory effects, whereas the percentage of more frequent words in the morphological family of the suffix was inhibitory. Our results are in line with previous studies conducted in French and in English. In conclusion, this database represents a valuable resource for studies on the effect of morphology in visual word processing in French.
This article presents the development of the “Hoosier Vocal Emotions Corpus,” a stimulus set of recorded pseudo-words based on the pronunciation rules of English. The corpus contains 73 controlled audio pseudo-words uttered by two actresses in five different emotions (i.e., happiness, sadness, fear, anger, and disgust) and in a neutral tone, yielding 1,763 audio files. In this article, we describe the corpus as well as a validation study of the pseudo-words. A total of 96 native English speakers completed a forced choice emotion identification task. All emotions were recognized better than chance overall, with substantial variability among the different tokens. All of the recordings, including the ambiguous stimuli, are made freely available, and the recognition rates and the full confusion matrices for each stimulus are provided in order to assist researchers and clinicians in the selection of stimuli. The corpus has unique characteristics that can be useful for experimental paradigms that require controlled stimuli (e.g., electroencephalographic or fMRI studies). Stimuli from this corpus could be used by researchers and clinicians to answer a variety of questions, including investigations of emotion processing in individuals with certain temperamental or behavioral characteristics associated with difficulties in emotion recognition (e.g., individuals with psychopathic traits); in bilingual individuals or nonnative English speakers; in patients with aphasia, schizophrenia, or other mental health disorders (e.g., depression); or in training automatic emotion recognition algorithms. The Hoosier Vocal Emotions Corpus is available at https://psycholinguistics.indiana.edu/hoosiervocalemotions.htm.
The LENA system has revolutionized research on language acquisition, providing both a wearable device to collect day-long recordings of children’s environments, and a set of automated outputs that process, identify, and classify speech using proprietary algorithms. This output includes information about input sources (e.g., adult male, electronics). While this system has been tested across a variety of settings, here we delve deeper into validating the accuracy and reliability of LENA’s automated diarization, i.e., tags of who is talking. Specifically, we compare LENA’s output with a gold standard set of manually generated talker tags from a dataset of 88 day-long recordings, taken from 44 infants at 6 and 7 months, which includes 57,983 utterances. We compare accuracy across a range of classifications from the original Lena Technical Report, alongside a set of analyses examining classification accuracy by utterance type (e.g., declarative, singing). Consistent with previous validations, we find overall high agreement between the human and LENA-generated speaker tags for adult speech in particular, with poorer performance identifying child, overlap, noise, and electronic speech (accuracy range across all measures: 0–92%). We discuss several clear benefits of using this automated system alongside potential caveats based on the error patterns we observe, concluding with implications for research using LENA-generated speaker tags.
Compared to early language development, later changes to the language system during orthography and literacy acquisition have not yet been researched in detail. We present a longitudinal corpus of texts on short picture stories written by German primary school children between grades 2 and 4 and grades 3 and 4. It includes 1,922 texts with 212,505 tokens (6,364 types) from 251 children. For each text, rich metadata is available, including age, grade and linguistic background (at least 60% of the children were multilingual). To our knowledge, our corpus is the largest longitudinal corpus of written texts by children at primary school age. Each word is included in its original spelling as well as in a normalized form (target hypothesis), specifying the intended word form, which we corrected for orthographic but not grammatical errors. Original and target word forms are aligned character-wise and the target word forms are enriched with phonological, syllabic, and morphological information. Additionally, for each target word form, we established key lexical variables, e.g., word frequency or summed bigram frequency, as specified in childLex. Where applicable, we also specify key features of German orthography (e.g., consonant doubling, vowel-lengthening <h>). Taken together, this information allows for a detailed assessment of the properties of words that tend to increase the likelihood of spelling errors. The corpus is available in different formats—as tab-delimited annotated token and type based lists, in an XML format, and via the corpus search tool ANNIS.
We present a new dataset of English word recognition times for a total of 62 thousand words, called the English Crowdsourcing Project. The data were collected via an internet vocabulary test in which more than one million people participated. The present dataset is limited to native English speakers. Participants were asked to indicate which words they knew. Their response times were registered, although at no point were the participants asked to respond as quickly as possible. Still, the response times correlate around .75 with the response times of the English Lexicon Project for the shared words. Also, the results of virtual experiments indicate that the new response times are a valid addition to the English Lexicon Project. This not only means that we have useful response times for some 35 thousand extra words, but we now also have data on differences in response latencies as a function of education and age.
Here we describe the Jena Speaker Set (JESS), a free database for unfamiliar adult voice stimuli, comprising voices from 61 young (18–25 years) and 59 old (60–81 years) female and male speakers uttering various sentences, syllables, read text, semi-spontaneous speech, and vowels. Listeners rated two voice samples (short sentences) per speaker for attractiveness, likeability, two measures of distinctiveness (“deviation”-based [DEV] and “voice in the crowd”-based [VITC]), regional accent, and age. Interrater reliability was high, with Cronbach’s α between .82 and .99. Young voices were generally rated as more attractive than old voices, but particularly so when male listeners judged female voices. Moreover, young female voices were rated as more likeable than both young male and old female voices. Young voices were judged to be less distinctive than old voices according to the DEV measure, with no differences in the VITC measure. In age ratings, listeners almost perfectly discriminated young from old voices; additionally, young female voices were perceived as being younger than young male voices. Correlations between the rating dimensions above demonstrated (among other things) that DEV-based distinctiveness was strongly negatively correlated with rated attractiveness and likeability. By contrast, VITC-based distinctiveness was uncorrelated with rated attractiveness and likeability in young voices, although a moderate negative correlation was observed for old voices. Overall, the present results demonstrate systematic effects of vocal age and gender on impressions based on the voice and inform as to the selection of suitable voice stimuli for further research into voice perception, learning, and memory.
The application of word associations has become increasingly widespread. However, the association norms produced by traditional free association tests tend not to exceed 10,000 stimulus words, making the number of associated words too small to be representative of the overall language. In this study we used text corpora totaling over 400 million Chinese words, along with a multitude of association measures, to automatically construct a Chinese Lexical Association Database (CLAD) comprising the lexical association of over 80,000 words. Comparison of the CLAD with a database of traditional Chinese word association norms shows that word associations extracted from large text corpora are similar in strength to those elicited from free association tests but contain a much greater number of associative word pairs. Additionally, the relatively small numbers of participants involved in the creation of traditional norms result in relatively coarse scales of association measurement, whereas the differentiation of association strengths is greatly enhanced in the CLAD. The CLAD provides researchers with a great supplement to traditional word association norms. A query website at www.chinesereadability.net/LexicalAssociation/CLAD/ affords access to the database.
The research of the word is still very much the research of the noun. Adjectives have been largely overlooked, despite being the second-largest word class in many languages and serving an important communicative function, because of the rich, nuanced qualifications they afford. Adjectives are also ideally suited to study the interface between cognition and emotion, as they naturally cover the entire range of lexicosemantic variables such as imageability (infinite–green), and affective variables such as valence (sad–happy). We illustrate this by showing how the centrality of words in the mental lexicon varies as a function of the words’ affective dimensions, using newly collected norms for 1,000 Dutch adjectives. The norms include the lexicosemantic variables age of acquisition, familiarity, concreteness, and imageability; the affective variables valence, arousal, and dominance; and a variety of distributional variables, including network statistics resulting from a large-scale word association study. The norms are freely available from https://osf.io/nyg8v/, for researchers studying adjectives specifically or for whom adjectives constitute convenient stimuli to study other topics, such as vagueness, inference, spatial cognition, or affective word processing.
The Large Database of English Compounds (LADEC) consists of over 8,000 English words that can be parsed into two constituents that are free morphemes, making it the largest existing database specifically for use in research on compound words. Both monomorphemic (e.g., wheel) and multimorphemic (e.g., teacher) constituents were used. The items were selected from a range of sources, including CELEX, the English Lexicon Project, the British Lexicon Project, the British National Corpus, and Wordnet, and were hand-coded as compounds (e.g., snowball). Participants rated each compound in terms of how predictable its meaning is from its parts, as well as the extent to which each constituent retains its meaning in the compound. In addition, we obtained linguistic characteristics that might influence compound processing (e.g., frequency, family size, and bigram frequency). To show the usefulness of the database in investigating compound processing, we conducted a number of analyses that showed that compound processing is consistently affected by semantic transparency, as well as by many of the other variables included in LADEC. We also showed that the effects of the variables associated with the two constituents are not symmetric. In short, LADEC provides the opportunity for researchers to investigate a number of questions about compounds that have not been possible to investigate in the past, due to the lack of sufficiently large and robust datasets. In addition to directly allowing researchers to test hypotheses using the information included in LADEC, the database will contribute to future compound research by allowing better stimulus selection and matching.
The BT Archives house the records of British Telecom, the world's oldest telecommunications company, which traces its history back to the formation of the Electric Telegraphy Company in 1846. Prior to its privatisation in 1984, BT was a public corporation (and before that a government department) and as a result all of the pre-privatisation material in the archives is in the public domain, making it ideal for academic research. Despite this legal availability, however, the physical availability of material in the archive was limited to two days a week in an archive space in Holborn, London. In 2011 the ‘New Connections' project was set up with the aim of making around half a million items from the public archives of British Telecom available in a new digital archive. As part of ‘New Connections', three academic research projects were funded, one of which was the creation and analysis of the British Telecom Correspondence Corpus (BTCC). The era that the archive covers makes it a potentially fascinating source of data for the linguistic study of business correspondence. The mid-nineteenth to latetwentieth century is a crucial period in the development of English business correspondence as the amount of business being conducted by letter increased massively during this period as a result of the Industrial Revolution, the introduction of the Penny Post, and increased access to education both in schools and through composition grammar guides. Despite its importance in the development of business correspondence, this period has received relatively little attention. The aim of constructing the British Telecom Correspondence Corpus was to start addressing this gap in available linguistic data and enable studies into the development of business correspondence from the mid-nineteenth to late-twentieth century.
WordNet, a large lexical database of English, was conceived as a model of human semantic organization. Evidence from timing experiments, association norms, and distributional properties of words supported a semantic network model in which words are interlinked via a small number of lexical and conceptual relations. Its large coverage and unique structure, which allows automatic systems to detect and quantify semantic relatedness among words, soon made WordNet an invaluable tool for natural language processing tasks. Information retrieval, document summarization, and machine translation crucially require word sense discrimination and disambiguation. Wordnets have been built in dozens of languages and for specific technical sublanguages, and the number of applications in research, language technology and pedagogy has grown. Although WordNet’s central focus has shifted from its psycholinguistic origins, its design, based on theories about the structure of the human mental lexicon, is validated as a sound approach to representing the meanings of words. (PsycINFO Database Record (c) 2019 APA, all rights reserved)
Our understanding of the mental lexicon, the way meaning is extracted from word forms, is almost entirely built on data from spoken languages. While there is much work demonstrating that in many ways the linguistic structure and psychological mechanisms for processing signed language and spoken language processing are the same, less is known about the signed language mental lexicon. In this dissertation, I examine the structure of the American Sign Language mental lexicon, and the ways meaning can be extracted from the manual/visual signal. In the third chapter of this dissertation I ask whether a single cognitive architecture might explain diverse behavioral patterns in signed and spoken language. Chen and Mirman (2012) presented a computational model of word processing that unified opposite effects of neighborhood density in speech production, perception, and written word recognition. Carreiras et al. (2008) demonstrate that neighborhood density effects in Spanish Sign Language (LSE) also vary depending on whether the neighbors share the same handshape or location. We present a spreading activation architecture that borrows the principles proposed by Chen and Mirman (2012), and show that if this architecture is elaborated to incorporate relatively minor facts about either 1) the time course of sign perception or 2) the frequency of sub-lexical units in sign languages, it produces data that match the experimental findings from sign languages. This work serves as a proof of concept that a single cognitive architecture could underlie both sign and word recognition. In the second chapter I present ASL-LEX, a lexical database for ASL that catalogues more than forty properties about almost 1,000 signs. The database includes, for example, information about each sign's iconicity, phonological make-up, and neighborhood density. I use this information to better understand the structure of the ASL lexicon, the distribution of each of these properties, and the relationships between these properties. This lexical database is the largest and most comprehensive database of ASL, and can be used by researchers to develop experiments and by educators to identify and support vocabulary development. In the fourth chapter, I use ASL-LEX to develop a tightly-controlled study of sign perception. I ask whether neighborhood density and sub-lexical frequency play a role in sign perception, and if the mechanisms of sign perception are affected by early language experience. Eighty deaf participants with varying early language backgrounds completed a lexical decision task. I find that neighborhood density inhibits sign perception in people with low early ASL exposure, but has no effect in people with high early ASL exposure. Location frequency inhibits sign perception in all people, but the effect is stronger in people with low early ASL exposure. This suggests that impoverished access to ASL early in life has lasting consequences for sign perception. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
The design of experimental tasks in psychology and linguistics requires using stimulus with properties and characteristics in standardized values. This allows predicting with higher accuracy the impact of the stimulus presentation. The lexical associative norms are instruments that determine the strength of association between two concepts. The most common method to construct these norms is to take a free response from a presentation of a cue word. The main goal of this study was to construct lexical associative norms of 407 Spanish words. 800 students from Ciudad de Córdoba, Argentina, participated in the study. Quantitative analyses were performed taking into account the number of valid answers, blank and non valid answers, and number of associates per item. A qualitative classification was performed according to the strength of association. Additionally, it is presented a group of psycholinguistic indexes for a better description of the items used. Correlation analysis demonstrated a strong and negative relation between the frequency of first and second associations and the number of associations per item. This study pretends to be highly useful in research in psychology and linguistic where it is required consulting the norms presented to the design of evaluation instruments. (PsycINFO Database Record (c) 2018 APA, all rights reserved)
The Database for Spoken German (Datenbank f{\"{u}}r Gesprochenes Deutsch, DGD2, http://dgd.ids-mannheim.de) is the central platform for publishing and disseminating spoken language corpora from the Archive of Spoken German (Archiv f{\"{u}}r Gesprochenes Deutsch, AGD, http://agd.ids-mannheim.de) at the Institute for the German Language in Mannheim. The corpora contained in the DGD2 come from a variety of sources, some of them in-house projects, some of them external projects. Most of the corpora were originally intended either for research into the (dialectal) variation of German or for studies in conversation analysis and related fields. The AGD has taken over the task of permanently archiving these resources and making them available for reuse to the research community. To date, the DGD2 offers access to 19 different corpora, totalling around 9000 speech events, 2500 hours of audio recordings or 8 million transcribed words. This paper gives an overview of the data made available via the DGD2, of the technical basis for its implementation, and of the most important functionalities it offers. The paper concludes with information about the users of the database and future plans for its development.
Naturalistic learner productions are an important empirical resource for SLA research. Some pioneering works have produced valuable second language (L2) resources supporting SLA research.1 One common limitation of these resources is the absence of individual longitudinal data for numerous speakers with different backgrounds across the proficiency spectrum, which is vital for understanding
To cite this version: Pollet Samvelian, Pegah Faghiri. Introducing PersPred, a syntactic and semantic database for Persian Complex Predicates. Abstract This paper introduces PersPred, the first manually elaborated syntactic and semantic database for Persian Complex Predicates (CPs). Beside their theoretical interest, Per-sian CPs constitute an important challenge in Persian lexicography and for NLP. The first delivery, PersPred 1 1 , contains 700 CPs, for which 22 fields of lexical, syntactic and semantic information are encoded. The semantic classification PersPred provides allows to account for the productivity of these combinations in a way which does justice to their compositionality without overlooking their id-iomaticity.
Since long it has been noted that cross-linguistically recurring polysemies can serve as an indi-cator of conceptual relations, and quite a few approaches to model and analyze such data have been proposed in the recent past. Although – given the nature of the data – it seems natural to model and analyze it with the help of network techniques, there are only a few approaches which make explicit use of them. In this paper, we show how the strict application of weighted network models helps to get more out of cross-linguistic polysemies than would be possible using approaches that are only based on item-to-item comparison. For our study we use a large dataset consisting of 1252 semantic items translated into 195 different languages covering 44 different language families. By analyz-ing the community structure of the network reconstructed from the data, we find that a majority of the concepts (68{\%}) can be separated into 104 large communities consisting of five and more nodes. These large communities almost exclusively constitute meaningful groupings of concepts into con-ceptual fields. They provide a valid starting point for deeper analyses of various topics in historical semantics, such as cognate detection, etymological analysis, and semantic reconstruction.
LAPSyD, the Lyon-Albuquerque Phonological Systems Database, is an online phonological database equipped with powerful query, mapping and visualization tools. It stems from the UPSID and WALS databases, enhanced with newly validated data not only covering segmental inventories but also syllable structures, stress and tonal systems. In its current version it covers around 700 languages and it is accessible at http://www.lapsyd.ddl.ish-lyon.cnrs.fr. This paper provides a description of the data structure in LAPSyD and the features of the interface. Brief illustrations of the types of analysis that can be done with this tool are provided, exploiting the ability to cross-reference data on segments, other phonological properties and language location. Copyright {\textcopyright} 2013 ISCA.
This paper describes the SubCat-Extractor as a novel tool to obtain verb subcategori-sation data from parsed German web corpora. The SubCat-Extractor is based on a set of detailed rules that go beyond what is directly accessible in the parses. The extracted subcategorisation database is represented in a compact but linguistically detailed and flexible format, comprising various aspects of verb information, complement information and sentence information , within a one-line-per-clause style. We describe the tool, the extraction rules and the obtained resource database, as well as actual and potential uses in computational linguistics.
This paper serves as an initial announcement of the avail- ability of a corpus of articulatory data called mngu0. This cor- pus will ultimately consist of a collection of multiple sources of articulatory data acquired from a single speaker: electro- magnetic articulography (EMA), audio, video, volumetric MRI scans, and 3D scans of dental impressions. This data will be provided free for research use. In this first stage of the release, we are making available one subset of EMA data, consisting of more than 1,300 phonetically diverse utterances recorded with a Carstens AG500 electromagnetic articulograph. Distribution of mngu0 will be managed by a dedicated “forum-style” web site. This paper both outlines the general goals motivating the distribution of the data and the creation of the mngu0 web fo- rum, and also provides a description of the EMA data contained in this initial release.
The lexical database dlexDB supplies in form of an online database frequency-based norms of numerous processrelated word properties for psychological and linguistic research. These values include well known variables such as printed frequency of word form and lemma as documented also in CELEX (Baayen, Piepenbrock und Gulikers, 1995). In addition, we compute new values like frequencies based on syllables, and morphemes as well as frequencies of character chains, and multiple word combinations. The statistics are based on the Kernkorpus des Digitalen Wörterbuchs der deutschen Sprache (DWDS) with over 100 million running words. We illustrate the validity of these norms with new results about fixation durations in sentence reading. (PsycINFO Database Record (c) 2016 APA, all rights reserved)
The article reports on the importance and impact of the electronic corpora in linguistics. It mentions that since the electronic corpus of English, problem exists for corpus compilers because of the different needs of linguists. It suggests that the problem lies on the database of the text descriptors that would guide researchers through the existing repositories to enable researchers to create their own corpora.
In this paper we introduce the first version of noWaC, a large web-based corpus of Bokm{\aa}l Norwegian currently containing about 700 million tokens. The corpus has been built by crawling, downloading and processing web documents in the .no top-level internet domain. The procedure used to collect the noWaC corpus is largely based on the techniques described by Ferraresi et al. (2008). In brief, first a set of "seed" URLs containing documents in the target language is collected by sending queries to commercial search engines (Google and Yahoo). The obtained seeds (overall 6900 URLs) are then used to start a crawling job using the Heritrix web-crawler limited to the .no domain. The downloaded documents are then processed in various ways in order to build a linguistic corpus (e.g. filtering by document size, language identification, duplicate and near duplicate detection, etc.).
Frequency, familiarity, and age of acquisition are important factors for word recognition that must be considered by researchers of language acquisition. Current psycholinguistic databases, based on studies of native English speakers, include objective frequency count, subjective rating of familiarity, and age of acquisition. One can easily employ those databases to obtain a stimulus list for one's studies. For word recognition researchers interested in non-native English speakers in of Taiwan, however, there is currently no existing database. In this study, we created a psycholinguistic database which includes subjective familiarity rating and age of acquisition for 3,080 English words. Participants were 120 college students in Taiwan. They were asked to make judgments about 4,000 stimulus words. For recognized stimulus words, participants gave a rating of familiarity and self-report of age of acquisition; for non-recognized words, participants were asked to move on to the next stimulus word. Further analysis of the database showed that familiarity index, age of acquisition, and number of syllables are important factors for the recognition of a word. Variance in word recognition ratings for each factor was explained and implications were discussed.
We describe the Leipzig Corpora collection (LCC), a freely available resource for corpora and corpus statistics covering more than 20 languages at the time being. Unified format and easy accessibility encourage incorporation of the data into many projects and render the collection a useful resource especially in multilingual settings and for small languages. The preparation of monolingual corpora of standard sizes from different sources (web, newspaper, Wikipedia) is described in detail.