1358 norm sets
<jats:p>The collections of the Internet Archive include many digitized historical sources. Many contain rich bibliographic data in a format called MARC. In this lesson, you'll learn how to use Python to automate the downloading of large numbers of MARC files from the Internet Archive and the parsing of MARC records for specific information such as authors, places of publication, and dates. The lesson can be applied more generally to other Internet Archive files and to MARC records found elsewhere.</jats:p>
<jats:p>Wikipedia is a well known free content, multilingual encyclopedia written collaboratively by contributors around the world. Anybody can edit an article using a wiki markup language that offers a simplified alternative to HTML. This encyclopedia is composed of millions of articles in different languages.</jats:p>
<jats:p>In this paper, we provide an overview of the new GloWbE Corpus — the Corpus of Global Web-based English. GloWbE is based on 1.9 billion words in 1.8 million web pages from 20 different English-speaking countries. Approximately 60 percent of the corpus comes from informal blogs, and the rest from a wide range of other genres and text types. Because of its large size, its architecture and interface, the corpus can be used to examine many types of variation among dialects, which might not be possible with other corpora — including variation in lexis, morphology, (medium- and low-frequency) syntactic constructions, variation in meaning, as well as discourse and its relationship to culture.</jats:p>
<jats:title>Abstract</jats:title><jats:p>Borrowing affixes may be rare compared to lexical borrowing, but it is not random. The current study describes regular patterns of affix borrowing in a database containing 649 borrowed affixes, challenging a number of previous claims about relative borrowability, in particular regarding inflectional categories. It is shown that borrowing affixes of all major nominal and verbal inflectional categories, including case markers and argument indexes, is well attested. Borrowing case markers, for instance, appears to be just as common as borrowing plural markers. By factoring in the “availability” for borrowing (i.e. whether a potential donor language has a relevant affix), it can be shown that nominal categories are far more frequently borrowed than verbal categories. Additionally, it is shown that sets of borrowed affixes often consist of interrelated sets of forms, e.g. forming paradigms, rather than being isolated forms from different morphosyntactic systems, in particular for the more tightly integrated inflectional subsystems. The frequency and systematicity by which inflectional affixes are borrowed calls for a reconsideration of the role of inflection in models of language contact.</jats:p>
<jats:p>RESUMO O presente estudo tem como objetivo descrever os desafios e soluções encontrados na compilação do Corpus de Português Escrito em Periódicos - CoPEP, que contém aproximadamente 40 milhões de palavras, é equilibrado entre as variedades português brasileiro e português europeu em número de palavras e cobre seis grandes áreas de conhecimento. Primeiramente, apresentaremos o contexto de criação do CoPEP, qual seja, a elaboração de um dicionário on-line de português para universitários, para o qual serviu como fonte primária de obtenção de evidências linguísticas. Assim, foram as características desse projeto lexicográfico que informaram os critérios de criação do desenho do CoPEP e as consequentes tomadas de decisão. A seguir, descreveremos a metodologia de aquisição de dados, com foco especial nos desafios enfrentados e nas soluções encontradas. Terminaremos com a descrição da fase final de compilação, na qual aplicamos uma série de procedimentos para obtenção de equilíbrio.</jats:p>
<jats:p>We examined the potential advantage of the lexical databases using subtitles and present SUBTLEX-PT, a new lexical database for 132,710 Portuguese words obtained from a 78 million corpus based on film and television series subtitles, offering word frequency and contextual diversity measures. Additionally we validated SUBTLEX-PT with a lexical decision study involving 1920 Portuguese words (and 1920 nonwords) with different lengths in letters ( M = 6.89, SD = 2.10) and syllables ( M = 2.99, SD = 0.94). Multiple regression analyses on latency and accuracy data were conducted to compare the proportion of variance explained by the Portuguese subtitle word frequency measures with that accounted by the recent written-word frequency database (Procura-PALavras; P-PAL; Soares, Iriarte, et al., 2014). As its international counterparts, SUBTLEX-PT explains approximately 15% more of the variance in the lexical decision performance of young adults than the P-PAL database. Moreover, in line with recent studies, contextual diversity accounted for approximately 2% more of the variance in participants' reading performance than the raw frequency counts obtained from subtitles. SUBTLEX-PT is freely available for research purposes (at http://p-pal.di.uminho.pt/about/databases ).</jats:p>
Feature stability, time and tempo of change, and the role of genealogy versus areality in creating linguistic diversity are important issues in current computational research on linguistic typology. This paper presents a database initiative, DiACL Typology, which aims to provide a resource for addressing these questions with specific of the extended Indo-European language area of Eurasia, the region with the best documented linguistic history. The database is pre-prepared for statistical and phylogenetic analyses and contains both linguistic typological data from languages spanning over four millennia, and linguistic metadata concerning geographic location, time period, and reliability of sources. The typological data has been organized according to a hierarchical model of increasing granularity in order to create datasets that are complete and representative. [ABSTRACT FROM AUTHOR], Copyright of PLoS ONE is the property of Public Library of Science and its content may not be copied o)
The Lesser Sunda Islands in eastern Indonesia cover a longitudinal distance of some 600 kilometres. They are the westernmost place where languages of the Austronesian family come into contact with a family of Papuan languages and constitute an area of high linguistic diversity. Despite its diversity, the Lesser Sundas are little studied and for most of the region, written historical records, as well as archaeological and ethnographic data are lacking. In such circumstances the study of relationships between languages through their lexicon is a unique tool for making inferences about human (pre-)history and tracing population movements. However, the lack of a collective body of lexical data has severely limited our understanding of the history of the languages and peoples in the Lesser Sundas. The LexiRumah database fills this gap by assembling lexicons of Lesser Sunda languages from published and unpublished sources, and making those lexicons available online in a consistent format. T)
Presents a study which aims to investigate SPALEX, a Spanish lexical decision database by focusing on native Spanish speakers at a global scale and with a vast amount of words, to provide a useful tool for researchers exploring the acquisition and processing of this language in native and foreign contexts. SPALEX contains data from a Spanish crowd-sourced lexical decision mega study. The authors collected the data through an online platform from May 12th, 2014 to December 19th, 2017. The majority of the data was acquired during the first month of the experiment, when an advertising campaign was done in order to attract the public’s attention. Participants also had the option of publishing their results via social networks, which led to attract more participants in a snow-ball sampling fashion. Additionally, the database contains information on participants that voluntarily provided information about their gender, age, country of origin, education level, handedness, native language, and best foreign language. In each experimental session, participants responded to 70 words and 30 non-words presented randomly and without repetition. Accuracy in SPALEX is expressed as 1 for correct answers and 0 for incorrect answers. Based on participants’ responses, the authors calculated percentage known, a measure of the percentage of participants that know a particular word. (PsycINFO Database Record (c) 2018 APA, all rights reserved)
We introduce a dataset for studying the evolution of words, constructed from WordNet and the Google Books Ngram Corpus. The dataset tracks the evolution of 4,000 synonym sets (synsets), containing 9,000 English words, from 1800 AD to 2000 AD. We present a supervised learning algorithm that is able to predict the future leader of a synset: the word in the synset that will have the highest frequency. The algorithm uses features based on a word’s length, the characters in the word, and the historical frequencies of the word. It can predict change of leadership (including the identity of the new leader) fifty years in the future, with an F-score considerably above random guessing. Analysis of the learned models provides insight into the causes of change in the leader of a synset. The algorithm confirms observations linguists have made, such as the trend to replace the -ise suffix with -ize, the rivalry between the -ity and -ness suffixes, and the struggle between economy (shorter words ar)
The Moral Foundations Dictionary (MFD) is a useful tool for applying the conceptual framework developed in Moral Foundations Theory and quantifying the moral meanings implicated in the linguistic information people convey. However, the applicability of the MFD is limited because it is available only in English. Translated versions of the MFD are therefore needed to study morality across various cultures, including non-Western cultures. The contribution of this paper is two-fold. We developed the first Japanese version of the MFD (referred to as the J-MFD) using a semi-automated method—this serves as a reference when translating the MFD into other languages. We next tested the validity of the J-MFD by analyzing open-ended written texts about the situations that Japanese participants thought followed and violated the five moral foundations. We found that the J-MFD correctly categorized the Japanese participants’ descriptions into the corresponding moral foundations, and that the Moral F)
Language is one the earliest capacities affected by cognitive change. To monitor that change longitudinally, we have developed a web portal for remote linguistic data acquisition, called Talk2Me, consisting of a variety of tasks. In order to facilitate research in different aspects of language, we provide baselines including the relations between different scoring functions within and across tasks. These data can be used to augment studies that require a normative model; for example, we provide baseline classification results in identifying dementia. These data are released publicly along with a comprehensive open-source package for extracting approximately two thousand lexico-syntactic, acoustic, and semantic features. This package can be applied arbitrarily to studies that include linguistic data. To our knowledge, this is the most comprehensive publicly available software for extracting linguistic features. The software includes scoring functions for different tasks. [ABSTRACT FROM)