1358 norm sets
The present article provides Spanish norms for name agreement, printed word frequency, word compound frequency, familiarity, imageability, visual complexity, age of acquisition, and word length (measured by syllables and phonemes) for 100 line drawings of actions taken from Druks and Masterson (2000). In addition, through a naming-time experiment carried out with a group of 54 Spanish students in a pool of 63 of these line drawings, we determined the best predictors of naming actions. In the multiple regression analysis, age of acquisition and name agreement emerged as the most important determinants of action-naming reaction time.
This study introduces the Tool for the Automatic Analysis of Cohesion (TAACO), a freely available text analysis tool that is easy to use, works on most operating systems (Windows, Mac, and Linux), is housed on a user's hard drive (rather than having an Internet interface), allows for the batch processing of text files, and incorporates over 150 classic and recently developed indices related to text cohesion. The study validates TAACO by investigating how its indices related to local, global, and overall text cohesion can predict expert judgments of text coherence and essay quality. The findings of this study provide predictive validation of TAACO and support the notion that expert judgments of text coherence and quality are either negatively correlated or not predicted by local and overall text cohesion indices, but are positively predicted by global indices of cohesion. Combined, these findings provide supporting evidence that coherence for expert raters is a property of global cohesion and not of local cohesion, and that expert ratings of text quality are positively related to global cohesion.
In emotional research, efficient designs often rely on successful emotion induction. For visual stimulation, the only reliable database available so far is the International Affective Picture System (IAPS). However, extensive use of these stimuli lowers the impact of the images by increasing the knowledge that participants have of them. Moreover, the limited number of pictures for specific themes in the IAPS database is a concern for studies centered on a specific emotion thematic and for designs requiring a lot of trials from the same kind (e.g., EEG recordings). Thus, in the present article, we present a new database of 730 pictures, the Geneva Affective PicturE Database, which was created to increase the availability of visual emotion stimuli. Four specific negative contents were chosen: spiders, snakes, and scenes that induce emotions related to the violation of moral and legal norms (human rights violation or animal mistreatment). Positive and neutral pictures were also included: Positive pictures represent mainly human and animal babies as well as nature sceneries, whereas neutral pictures mainly depict inanimate objects. The pictures were rated according to valence, arousal, and the congruence of the represented scene with internal (moral) and external (legal) norms. The constitution of the database and the results of the picture ratings are presented.
Lexvo.org brings information about languages, words, and other linguistic entities to the Web of Linked Data. It defines URIs for terms, languages, scripts, and characters, which are not only highly interconnected but also linked to a variety of resources on the Web. Additionally, new datasets are being published to contribute to the emerging Linked Data Cloud of Language-Related information.
This paper describes the Atlante Sintattico d'Italia, Syntactic Atlas of Italy (ASIt) linguistic linked dataset. ASIt is a scientific project aiming to account for minimally different variants within a sample of closely related languages; it is part of the Edisyn network, the goal of which is to establish a European network of researchers in the area of language syntax that use similar standards with respect to methodology of data collection, data storage and annotation, data retrieval and cartography. In this context, ASIt is defined as a curated database which builds on dialectal data gathered during a twenty-year-long survey investigating the distribution of several grammatical phenomena across the dialects of Italy. Both the ASIt linguistic linked dataset and the Resource Description Framework Schema (RDF/S) on which it is based are publicly available and released with a Creative Commons license (CC BY-NC-SA 3.0). We report the characteristics of the data exposed by ASIt, the statistics about the evolution of the data in the last two years, and the possible usages of the dataset, such as the generation of linguistic maps. {\textcopyright} 2012 - IOS Press and the authors. All rights reserved.
This paper presents a rule-based approach for generating a large phonetic database for Romanian. The knowledge base is developed by means of the GRAALAN (Grammar Abstract Language) system. By inspecting dictionaries and corpora, we generate a phonetic database over 100,000 lemmas. Our database has a high degree of accuracy ensured by our rule-based method applied for generating phonetic transcriptions.
As human activity and interaction increasingly take place online, the digital residues of these activities provide a valuable window into a range of psychological and social processes. A great deal of progress has been made toward utilizing these opportunities; however, the complexity of managing and analyzing the quantities of data currently available has limited both the types of analysis used and the number of researchers able to make use of these data. Although fields such as computer science have developed a range of techniques and methods for handling these difficulties, making use of those tools has often required specialized knowledge and programming experience. The Text Analysis, Crawling, and Interpretation Tool (TACIT) is designed to bridge this gap by providing an intuitive tool and interface for making use of state-of-the-art methods in text analysis and large-scale data management. Furthermore, TACIT is implemented as an open, extensible, plugin-driven architecture, which will allow other researchers to extend and expand these capabilities as new methods become available.
Despite their relatively low sampling factor, the freely available, randomly sampled status streams of Twitter are very useful sources of geographically embedded social network data. To statistically analyze the information Twitter provides via these streams, we have collected a year's worth of data and built a multi-terabyte relational database from it. The database is designed for fast data loading and to support a wide range of studies focusing on the statistics and geographic features of social networks, as well as on the linguistic analysis of tweets. In this paper we present the method of data collection, the database design, the data loading procedure and special treatment of geo-tagged and multi-lingual data. We also provide some SQL recipes for computing network statistics.
Word lists are most commonly used in the investigation of human memory. To prevent transfer effects, repeated measures of memory for words require multiple lists of different words. Yet, the psycholinguistic properties of all word lists employed should match as closely as possible to avoid confounding with the independent variable(s) in question. Although comprehensive databases for word norms exist, to our knowledge no tool is available that automates the creation of such equivalent word lists. Instead, matching different lists is often accomplished prima facie. We have therefore developed a Windows program called EQUIWORD that completely automates the creation of word lists that are truly parallel with respect to a wide range of attributes. EQUIWORD takes psycholinguistic databases of different formats as input and computes several coefficients of distance for every possible word pairing. Program output consists of a list of all word pairs sorted according to their distance. On that basis, creating equivalent word lists is simply done by selecting the pairs with the lowest distance coefficients.
Although many facial and vocal databases are available for research, very few of them have controlled the range of attractiveness of the stimuli that they offer. To fill this gap, we created the GEneva Faces and Voices (GEFAV) database, providing standardized faces (static and dynamic neutral, smiling) and voices (speaking sentences, vowels) of young European adults. A total of 61 women and 50 men 18-35 years old agreed to be part of the GEFAV stimuli, and two rating studies involving 285 participants provided evaluations of the facial and vocal samples. The final set of stimuli was satisfactory in terms of attractiveness range (wide and rather symmetrical distribution over the attractiveness continuum) and the reliability of the ratings (high consistency between the two rating studies, high interrater agreement in the final rating study). Moreover, the database showed an adequate validity, since a series of findings described by earlier research on human attractiveness were confirmed-namely, that facial and vocal attractiveness are predicted by femininity and health in women, and by masculinity, dominance, and trustworthiness in men. In future studies, the GEFAV stimuli may be used intact or transformed, individually or in multimodal combinations, to investigate a wide range of mechanisms, such as the behavioral, neuropsychological, and neurophysiological processes involved in social cognition.
We present a database of 858 German words from the semantic fields of authority and community, which represent core dimensions of human sociality. The words were selected on the basis of co-occurrence profiles of representative keywords for these semantic fields. All words were rated along five dimensions, each measured by a bipolar semantic-differential scale: Besides the classic dimensions of affective meaning (valence, arousal, and potency), we collected ratings of authority and community with newly developed scales. The results from cluster, correlational, and multiple regression analyses on the rating data suggest a robust negativity bias for authority valuation among German raters recruited via university mailing lists, whereas community ratings appear to be rather unrelated to the well-established affective dimensions. Furthermore, our data involve a strong overall negative correlation-rather than the classical U-shaped distribution-between valence and arousal for socially relevant concepts. Our database provides a valuable resource for research questions at the intersection of cognitive neuroscience and social psychology. It can be downloaded as supplemental materials with this article.
This study presents a database of 500 words from five semantic categories: animals, body parts, furniture, clothing, and intelligence. Each category contains 100 words, and data on lexical availability, age of acquisition, imageability, typicality, concept familiarity, written word frequency, and word length in number of syllables are provided with each word. The full set of norms may be downloaded from www.psychonomic.org/archive.
Normative databases containing psycholinguistic variables are commonly used to aid stimulus selection for investigations into language and other cognitive processes. Norms exist for many languages, but not for Thai. The aim of the present research, therefore, was to obtain Thai normative data for the BOSS, a set of 480 high resolution color photographic images of real objects (Brodeur et al. in PLoS ONE 5(5), 2010. https://doi.org/10.1371/journal.pone.0010773 ). Norms were provided by 584 Thai university students on eight dimensions: name agreement, object familiarity, visual complexity, category agreement, image agreement, two types of manipulability (graspability and mimeability), and age of acquisition. The results revealed comparatively similar levels of name agreement to Brodeur et al. especially when unfamiliar items were factored out. The pattern of intercorrelations among the Thai psycholinguistic norms was comparable to previous studies and our cross-linguistic correlations were robust for the same set of pictures in English and French. Conjointly, the findings extend the relevancy of the BOSS to Thailand, supporting this photographic resource for investigations of language and other cognitive processes in monolingual, multilingual, and brain-impaired populations.
In this study, we aimed to provide a large-scale set of psycholinguistic norms for 3,314 traditional Chinese characters, along with their naming reaction times (RTs), collected from 140 Chinese speakers. The lexical and semantic variables in the database include frequency, regularity, familiarity, consistency, number of strokes, homophone density, semantic ambiguity rating, phonetic combinability, semantic combinability, and the number of disyllabic compound words formed by a character. Multiple regression analyses were conducted to examine the predictive powers of these variables for the naming RTs. The results demonstrated that these variables could account for a significant portion of variance (55.8{\%}) in the naming RTs. An additional multiple regression analysis was conducted to demonstrate the effects of consistency and character frequency. Overall, the regression results were consistent with the findings of previous studies on Chinese character naming. This database should be useful for research into Chinese language processing, Chinese education, or cross-linguistic comparisons. The database can be accessed via an online inquiry system (http://ball.ling.sinica.edu.tw/namingdatabase/index.html).
A major obstacle for the design of rigorous, reproducible studies in moral psychology is the lack of suitable stimulus sets. Here, we present the Socio-Moral Image Database (SMID), the largest standardized moral stimulus set assembled to date, containing 2,941 freely available photographic images, representing a wide range of morally (and affectively) positive, negative and neutral content. The SMID was validated with over 820,525 individual judgments from 2,716 participants, with normative ratings currently available for all images on affective valence and arousal, moral wrongness, and relevance to each of the five moral values posited by Moral Foundations Theory. We present a thorough analysis of the SMID regarding (1) inter-rater consensus, (2) rating precision, and (3) breadth and variability of moral content. Additionally, we provide recommendations for use aimed at efficient study design and reproducibility, and outline planned extensions to the database. We anticipate that the SMID will serve as a useful resource for psychological, neuroscientific and computational (e.g., natural language processing or computer vision) investigations of social, moral and affective processes. The SMID images, along with associated normative data and additional resources are available at https://osf.io/2rqad/.
Despite flourishing research on the relationship between emotion and literal language, and despite the pervasiveness of figurative expressions in communication, the role of figurative language in conveying affect has been underinvestigated. This study provides affective and psycholinguistic norms for 619 German idiomatic expressions and explores the relationships between affective and psycholinguistic idiom properties. German native speakers rated each idiom for emotional valence, arousal, familiarity, semantic transparency, figurativeness, and concreteness. They also described the figurative meaning of each idiom and rated how confident they were about the attributed meaning. The results showed that idioms rated high in valence were also rated high in arousal. Negative idioms were rated as more arousing than positive ones, in line with results from single words. Furthermore, arousal correlated positively with figurativeness (supporting the idea that figurative expressions are more emotionally engaging than literal expressions) and with concreteness and semantic transparency. This suggests that idioms may convey a more direct reference to sensory representations, mediated by the meanings of their constituting words. Arousal correlated positively with familiarity. In addition, positive idioms were rated as more familiar than negative idioms. Finally, idioms without a literal counterpart were rated as more emotionally valenced and arousing than idioms with a literal counterpart. Although the meanings of ambiguous idioms were less correctly defined than those of unambiguous idioms, ambiguous idioms were rated as more concrete than unambiguous ones. We also discuss the relationships between the various psycholinguistic variables characterizing idioms, with reference to the literature on idiom structure and processing.
Research on the emotional, cognitive, and social determinants of moral judgment has surged in recent years. The development of moral foundations theory (MFT) has played an important role, demonstrating the breadth of morality. Moral psychology has responded by investigating how different domains of moral judgment are shaped by a variety of psychological factors. Yet, the discipline lacks a validated set of moral violations that span the moral domain, creating a barrier to investigating influences on judgment and how their neural bases might vary across the moral domain. In this paper, we aim to fill this gap by developing and validating a large set of moral foundations vignettes (MFVs). Each vignette depicts a behavior violating a particular moral foundation and not others. The vignettes are controlled on many dimensions including syntactic structure and complexity making them suitable for neuroimaging research. We demonstrate the validity of our vignettes by examining respondents' classifications of moral violations, conducting exploratory and confirmatory factor analysis, and demonstrating the correspondence between the extracted factors and existing measures of the moral foundations. We expect that the MFVs will be beneficial for a wide variety of behavioral and neuroimaging investigations of moral cognition.
The main purpose of this study was to report age-based subjective age-of-acquisition (AoA) norms for 600 Turkish words. A total of 115 children, 100 young adults, 115 middle-aged adults, and 127 older adults provided AoA estimates for 600 words on a 7-point scale. The intraclass correlations suggested high reliability, and the AoA estimates were highly correlated across the four age groups. Children gave earlier AoA estimates than the three adult groups; this was true for high-frequency as well as low-frequency words. In addition to the means and standard deviations of the AoA estimates, we report word frequency, concreteness, and imageability ratings, as well as word length measures (numbers of syllables and letters), for the 600 words as supplemental materials. The present ratings represent a potentially useful database for researchers working on lexical processing as well as other aspects of cognitive processing, such as autobiographical memory.
Humans appear to rely on spatial mappings to describe and represent concepts. In particular, conceptual cueing refers to the effect whereby after reading or hearing a particular word, the location of observers' visual attention in space can be systematically shifted in a particular direction. For example, words such as "sun" and "happy" orient attention upwards, whereas words such as "basement" and "bitter" orient attention downwards. This area of research has garnered much interest, particularly within the embodied cognition framework, for its potential to enhance our understanding of the interaction between abstract cognitive processes such as language and basic visual processes such as attention and stimulus processing. To date, however, this area has relied on subjective classification criteria to determine whether words ought to be classified as having a meaning that implies "up" or "down." The present study, therefore, provides a set of 498 items that have each been systematically rated by over 90 participants, providing refined, continuous measures of the extent to which people associate given words with particular spatial dimensions. The resulting database provides an objective means to aid item-selection for future research in this area.
Studies of semantic variables (e.g., concreteness) and affective variables (i.e., valence and arousal) have traditionally tended to run in different directions. However, in recent years there has been growing interest in studying the relationship, as well as the potential overlaps, between the two. This article describes a database that provides subjective ratings for 1,400 Spanish words for valence, arousal, concreteness, imageability, context availability, and familiarity. Data were collected online through a process involving 826 university students. The results showed a high interrater reliability for all of the variables examined, as well as high correlations between our affective and semantic values and norms currently available in other Spanish databases. Regarding the affective variables, the typical quadratic correlation between valence and arousal ratings was obtained. Likewise, significant correlations were found between the lexico-semantic variables. Importantly, we obtained moderate negative correlations between emotionality and both concreteness and imageability. This is in line with the claim that abstract words have more affective associations than concrete ones (Kousta, Vigliocco, Vinson, Andrews, {\&} Del Campo, 2011). The present Spanish database is suitable for experimental research into the effects of both affective properties and lexico-semantic variables on word processing and memory.
Synesthesia is a neurological phenomenon in which certain types of stimuli elicit involuntary perceptions in an unrelated pathway. A common type of synesthesia is grapheme–color synesthesia, in which the visual perception of letters and numbers stimulates the perception of a specific color. Previous studies have often collected relatively small numbers of grapheme–color associations per synesthete, but the accumulation of a large quantity of data has greater promise for uncovering the mechanisms underlying synesthetic association. In this study, we therefore collected large samples of data from a total of eight synesthetes. All told, we obtained over 1000 synesthetic colors associated with Japanese kanji characters from each of two synesthetes, over 100 synesthetic colors form each of three synesthetes, and about 80 synesthetic colors associated with Japanese hiragana, Latin letters, and Arabic numerals from each of three synesthetes. We then compiled the data into a database, called the KANJI-Synesthetic Colors Database (K-SCD), which has a total of 5122 colors for 483, 46, and 46 Japanese kanji, hiragana, and katakana characters, respectively, as well as for 26 Latin letters and ten Arabic numerals. In addition to introducing the K-SCD, this article demonstrates the database's merits by using two examples, in which two new rules for synesthetic association, “shape similarity” and “synesthetic color clustering,” were found. The K-SCD is publicly accessible (www.cv.jinkan.kyoto-u.ac.jp/site/uploads/K-SCD.xlsm) and will be a valuable resource for those who wish to conduct statistical analyses using a rich dataset in order to uncover the rules governing synesthetic association and to understand its mechanisms.
The LSE-Sign database is a free online tool for selecting Spanish Sign Language stimulus materials to be used in experiments. It contains 2,400 individual signs taken from a recent standardized LSE dictionary, and a further 2,700 related nonsigns. Each entry is coded for a wide range of grammatical, phonological, and articulatory information, including handshape, location, movement, and non-manual elements. The database is accessible via a graphically based search facility which is highly flexible both in terms of the search options available and the way the results are displayed. LSE-Sign is available at the following website: http://www.bcbl.eu/databases/lse/.
In the speeded word fragment completion task, participants have to complete fragments such as tom{\_}to as quickly and accurately as possible. Previous work has shown that this paradigm can successfully capture subtle priming effects (Heyman, De Deyne, Hutchison, {\&} Storms Behavior Research Methods, 47, 580-606, 2015). In addition, it has several advantages over the widely used lexical decision task. That is, the speeded word fragment completion task is more efficient, more engaging, and easier. Given its potential, we conducted a study to gather speeded word fragment completion norms. The goal of this megastudy was twofold. On the one hand, it provides a rich database of over 8,000 stimuli, which can, for instance, be used in future research to equate stimuli on baseline response times. On the other hand, the aim was to gain insight into the underlying processes of the speeded word fragment completion task. To this end, item-level regression and mixed-effects analyses were performed on the response latencies using 23 predictor variables. Since all items were selected from the Dutch Lexicon Project (Keuleers, Diependaele, {\&} Brysbaert Frontiers in Psychology, 1, 174, 2010), we ran the same analyses on lexical decision latencies to compare the two tasks. Overall, the results revealed many similarities, but also some remarkable differences, which are discussed. We propose that both tasks are complementary when examining visual word recognition. The article ends with a discussion of potential process models of the speeded word fragment completion task.
This article presents K-SPAN (Korean Surface Phonetics and Neighborhoods), a database of surface phonetic forms and several measures of phonological neighborhood density for 63,836 Korean words. Currently publicly available Korean corpora are limited by the fact that they only provide orthographic representations in Hangeul, which is problematic since phonetic forms in Korean cannot be reliably predicted from orthographic forms. We describe the method used to derive the surface phonetic forms from a publicly available orthographic corpus of Korean, and report on several statistics calculated using this database; namely, segment unigram frequencies, which are compared to previously reported results, along with segment-based and syllable-based neighborhood density statistics for three types of representation: an "orthographic" form, which is a quasi-phonological representation, a "conservative" form, which maintains all known contrasts, and a "modern" form, which represents the pronunciation of contemporary Seoul Korean. These representations are rendered in an ASCII-encoded scheme, which allows users to query the corpus without having to read Korean orthography, and permits the calculation of a wide range of phonological measures.
In the present study, we introduce affective norms for a new set of Spanish words, the Madrid Affective Database for Spanish (MADS), that were scored on two emotional dimensions (valence and arousal) and on five discrete emotional categories (happiness, anger, sadness, fear, and disgust), as well as on concreteness, by 660 Spanish native speakers. Measures of several objective psycholinguistic variables—grammatical class, word frequency, number of letters, and number of syllables—for the words are also included. We observed high split-half reliabilities for every emotional variable and a strong quadratic relationship between valence and arousal. Additional analyses revealed several associations between the affective dimensions and discrete emotions, as well as with some psycholinguistic variables. This new corpus complements and extends prior databases in Spanish and allows for designing new experiments investigating the influence of affective content in language processing under both dimensional and discrete theoretical conceptions of emotion. These norms can be downloaded as supplemental materials for this article from www.dropbox.com/s/o6dpw3irk6utfhy/Hinojosa{\%}20et{\%}20al{\_}Supplementary{\%}20materials.xlsx?dl=0.
The two main theoretical accounts of the human affective space are the dimensional perspective and the discrete-emotion approach. In recent years, several affective norms have been developed from a dimensional perspective, including ratings for valence and arousal. In contrast, the number of published datasets relying on the discrete-emotion approach is much lower. There is a need to fill this gap, considering that discrete emotions have an effect on word processing above and beyond those of valence and arousal. In the present study, we present ratings from 1,380 participants for a set of 2,266 Spanish words in five discrete emotion categories: happiness, anger, fear, disgust, and sadness. This will be the largest dataset published to date containing ratings for discrete emotions. We also present, for the first time, a fine-grained analysis of the distribution of words into the five emotion categories. This analysis reveals that happiness words are the most consistently related to a single, discrete emotion category. In contrast, there is a tendency for many negative words to belong to more than one discrete emotion. The only exception is disgust words, which overlap least with the other negative emotions. Normative valence and arousal data already exist for all of the words included in this corpus. Thus, the present database will allow researchers to design studies to contrast the predictions of the two most influential theoretical perspectives in this field. These studies will undoubtedly contribute to a deeper understanding of the effects of emotion on word processing.
Houses have often been used as comparison stimuli in face-processing studies because of the many attributes they share with faces (e.g., distinct members of a basic category, consistent internal features, mono-orientation, and relative familiarity). Despite this, no large, well-controlled databases of photographs of houses that have been developed for research use currently exist. To address this gap, we photographed 100 houses and carefully edited these images. We then asked 41 undergraduate students (18 to 31 years of age) to rate each house on three dimensions: typicality, likeability, and face-likeness. The ratings had a high degree of face validity, and analyses revealed a significant positive correlation between typicality and likeability. We anticipate that this stimulus set (i.e., the DalHouses) and the associated ratings will prove useful to face-processing researchers by minimizing the effort required to acquire stimuli and allowing for easier replication and extension of studies. The photographs of all 100 houses and their ratings data can be obtained at http://dx.doi.org/10.6084/m9.figshare.1279430.
Emotions are highly influential to many psychological processes. Indeed, research employing emotional stimuli is rapidly escalating across the field of psychology. However, challenges remain regarding discrete evocation of frequently co-elicited emotions such as amusement and happiness, or anger and disgust. Further, as much contemporary work in emotion employs college students, we sought to additionally evaluate the efficacy of film clips to discretely elicit these more challenging emotions in a young adult population using an online medium. The internet is an important tool for investigating responses to emotional stimuli, but validations of emotionally evocative film clips across laboratory and web-based settings are limited in the literature. An additional obstacle is identifying stimuli amidst the numerous film clip validation studies. During our investigation, we recognized the lack of a categorical database to facilitate rapid identification of useful film clips for individual researchers' unique investigations. Consequently, here we also sought to produce the first compilation of such stimuli into an accessible and comprehensive catalog. We based our catalog upon prior work as well as our own, and identified 24 articles and 295 film clips from four decades of research. We present information on the validation of these clips in addition to our own research validating six clips using online administration settings. The results of our search in the literature and our own study are presented in tables designed to facilitate and improve a selection of highly valid film stimuli for future research.
This article presents subjective rating norms for a new set of Stills And Videos of facial Expressions-the SAVE database. Twenty nonprofessional models were filmed while posing in three different facial expressions (smile, neutral, and frown). After each pose, the models completed the PANAS questionnaire, and reported more positive affect after smiling and more negative affect after frowning. From the shooting material, stills and 5 s and 10 s videos were edited (total stimulus set = 180). A different sample of 120 participants evaluated the stimuli for attractiveness, arousal, clarity, genuineness, familiarity, intensity, valence, and similarity. Overall, facial expression had a main effect in all of the evaluated dimensions, with smiling models obtaining the highest ratings. Frowning expressions were perceived as being more arousing, clearer, and more intense, but also as more negative than neutral expressions. Stimulus presentation format only influenced the ratings of attractiveness, familiarity, genuineness, and intensity. The attractiveness and familiarity ratings increased with longer exposure times, whereas genuineness decreased. The ratings in the several dimensions were correlated. The subjective norms of facial stimuli presented in this article have potential applications to the work of researchers in several research domains. From our database, researchers may choose the most adequate stimulus presentation format for a particular experiment, select and manipulate the dimensions of interest, and control for the remaining dimensions. The full stimulus set and descriptive results (means, standard deviations, and confidence intervals) for each stimulus per dimension are provided as supplementary material.
Differences between norm ratings collected when participants are asked to consider more than one picture characteristic are contrasted with the traditional methodological approaches of collecting ratings separately for image constructs. We present data that suggest that reporting normative data, based on methodological procedures that ask participants to consider multiple image constructs simultaneously, could potentially confounded norm data. We provide data for two new image constructs, beauty and the extent to which participants encountered the stimuli in their everyday lives. Analysis of this data suggests that familiarity and encounter are tapping different image constructs. The extent to which an observer encounters an object predicts human judgments of visual complexity. Encountering an image was also found to be an important predictor of beauty, but familiarity with that image was not. Taken together, these results suggest that continuing to collect complexity measures from human judgments is a pointless exercise. Automated measures are more reliable and valid measures, which are demonstrated here as predicting human preferences.
Pictures are often used in studies on memory, perception, and language; normative data are thus needed for such visual stimuli. In the present study, we aimed to obtain normative data for a set of 272 black-and-white pictures from middle-aged and elderly Persian speakers. A total of 206 volunteers were divided into two groups: a middle-aged (40-59 years old) group and an elderly (60 years old and over) group. The groups had similar characteristics in terms of education. Norms for every picture were developed to provide measures of name agreement, image agreement, conceptual familiarity, age of acquisition, and visual complexity. The results revealed that all of these measures vary with age, except for conceptual familiarity.
Silent-letter endings are often claimed to be a major source of inconsistency in the French orthography. In this report, we introduce Silex, a database designed to facilitate the study of spelling performance in general, and silent-letter endings in particular. It was derived from two large and recent corpora based on child- and adult-targeted material. Silex consists of three kinds of Excel workbooks: a set of Stimuli Selector workbooks that allow researchers to select words based on a variety of statistics and word characteristics; a Table Generator workbook that allows researchers to build consistency distribution tables by selecting specific phonological or orthographic units; and a Master File workbook, from which all statistics were derived, and that allows researchers to compute other statistics. Silex is different from existing databases in the manner that silent-letter endings were coded and how consistency indices were computed. Importantly, Silex provides unconditional- and conditional-consistency indices for silent-letter endings. To demonstrate the utility of Silex, we first described the silent-letter phenomenon in French. We found that, at minimum, 28 {\%} of French words end with a silent letter. Moreover, silent-letter endings are usually t, e, s, x, or d, and the occurrence of these letters is conditioned by the phonological ending of words. Second, we showed how Silex could prove useful for the development of theoretical models and for empirical studies. The novel information provided in Silex as well as the flexibility of this database should enable researchers to advance our understanding of developing and skilled spelling performance.
Lexical frequency is one of the strongest predictors of word processing time. The frequencies are often calculated from book-based corpora, or more recently from subtitle-based corpora. We present new frequencies based on Twitter, blog posts, or newspapers for 66 languages. We show that these frequencies predict lexical decision reaction times similar to the already existing frequencies, or even better than them. These new frequencies are freely available and may be downloaded from http://worldlex.lexique.org .
We present age-of-acquisition (AoA) ratings for 30,121 English content words (nouns, verbs, and adjectives). For data collection, this megastudy used the Web-based crowdsourcing technology offered by the Amazon Mechanical Turk. Our data indicate that the ratings collected in this way are as valid and reliable as those collected in laboratory conditions (the correlation between our ratings and those collected in the lab from U.S. students reached .93 for a subsample of 2,500 monosyllabic words). We also show that our AoA ratings explain a substantial percentage of the variance in the lexical-decision data of the English Lexicon Project, over and above the effects of log frequency, word length, and similarity to other words. This is true not only for the lemmas used in our rating study, but also for their inflected forms. We further discuss the relationships of AoA with other predictors of word recognition and illustrate the utility of AoA ratings for research on vocabulary growth.
The most important forms of idioms in Chinese, chengyus (CYs), have a fixed length of four Chinese characters. Most CYs are joined structures of two, two-character words—subject–verb units (SVs), verb–object units (VOs), structures of modification (SMs), or verb–verb units—or of four, one-character words. Both the first and second pairs of words in a four-word CY form an SV, a VO, or an SM. In the present study, normative measures were obtained for knowledge, familiarity, subjective frequency, age of acquisition, predictability, literality, and compositionality for 350 CYs, and the influences of the CYs' syntactic structures on the descriptive norms were analyzed. Consistent with previous studies, all of the norms yielded a high reliability, and there were strong correlations between knowledge, familiarity, subjective frequency, and age of acquisition, and between familiarity and predictability. Unlike in previous studies (e.g., Libben {\&} Titone in Memory {\&} Cognition, 36, 1103–1121, 2008), however, we observed a strong correlation between literality and compositionality. In general, the results seem to support a hybrid view of idiom representation and comprehension. According to the evaluation scores, we further concluded that CYs consisting of just one SM are less likely to be decomposable than those with a VOVO composition, and also less likely to be recognized through their constituent words, or to be familiar to, known by, or encountered by users. CYs with an SMSM composition are less likely than VOVO CYs to be decomposable or to be known or encountered by users. Experimental studies should investigate how a CY's syntactic structure influences its representation and comprehension.
Gestures are commonly used together with spoken language in human communication. One major limitation of gesture investigations in the existing literature lies in the fact that the coding of forms and functions of gestures has not been clearly differentiated. This paper first described a recently developed Database of Speech and GEsture (DoSaGE) based on independent annotation of gesture forms and functions among 119 neurologically unimpaired right-handed native speakers of Cantonese (divided into three age and two education levels), and presented findings of an investigation examining how gesture use was related to age and linguistic performance. Consideration of these two factors, for which normative data are currently very limited or lacking in the literature, is relevant and necessary when one evaluates gesture employment among individuals with and without language impairment. Three speech tasks, including monologue of a personally important event, sequential description, and story-telling, were used for elicitation. The EUDICO Linguistic ANnotator (ELAN) software was used to independently annotate each participant's linguistic information of the transcript, forms of gestures used, and the function for each gesture. About one-third of the subjects did not use any co-verbal gestures. While the majority of gestures were non-content-carrying, which functioned mainly for reinforcing speech intonation or controlling speech flow, the content-carrying ones were used to enhance speech content. Furthermore, individuals who are younger or linguistically more proficient tended to use fewer gestures, suggesting that normal speakers gesture differently as a function of age and linguistic performance.
Databases containing lexical properties on any given orthography are crucial for psycholinguistic research. In the last ten years, a number of lexical databases have been developed for Greek. However, these lack important part-of-speech information. Furthermore, the need for alternative procedures for calculating syllabic measurements and stress information, as well as combination of several metrics to investigate linguistic properties of the Greek language are highlighted. To address these issues, we present a new extensive lexical database of Modern Greek (GreekLex 2) with part-of-speech information for each word and accurate syllabification and orthographic information predictive of stress, as well as several measurements of word similarity and phonetic information. The addition of detailed statistical information about Greek part-of-speech, syllabification, and stress neighbourhood allowed novel analyses of stress distribution within different grammatical categories and syllabic lengths to be carried out. Results showed that the statistical preponderance of stress position on the pre-final syllable that is reported for Greek language is dependent upon grammatical category. Additionally, analyses showed that a proportion higher than 90{\%} of the tokens in the database would be stressed correctly solely by relying on stress neighbourhood information. The database and the scripts for orthographic and phonological syllabification as well as phonetic transcription are available at http://www.psychology.nottingham.ac.uk/greeklex/.
Researchers studying a range of psychological phenomena (e.g., theory of mind, emotion, stereotyping and prejudice, interpersonal attraction, etc.) sometimes employ photographs of people as stimuli. In this paper, we introduce the Chicago Face Database, a free resource consisting of 158 high-resolution, standardized photographs of Black and White males and females between the ages of 18 and 40 years and extensive data about these targets. In Study 1, we report pre-testing of these faces, which includes both subjective norming data and objective physical measurements of the images included in the database. In Study 2 we surveyed psychology researchers to assess the suitability of these targets for research purposes and explored factors that were associated with researchers' judgments of suitability. Instructions are outlined for those interested in obtaining access to the stimulus set and accompanying ratings and measures.
Developed a confusion matrix using computer-generated continuous-line uppercase letters. Data were gathered from 20 university students. The letters were presented in 1 of 5 possible positions on the circumference of an imaginary circle at 2.75 deg around the fixation point. The resulting matrix is compared with those of J. T. Townsend (see record 1971-28051-001) and G. C. Gilmore et al (1979) and with D. J. Mewhort and M. L. Dow's analysis of the data of Gilmore et al.
Cognitive theories in visual attention and perception, categorization, and memory often critically rely on concepts of similarity among objects, and empirically require measures of “sameness” among their stimuli. For instance, a researcher may require similarity estimates among multiple exemplars of a target category in visual search, or targets and lures in recognition memory. Quantifying similarity, however, is challenging when everyday items are the desired stimulus set, particularly when researchers require several different pictures from the same category. In this article, we document a new multidimensional scaling database with similarity ratings for 240 categories, each containing color photographs of 16–17 exemplar objects. We collected similarity ratings using the spatial arrangement method. Reports include: the multidimensional scaling solutions for each category, up to five dimensions, stress and fit measures, coordinate locations for each stimulus, and two new classifications. For each picture, we categorized the item's prototypicality, indexed by its proximity to other items in the space. We also classified pairs of images along a continuum of similarity, by assessing the overall arrangement of each MDS space. These similarity ratings will be useful to any researcher that wishes to control the similarity of experimental stimuli according to an objective quantification of “sameness.”
Many experimental research designs require images of novel objects. Here we introduce the Novel Object and Unusual Name (NOUN) Database. This database contains 64 primary novel object images and additional novel exemplars for ten basic- and nine global-level object categories. The objects' novelty was confirmed by both self-report and a lack of consensus on questions that required participants to name and identify the objects. We also found that object novelty correlated with qualifying naming responses pertaining to the objects' colors. The results from a similarity sorting task (and a subsequent multidimensional scaling analysis on the similarity ratings) demonstrated that the objects are complex and distinct entities that vary along several featural dimensions beyond simply shape and color. A final experiment confirmed that additional item exemplars comprised both sub- and superordinate categories. These images may be useful in a variety of settings, particularly for developmental psychology and other research in the language, categorization, perception, visual memory, and related domains.
We present a collection of association norms for 246 German depictable compound nouns and their constituents, comprising 58,652 association tokens distributed over 26,004 stimulus–associate pair types. Analyses of the data revealed that participants mainly provided noun associates, followed by adjective and verb associates. In corpus analyses, co-occurrence values for compounds and their associates were below those for nouns in general and their associates. The semantic relations between compound stimuli and their associates were more often co-hyponymy and hypernymy and less often hyponymy than for associations to nouns in general. Finally, we found a moderate correlation between the overlap of the associations to compounds and their constituents and the degree of semantic transparency. These data represent a collection of associations to German compound nouns and their constituents that constitute a valuable resource concerning the lexical semantic properties of the compound stimuli and the semantic relations between the stimuli and their associates. More specifically, the norms can be used for stimulus selection, hypothesis testing, and further research on morphologically complex words. The norms are available in text format (utf-8 encoding) as supplemental materials.
The Affective Norms for Polish Short Texts (ANPST) dataset (Imbir, 2016d) is a list of 718 affective sentence stimuli with known affective properties with respect to subjectively perceived valence, arousal, dominance, origin, subjective significance, and source. This article examines the reliability of the ANPST and the impact of population type and sex on affective ratings. The ANPST dataset was introduced to provide a recognized method of eliciting affective states with linguistic stimuli more complex than single words and that included contextual information and thus are less ambiguous in interpretation than single word. Analysis of the properties of the ANPST dataset showed that norms collected are reliable in terms of split-half estimation and that the distributions of ratings are similar to those obtained in other affective norms studies. The pattern of correlations was the same as that found in analysis of an affective norms dataset for words based on the same six variables. Female psychology students' valence ratings were also more polarized than those of their female student peers studying other subjects, but arousal ratings were only higher for negative words. Differences also appeared for all other measured dimensions. Women's valence ratings were found to be more polarized and arousal ratings were higher than those made by men, and differences were also present for dominance, origin, and subjective significance. The ANPST is the first Polish language list of sentence stimuli and could easily be adapted for other languages and cultures.
The present study provides norms for Deese/Roediger-McDermott (DRM) lists that were used to create false memories in native speakers of Italian. The word lists reported in this article are based on the DRM lists that have been used extensively to examine illusory memories in English speakers (Deese in Journal of Experimental Psychology, 58, 17-22, 1959; Roediger {\&} McDermott in Journal of Experimental Psychology: Learning, Memory, {\&} Cognition, 21, 803-814, 1995). We translated the 24 critical lures from 24 English DRM lists and created semantically associated Italian word lists that were then normed with native Italian speakers. Overall, the participants recalled 63{\%} of the list items and 22{\%} of the critical lures with the word lists developed. In addition, 56{\%} of the list items and 82{\%} of the critical lures were recognized by the participants. The present study provides a set of Italian lists that can be used by researchers interested in evaluating false memories in Italian-speaking participants.
Since the work of Taft and Forster (1976), a growing literature has examined how English compound words are recognized and organized in the mental lexicon. Much of this research has focused on whether compound words are decomposed during recognition by manipulating the word frequencies of their lexemes. However, many variables may impact morphological processing, including relational semantic variables such as semantic transparency, as well as additional form-related and semantic variables. In the present study, ratings were collected on 629 English compound words for six variables [familiarity, age of acquisition (AoA), semantic transparency, lexeme meaning dominance (LMD), imageability, and sensory experience ratings (SER)]. All of the compound words selected for this study are contained within the English Lexicon Project (Balota et al., 2007), which made it possible to use a regression approach to examine the predictive power of these variables for lexical decision and word naming performance. Analyses indicated that familiarity, AoA, imageability, and SER were all significant predictors of both lexical decision and word naming performance when they were added separately to a model containing the length and frequency of the compounds, as well as the lexeme frequencies. In addition, rated semantic transparency also predicted lexical decision performance. The database of English compound words should be beneficial to word recognition researchers who are interested in selecting items for experiments on compound words, and it will also allow researchers to conduct further analyses using the available data combined with word recognition times included in the English Lexicon Project.
This article presents valence/pleasantness, activity/arousal, power/dominance, origin, subjective significance, and source-of-experience norms for 1,586 Polish words (primarily nouns), adapted from the Affective Norms for English Words list (1,040 words) and from my own previous research (546 words), regarding the duality-of-mind approach for emotion formation. This is a first attempt at creating affective norms for Polish words. The norms are based on ratings by a total of 1,670 college students (852 females and 818 males) from different Warsaw universities and academies, studying various disciplines in equal proportions (humanities, engineering, and social and natural sciences) using a 9-point Likert Self-Assessment Manikin scale. Each participant assessed 240 words on six different scales (40 words per scale) using a paper-and-pencil group survey procedure. These affective norms for Polish words are a valid and useful tool that will allow researchers to use standard, well-known verbal materials comparable to the materials used in other languages (English, German, Portuguese, Spanish, French, Dutch, etc.). The normative values of the Polish adaptation of affective norms are included in the online supplemental materials for this article.
Knowledge of thematic relations is an area of increased interest in semantic memory research because it is crucial to many cognitive processes. One methodological issue that researchers face is how to identify pairs of thematically related concepts that are well-established in semantic memory for most people. In this article, we review existing methods of assessing thematic relatedness and provide thematic relatedness production norming data for 100 object concepts. In addition, 1,174 related concept pairs obtained from the production norms were classified as reflecting one of the five subtypes of relations: attributive, argument, coordinate, locative, and temporal. The database and methodology will be useful for researchers interested in the effects of thematic knowledge on language processing, analogical reasoning, similarity judgments, and memory. These data will also benefit researchers interested in investigating potential processing differences among the five types of semantic relations.
Recent studies suggest that performance attendant on visual word perception is affected not only by feedforward inconsistency (i.e., multiple ways to pronounce a spelling) but also by feedback inconsistency (i.e., multiple ways to spell a pronunciation). In the present study, we provide a statistical analysis of these types of inconsistency for all monosyllabic English words. This database can be used as a tool for controlling, selecting, and constructing stimulus materials for psycholinguistic and neuropsychological research. Such large-scale statistical analyses are necessary devices for developing metrics of inconsistency, for generating hypotheses for psycholinguistic experiments, and for building models of word perception, speech perception, and spelling.
Parent report has proven a valid and cost-effective means of evaluating early child language. Norming datasets for these instruments, which provide the basis for standardized comparisons of individual children to a population, can also be used to derive norms for the acquisition of individual words in production and comprehension and also early gestures and symbolic actions. These lexical norms have a wide range of uses in basic research, assessment and intervention. In addition, cross-linguistic comparisons of lexical development are greatly facilitated by the availability of norms from diverse languages. This report describes the development of CLEX, a new web-based cross-linguistic database for lexical data from adaptations of the MacArthur-Bates Communicative Development Inventories. CLEX provides tools for a range of analyses within and across languages. It is designed to incorporate additional language datasets easily, and to permit users to define mappings between lexical items in pairs of languages for more specific cross-linguistic comparisons.
The subjective familiarity of 40 homophone pairs was examined. The homophones consisted of monosyllabic English words (on one reading) and male first names (on the other)—for example,art andArt. Subjects heard these homophones embedded in two kinds of lists, one with 40 unambiguous words and one with 40 unambiguous names. Ratings were made for familiarity as words and as names. These correlated significantly with the log of printed frequency (.63 for words, .53 for names). In a final task, just the homophones were presented, and the subjects were asked for a comparative rating of whether the word usage or the name usage was more familiar. This direct comparison correlated well (.91) with the difference between the ratings for the name and word familiarities, but less well (.55) with the differences between the printed frequencies of the word and name meanings. This indicates either consistent biases in the judgments or true differences between printed frequencies and subjective familiarity.